Method and device for video processing and medium
By determining the intra-frame mode derivation information of the decoder side of the video unit's neighboring video units for encoding and decoding, the problem of insufficient encoding and decoding efficiency in the existing technology is solved, and more efficient encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202480047205.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-15
- Filing Date
- 2024-07-12
- Publication Date
- 2026-02-13
AI Technical Summary
The efficiency of existing video encoding and decoding technologies needs to be further improved.
By determining the decoder-side intra-frame mode derivation (DIMD) information of the neighboring video units of a video unit and performing encoding and decoding based on this information, encoding and decoding efficiency and performance are improved.
It improves the efficiency and performance of video encoding and decoding.
Smart Images

Figure CN121533009A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to video processing technology, and more particularly, to adaptive derivation of weight parameters for video coding. BACKGROUND
[0002] Nowadays, digital video capability is being applied to various aspects of people's life. Various types of video compression techniques have been proposed for video coding / decoding, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, the coding efficiency of video coding techniques is generally expected to be further improved. SUMMARY
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method comprises: determining, for a conversion between a video unit of a video and a bitstream of the video, decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of the video unit, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode; and performing the conversion based on the DIMD information. In this way, it can improve the coding efficiency and coding performance.
[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions to cause a processor to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by a video processing apparatus. The method comprises: determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of a video unit of a video, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode; and generating the bitstream based on the DIMD information.
[0008] In a fifth aspect, a method for storing a bitstream of a video is presented. The method includes determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of a video unit of the video, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode; generating the bitstream based on the DIMD information; and storing the bitstream in a non-transitory computer-readable medium.
[0009] This Summary is intended to introduce some of the concepts that are further described in the detailed description below. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended for use in limiting the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to the like elements throughout. In the example embodiments of the present disclosure, like reference numerals refer to like elements throughout.
[0011] Figure 1 A block diagram of an example video coding system is shown in accordance with some embodiments of the present disclosure; Figure 2 A block diagram of a first example video encoder is shown in accordance with some embodiments of the present disclosure; Figure 3 A block diagram of an example video decoder is shown in accordance with some embodiments of the present disclosure; Figure 4 An example of an encoder block diagram of VVC is shown; Figure 5 65 intra prediction modes are shown; Figure 6A and Figure 6B Reference samples for wide-angle intra prediction are shown; Figure 7 Discontinuity issues in cases where the direction exceeds 45° are shown; Figure 8A and Figure 8B MMVD search points are shown; Figure 9 A diagram for symmetric MVD mode is shown; Figure 10 Extended CU regions used in BDOF are shown; Figure 11A Control point based affine motion model for a parameter a affine model is shown; Figure 11B Control point based affine motion model for a 6 parameter affine model is shown; Figure 12 affine MVFs are shown for each sub-block; Figure 13 positions of inherited affine motion prediction values are shown; Figure 14 control point motion vector inheritance is shown; Figure 15 positions of candidate positions for constructed affine Merge mode are shown; Figure 16 a schematic diagram of motion vector usage for the proposed combined method is shown; Figure 17 sub-block MVs VSB and pixel delta v(i,j) (arrows) are shown; Figure 18A and Figure 18B SbTMVP process in VVC is shown respectively; Figure 19 local illumination compensation is shown; Figure 20 no down-sampling for short side is shown; Figure 21 decoder-side motion vector refinement is shown; Figure 22 diamond-shaped region in search region is shown; Figure 23 positions of spatial Merge candidates are shown; Figure 24 candidate pairs considered for redundancy check for spatial Merge candidates are shown; Figure 25 a schematic diagram of motion vector scaling for temporal Merge candidates is shown; Figure 26 candidate positions, C0 and C1, for temporal Merge candidates are shown; Figure 27 VVC spatial neighboring blocks of current block are shown; Figure 28 a schematic diagram of virtual blocks in the i-th round of search is shown; Figure 29 an example of GPM partitioning grouped by the same angle is shown; Figure 30 uni-prediction MV selection for geometric partition mode is shown; Figure 31 an example generation of bending weight using geometric partition mode is shown; Figure 32 spatial neighboring blocks used to derive spatial Merge candidates are shown; Figure 33 Template matching is shown to be performed on a search region around the initial MV; Figure 34 A schematic diagram of a subblock for which OBMC is applied is shown; Figure 35 SBT positions, types and transform types are shown; Figure 36 Neighboring samples used for computing SAD are shown; Figure 37 Neighboring samples used for computing SAD for sub-CU level motion information are shown; Figure 38 A sorting process is shown; Figure 39 A reordering process in the encoder is shown; Figure 40 A reordering process in the decoder is shown; Figure 41 A schematic diagram of an extended reference region is shown; Figure 42 An IBC reference region depending on the current CU position is shown; Figure 43 An example of symmetry in screen content pictures is shown; Figure 44A A schematic diagram of BV adjustment for horizontal flipping is shown; Figure 44B A schematic diagram of BV adjustment for vertical flipping is shown; Figure 45 The used intra template matching search region is shown; Figure 46 A schematic diagram of a template region is shown; Figure 47 The spatial part of the convolution filter is shown; Figure 48 The reference region (and its padding) used for deriving filter coefficients is shown; Figure 49 Four Sobel-based gradient modes for GLM are shown; Figure 50 Target samples, template samples and reference samples of the template used in DIMD are shown; Figure 51 The proposed intra block decoding process is shown; Figure 52 HoG computation from a template of 3 pixels width is shown; Figure 53 Prediction fusion by weighted average of two HoG modes and a plane is shown; Figure 54 Spatial GPM candidates are shown; Figure 55 GPM templates are shown; Figure 56 GPM blending is shown; Figures 57A-57I Templates used for fusion are shown, respectively; Figure 58 A flowchart of a method for video processing according to an embodiment of the disclosure is shown; and Figure 59 A block diagram of a computing device in which various embodiments of the disclosure can be implemented is shown.
[0012] In all of the drawings, like or similar reference numerals are generally employed to refer to like or similar elements throughout the several views. DETAILED DESCRIPTION
[0013] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the embodiments are described for illustrative purposes only and help the person skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the ways described below, the disclosure described herein can be implemented in various ways.
[0014] In the following description and claims, unless otherwise defined, all scientific and technical terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0015] References in the present disclosure to “one embodiment”, “an embodiment”, “example embodiments”, etc. indicate that the embodiment described can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that the impact, if any, of such feature, structure, or characteristic relating to other embodiments is well within the knowledge of those skilled in the art.
[0016] It should be understood that although the terms “first” and “second” and the like can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes" and / or "including," when used herein, specify the presence of stated features, elements and / or components, but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0018] Example Environment Figure 1 FIG. 1 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As illustrated, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 can also be referred to as a video encoding device, and destination device 120 can also be referred to as a video decoding device. In operation, source device 110 can be configured to generate encoded video data, and destination device 120 can be configured to decode the encoded video data generated by source device 110. Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 can include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system to generate video data, and / or a combination thereof.
[0020] Video data can include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms a coded representation of the video data. The bitstream can include coded pictures and associated data. A coded picture is a coded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via I / O interface 116 through network 130A. The encoded video data can also be stored onto a storage medium / server 130B for access by destination device 120.
[0021] Destination device 120 can include I / O interface 126, video decoder 124, and display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120, which is configured to interface with an external display device.
[0022] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other existing and / or further standards.
[0023] Figure 2 is a block diagram illustrating an example of a video encoder 200 that can be Figure 1 the example of video encoder 114 in system 100 shown.
[0024] Video encoder 200 can be configured to implement any or all of the techniques of this disclosure. In Figure 2 example, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared between the components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, video encoder 200 can include partitioning unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.
[0026] In other examples, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components, such as motion estimation unit 204 and motion compensation unit 205, can be integrated, these components are shown separately in Figure 2 the example for illustrative purposes.
[0028] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[0029] Mode selection unit 203 can select one of a plurality of coding modes (intra- or inter-coding), e.g., based on the error results, and provide the resulting intra- or inter-coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 can select a combined intra- inter prediction (CIIP) mode, in which the prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 can also select a resolution for a motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0030] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from cache 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from cache 213 other than the picture associated with the current video block.
[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" can refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Further, as used herein, a "P slice" and a "B slice" can refer, in some aspects, to portions of a picture composed of macroblocks that are not dependent on macroblocks in the same picture.
[0032] In some examples, motion estimation unit 204 can perform uni-directional prediction for a current video block, and motion estimation unit 204 can search a reference picture of list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference picture in list 0 or list 1 that contains the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0033] Alternatively, in other examples, the motion estimation unit 204 can perform bi-prediction for the current video block. The motion estimation unit 204 can search the reference pictures in List 0 for a reference video block for the current video block and also search the reference pictures in List 1 for another reference video block for the current video block. The motion estimation unit 204 can then generate reference indices that indicate the reference pictures in List 0 and List 1 that contain the reference video blocks and motion vectors that indicate the spatial displacements between the reference video blocks and the current video block. The motion estimation unit 204 can output the reference indices and the motion vectors for the current video block as the motion information for the current video block. The motion compensation unit 205 can generate the predicted video block for the current video block based on the reference video blocks indicated by the motion information for the current video block.
[0034] In some examples, the motion estimation unit 204 can output a full set of motion information for use in the decoding process at the decoder. Alternatively, in some embodiments, the motion estimation unit 204 can refer to the motion information of another video block to signal the motion information of the current video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0035] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block to the video decoder 300, the value indicating that the current video block has the same motion information as another video block.
[0036] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0038] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0039] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0040] In other examples, such as in a skip mode, there can be no residual data for the current video block, and residual generation unit 207 can not perform the subtraction operation.
[0041] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] Quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block after transform processing unit 208 generates the transform coefficient video blocks associated with the current video block.
[0043] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in buffer 213.
[0044] Loop filtering operations can be performed to reduce video blockiness artifacts in the video block after reconstruction unit 212 reconstructs the video block.
[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0046] Figure 3 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 1 of the video decoder 124 in the system 100 shown.
[0047] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3In examples of the video decoder 300, the video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] In Figure 3 In examples of the video decoder 300, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200.
[0049] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine this information, for example, by performing AMVP and Merge modes. AMVP is used, including deriving a number of most probable candidates based on neighboring PBs' data and reference pictures. The motion information typically includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from spatially or temporally neighboring blocks.
[0050] The motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter to be used with sub-pixel precision can be included in the syntax elements.
[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during encoding of the video block to calculate interpolation for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 from the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 can use at least some of the syntax information to determine the size of blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information for decoding the coded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture, or can also be a region of a picture.
[0053] Intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form a prediction block from spatial neighboring blocks. Dequantization unit 304 dequantizes (i.e., inverse quantizes) quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0054] Reconstruction unit 306 can obtain a decoded block, for example, by adding the residual block to the corresponding prediction block generated by motion compensation unit 302 or intra prediction unit 303. If desired, a deblocking filter can also be applied to filter the decoded block in order to remove blocking artifacts. The decoded video blocks are then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and which also produces decoded video for presentation on a display device.
[0055] Some example embodiments of the present disclosure will be described in detail below. It should be understood that the use of section headings in this document is for ease of understanding only and is not to be used to limit the embodiments disclosed in that section only to that section. Furthermore, although certain embodiments are described with reference to a multi-function video codec or other specific video codec, the disclosed techniques are applicable to other video coding technologies as well. Moreover, although some embodiments describe video coding steps in detail, it will be appreciated that a decoder will implement corresponding decoding steps that undo the coding steps. Furthermore, the term "video processing" includes video coding or compression, video decoding or decompression, and video transcoding, where video pixels are represented from one compression format to another or at different compression bit rates.
[0056] 1. BRIEF OVERVIEW This disclosure relates to video encoding and decoding technologies. Specifically, it relates to the derivation of weighting parameters and how to apply them to encoding and decoding tools, as well as other encoding and decoding tools in image / video encoding and decoding. It can be applied to existing video encoding and decoding standards such as HEVC or VVC. It is also applicable to future video encoding and decoding standards or video codecs.
[0057] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between ITU-T VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, aiming to reduce the bitrate by 50% compared to HEVC. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5) are investigating the potential need for standardization of future video codec technologies with compression capabilities significantly exceeding the current VVC standard. Such future standardization efforts could take the form of multiple additional extensions to VVC or entirely new standards. These groups are jointly conducting this exploration in a collaborative effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. New codec features and coding methods implemented in Enhanced Compression Model (ECM) software, as potential enhanced video codec technologies exceeding VVC capabilities, are being explored in a coordinated manner by the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG.
[0058] 2.1. Encoding and decoding process of a typical video codec Figure 4An example of the encoder block diagram of VVC is shown, which contains three in-loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF which uses a pre-defined filter, SAO and ALF utilize the original samples of the current picture, by adding an offset and by applying a Finite Impulse Response (FIR) filter, respectively, and utilize coded side information to signal the offset and filter coefficients to reduce the mean square error between the original and the reconstructed samples. ALF is located at the last processing stage of each picture and can be considered as a tool that tries to capture and fix artifacts caused by previous stages.
[0059] 2.2. Intra mode coding with 67 intra prediction modes To capture arbitrary edge directions presented in natural videos, as Figure 5 shown, the number of directional intra modes is extended from 33 used in HEVC to 65, and the planar mode and DC mode remain the same. These denser directional intra prediction modes are applied to all block sizes and for both luma intra prediction and chroma intra prediction.
[0060] In HEVC, each intra coded block has a square shape and its each side has a length of power of 2. Therefore, no partitioning operation is needed to generate an intra prediction value using the DC mode. In VVC, a block can have a rectangular shape and in general case a partitioning operation is needed for each block. To avoid the partitioning operation for DC prediction, only the longer side is used to calculate the average value for non-square blocks.
[0061] 2.2.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. The regular angular intra prediction directions are defined as clockwise direction from 45 degrees to -135 degrees. In VVC, for non-square blocks, several regular angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of wide-angle modes after parsing. The total number of intra prediction modes remains the same, i.e., 67, and the intra mode coding method remains the same.
[0062] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 6A and Figure 6B shown.
[0063] The number of replaced modes in wide-angle directional modes depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 1.
[0064] Table 1 - Intra prediction modes replaced by wide-angle modes
[0065] As Figure 7 shown in FIG. 1, in the case of wide-angle intra prediction, two vertically neighboring prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap p α If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that satisfy this condition, i.e., [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted by these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the regular prediction modes with the non-fractional mode design in the wide-angle modes.
[0066] In VVC, 4:2:2 and 4:4:4 chroma formats are supported in addition to 4:2:0. The chroma derivation mode (DM) derivation table for 4:2:2 chroma format is originally ported from HEVC, which extends the number of entries from 35 to 67 to align with the extension of intra prediction modes. Since HEVC specification does not support prediction angles lower than -135 degrees and higher than 45 degrees, the luma intra prediction modes with range from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for 4:2:2 chroma format is updated by replacing some values of the entries of the mapping table to convert the prediction angles for chroma blocks more accurately.
[0067] 2.3. Inter prediction For each inter prediction CU, the motion parameters consist of motion vectors, reference picture indices and reference picture list usage indices, and additional information required for inter prediction sample generation using new coding features of VVC. The motion parameters can be signaled in an explicit or implicit way. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. Merge mode is specified, in which the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates and additional scheduling introduced in VVC. Merge mode can be applied to any inter prediction CU, not only for skip mode. An alternative to Merge mode is the explicit transmission of motion parameters, in which the motion vectors, the corresponding reference picture indices for each reference picture list and the reference picture list usage flags and other required information are explicitly signaled for each CU.
[0068] 2.4. Intra block copy (IBC) Intra block copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed inside the current picture. The luma block vector of an IBC coded CU is in integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pel motion vector precision and 4-pel motion vector precision. An IBC coded CU is considered as a third prediction mode different from intra or inter prediction modes. IBC mode is applicable to a CU whose width and height are less than or equal to 64 luma samples.
[0069] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check on a block whose width or height is not larger than 16 luma samples. For non-Merge mode, the block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a block matching based local search will be performed.
[0070] In hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key computation is based on 4x4 sub-blocks for each location in the current picture. For a larger size current block, the hash key is determined to match the hash key of a reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference locations. If multiple reference blocks' hash keys are found to match the hash key of the current block, the block vector cost of each matched reference is computed and the one with the minimum cost is selected.
[0071] In block matching search, the search range is set to cover both the previous CTU and the current CTU.
[0072] At the CU level, the IBC mode is signaled with a flag and it can be signaled as IBC AMVP mode or IBC skip / Merge mode as follows: - IBC skip / Merge mode: a Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC coded blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates and paired candidates.
[0073] - IBC AMVP mode: Block vector difference is coded in the same way as motion vector difference. Block vector prediction method uses two candidates as the predictor, one from the left neighboring block and one from the above neighboring block (if IBC coded). When either neighboring block is unavailable, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.
[0074] 2.5. IBC motion candidates The term“block” can represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB, or a video processing unit comprising a plurality of samples / pixels. A block can be rectangular or non-rectangular.
[0075] For IBC coded blocks, a block vector (BV) is used to indicate the displacement from the current block to a reference block that has already been reconstructed inside the current picture.
[0076] W and H are the width and height of the current block (e.g., luma block).
[0077] The non-adjacent spatial candidate of the current coding block is the adjacent spatial candidate of the virtual block in the i-th round of search (as shown in FIG. 1). For the i-th round of search, the width and height of the virtual block are calculated by the following equations: newWidth = i x 2 x gridX + W, newHeight = i x 2 x gridY + H. Obviously, if the search round i is 0, the virtual block is the current block. Figure 9
[0078] In the following, the BV predictor is also a BV candidate. The skip mode is also the Merge mode.
[0079] The BV candidates can be divided into several groups according to some criteria. Each group is called a sub-group. For example, the adjacent spatial and temporal BV candidates can be as a first sub-group, and the remaining BV candidates can be as a second sub-group; in another example, it can also put the first N (N > 2) BV candidates as a first sub-group, the following M (M > 2) BV candidates as a second sub-group, and the remaining BV candidates as a third sub-group.
[0080] 2.6. Merge mode with MVD (MMVD) In addition to the Merge mode (in which the implicitly derived motion information is directly used for the prediction sample generation of the current CU), the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag to specify whether the MMVD mode is used for the CU.
[0081] In MMVD, after a Merge candidate is selected, it is further refined by the signaled MVD information. The further information includes a Merge candidate flag, an index that specifies the motion amplitude, and an index that indicates the motion direction. In MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. A MMVD candidate flag is signaled to specify which one is used between the first Merge candidate and the second Merge candidate.
[0082] The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point. As shown in Figure 8A and Figure 8B , the offset is added to the horizontal component or the vertical component of the starting MV. The relationship of the distance index to the predefined offset is specified as Table 2.
[0083] Table 2 - Relationship of distance index to predefined offset
[0084] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 3. It should be noted that the meaning of the MVD sign can change depending on the information of the starting MV. When the starting MV is a uni-prediction MV or bi-prediction MV, and both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the sign designation in Table 3 is added to the sign of the MV offset of the starting MV. When the starting MV is a bi-prediction MV, and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other reference POC is less than the POC of the current picture), and the difference of POCs in list 0 is greater than the difference of POCs in list 1, the sign designation in Table 3 is added to the sign of the MV offset of the list 0 MV component of the starting MV, and the sign for the list 1 MV has the opposite value. Otherwise, if the difference of POCs in list 1 is greater than the difference of POCs in list 0, the sign designation in Table 3 is added to the sign of the MV offset of the list 1 MV component of the starting MV, and the sign for the list 0 MV has the opposite value.
[0085] The MVD is scaled according to the difference of POCs in each direction. If the difference of POCs in the two lists is the same, no scaling is needed. Otherwise, if the difference of POCs in list 0 is greater than the difference of POCs in list 1, the MVD of list 1 is scaled by defining the POC difference of L0 as td, and the POC difference of L1 as tb, as Figure 9The MVD of List 0 is scaled in the same way if the POC difference of L1 is greater than the POC difference of L0. If the starting MV is uni-predicted, the MVD is added to the available MV.
[0086] Table 3 - Sign of MV offset specified by direction index
[0087] 2.7. Symmetric MVD coding In VVC, in addition to normal uni-predicted and bi-predicted mode MVD signaling, a symmetric MVD mode for bi-predicted MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indices of both List 0 and List 1 and the MVD of List 1 are not signaled but derived.
[0088] The decoding process of the symmetric MVD mode is as follows: 1. At the slice level, the variables BiDirPredFlag, RefIdxSymL0 and RefIdxSymL1 are derived as follows: - If mvd_l1_zero_flag is equal to 1, BiDirPredFlag is set equal to 0.
[0089] - Otherwise, if the closest reference picture in List 0 and the closest reference picture in List 1 form a pair of forward and backward reference pictures or a pair of backward and forward reference pictures, BiDirPredFlag is set to 1 and both List 0 and List 1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.
[0090] 2. At the CU level, if the CU is bi-predicted coded and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used or not is explicitly signaled.
[0091] When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. The reference indices of List 0 and List 1 are set equal to the pair of reference pictures, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown as the following equation.
[0092]
[0093] In the encoder, symmetric MVD motion estimation starts from the initial MV estimation. A set of initial MV candidates, including the MV obtained from the uni-prediction search, the MV obtained from the bi-prediction search, and the MV from the AMVP list. One with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.
[0094] 2.8. Bi-directional optical flow (BDOF) The bi-directional optical flow (BDOF) tool is included in VVC. BDOF (previously called BIO) is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version, which requires much less computation, especially in terms of the number of multiplications and the size of the multiplier.
[0095] BDOF is used to refine the bi-predicted signal of a CU at the 4x4 subblock level. BDOF is applied to a CU if it satisfies all the following conditions: - The CU is coded using the "true" bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order.
[0096] - The distance (i.e., POC difference) from the two reference pictures to the current picture is the same.
[0097] - Both reference pictures are short-term reference pictures.
[0098] - The CU is not coded using the affine mode or the SbTMVP Merge mode.
[0099] - The CU has more than 64 luma samples.
[0100] - The CU height and CU width are both greater than or equal to 8 luma samples.
[0101] - The BCW weight index indicates equal weights.
[0102] - WP is not enabled for the current CU.
[0103] - CIIP mode is not used for the current CU.
[0104] BDOF is only applied to the luma component. As the name indicates, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4x4 subblock, the motion refinement is computed by minimizing the difference between the L0 predicted samples and the L1 predicted samples. The motion refinement is then used to adjust the bi-predicted sample values in the 4x4 subblock. The following steps are applied in the BDOF process.
[0105] First, the horizontal and vertical gradients of the two prediction signals and , It is calculated by directly calculating the difference between two neighboring sample points, that is,
[0106] in It is a list The predicted signal in the coordinates The sample value at that location, Furthermore, shift1 is calculated based on the luminance bit depth bitDepth as shift1 = max(6, bitDepth-6).
[0107] Then, the autocorrelation and cross-correlation of the gradients. , , , and Calculated as
[0108] in
[0109] in It is a 6x6 window surrounding a 4x4 sub-block, and and The values are set to min(1, bitDepth) respectively. 11) and min(4, bitDepth) 8).
[0110] Motion refinement Then, the following formula is derived using cross-correlation and autocorrelation terms:
[0111] in , , . It is a floor function, and .
[0112] Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4×4 sub-block:
[0113] Finally, the BDOF samples of CU are calculated by adjusting the bidirectional prediction samples as follows:
[0114] These values are chosen such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit width of intermediate parameters in the BDOF process is kept within 32 bits.
[0115] To derive gradient values, some of the prediction samples outside the current CU boundary in List ) need to be generated. As shown in Figure 10 BDOF in VVC uses one extended row / column around the CU boundary. To control the computational complexity of generating the out-of-boundary prediction samples, the prediction samples in the extended area (white positions) are generated by directly taking the reference samples at the nearby integer positions (using floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the prediction samples inside the CU (gray positions). These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample and gradient values outside the CU boundary are needed, they are padded (i.e., repeated) from their nearest neighbors.
[0116] When the width and / or height of a CU is larger than 16 luma samples, it is divided into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries are treated as the CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is smaller than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8 * W * (H >> 1)), where W indicates the sub-block width and H indicates the sub-block height. To avoid the extra complexity of SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here.
[0117] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bi-directional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is one for either of the two reference pictures, BDOF is also disabled. BDOF is also disabled when the CU is coded with symmetric MVD mode or CIIP mode.
[0118] 2.9. Inter-intra joint prediction (CIIP) 2.10. Affine motion compensation prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 11A and Figure 11B As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0119] For a 4-parameter affine motion model, the location of the sample point in the block ( x, y The motion vector at point () is derived as:
[0120] For a 6-parameter affine motion model, the location of the sample points in the block ( x, y The motion vector at point () is derived as:
[0121] in( mv 0x , mv 0y ) is the motion vector of the upper left control point, ( mv 1x , mv 1y ) is the motion vector of the upper right control point, and ( mv 2x , mv 2y ) is the motion vector of the lower left control point.
[0122] To simplify motion compensation prediction, block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation (e.g., ...). Figure 12 (As shown), and rounded to 1 / 16 fractional precision. Then a motion-compensated interpolation filter is applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks.
[0123] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.
[0124] 2.10.1. Affine Merge Prediction The AF_MERGE mode can be applied to a CU whose width and height are both greater than or equal to 8. In this mode, the CPMVs of the current CU are generated based on the motion information of spatial neighboring CUs. There can be up to five CPMV candidates, and one to be used for the current CU is indicated by a signaled index. The following three types of CPMV candidates are used to form the affine merge candidate list: - Inherited affine merge candidate inferred from the CPMVs of neighboring CUs.
[0125] - Constructed affine merge candidate CPMVP derived using the translational MVs of neighboring CUs.
[0126] - Zero MV.
[0127] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks, one from the left neighboring CU and one from the top neighboring CU. The candidate blocks are shown in Figure 13 . For the left prediction, the scan order is A0->A1, and for the top prediction, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No de-duplication check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidates in the affine merge list of the current CU. As shown, if the neighboring bottom-left block A is coded in affine mode, the motion vectors of the top-left, top-right, and bottom-left corners of the CU containing block A , and are obtained. When block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated according to and . When block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated according to , and . Figure 14 Inherited control point motion vectors are shown.
[0128] Constructed affine candidate means that the candidate is constructed by combining the neighboring translational motion information of each control point. The motion information for a control point is derived from the specified spatial and temporal neighbors as shown in Figure 15 . The CPMVs k(k = 1, 2, 3, 4) denotes the k-th control point. For CPMV1, the B2->B3->A2 block is checked and the MV of the first available block is used. For CPMV2, the B1->B0 block is checked and for CPMV3, the A1->A0 block is checked. If TMVP is available, it is used as CPMV4.
[0129] After obtaining the MVs of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, { CPMV1, CPMV2}, { CPMV1, CPMV3}.
[0130] The combinations of 3 CPMVs construct 6-parameter affine Merge candidates and the combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the related combination of control point MVs is discarded.
[0131] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list.
[0132] 2.10.2. Affine AMVP prediction The affine AMVP mode can be applied to a CU whose width and height are both greater than or equal to 16. An affine flag at CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used or not, and then another flag is signaled to indicate whether 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMVs of the current CU and their prediction values CPMVPs are signaled in the bitstream. The affine AMVP candidate list size is 2 and is generated by using the following four types of CPVM candidates in order: - inherited affine AMVP candidates inferred from the CPMVs of neighboring CUs; - constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs; - translational MVs from neighboring CUs; - zero MV.
[0133] The checking order of inherited affine AMVP candidates is the same as that of inherited affine Merge candidates. The only difference is that for AMVP candidates, only affine CUs with the same reference picture as in the current block are considered. No de-duplication process is applied when inserting inherited affine motion predictors into the candidate list.
[0134] Constructed AMVP candidates are derived from Figure 15 the specified spatial neighbors as shown in FIG. 6. The same checking order as done in affine Merge candidate construction is used. In addition, the reference picture index of the neighboring blocks is also checked. The first block in the checking order that is inter coded and has the same reference picture as in the current CU is used. Only when the current CU is coded with 4-parameter affine mode, and mv 0 and mv 1 are available, they are added as one candidate in the affine AMVP list. When the current CU is coded with 6-parameter affine mode, and all three CPMVs are available, they are added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable.
[0135] If the affine AMVP list candidates are still less than 2 after the inherited affine AMVP candidates and the constructed AMVP candidate are checked, when available, mv 0, mv 1 and mv 2 will be added in order as translational MVs to predict all control point MVs of the current CU. Finally, if the affine AMVP list is still not full, a zero MV is used to fill the affine AMVP list.
[0136] 2.10.3. Affine Motion Information Storage In VVC, the CPMVs of an affine CU are stored in a separate buffer. The stored CPMVs are only used to generate inherited CPMVPs for the most recently coded CU in affine Merge mode and affine AMVP mode. The sub-block MVs derived from the CPMVs are used for motion compensation, MV derivation for the translational MVs of the Merge / AMVP list and deblocking.
[0137] To avoid picture row buffering for additional CPMVs, the affine motion data inheritance from a CU of the above CTU is handled differently from that from regular neighboring CUs. If the candidate CU for affine motion data inheritance is in the above CTU row, the bottom-left and bottom-right sub-block MVs in the row buffer, not the CPMVs, are used for affine MVP derivation. In this way, the CPMVs are only stored in the local buffer. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to the 4-parameter model. As Figure 16As shown, along the top CTU boundary, the motion vectors of the lower left and lower right sub-blocks of the CU are used for the affine inheritance of the CU in the bottom CTU.
[0138] 2.10.4. Refinement of Optical Flow Prediction for Affine Modes Compared to pixel-based motion compensation, sub-block-based affine motion compensation saves memory access bandwidth and reduces computational complexity, at the cost of reduced prediction accuracy. To achieve finer-grained motion compensation, Prediction Refinement (PROF) using optical flow is used to refine sub-block-based affine motion compensation predictions without increasing memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the difference derived from the optical flow equation. PROF is described in the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate sub-block predictions. .
[0139] Step 2) Calculate the spatial gradient of the sub-block prediction at each sample location using a 3-tap filter [-1, 0, 1]. and The gradient calculation is exactly the same as the gradient calculation in BDOF.
[0140]
[0141]
[0142] This is used to control the precision of the gradient. The sub-block (i.e., 4x4) prediction is expanded by one sample point on each side for gradient computation. To avoid additional memory bandwidth and additional interpolation computation, these expanded samples on the expanded boundaries are copied from the nearest integer pixel position in the reference image.
[0143] Step 3) Brightness prediction refinement is calculated using the following optical flow equation.
[0144]
[0145] in It refers to the location of the sample points. The calculated sample MV (by (representation) and sample points The difference between the MVs of the sub-blocks of the same sub-block, such as Figure 17 As shown. It is quantized in units of 1 / 32 brightness sample precision.
[0146] Since the affine model parameters and the sample point positions relative to the sub-block center do not change from one sub-block to another, therefore may be computed for the first sub-block and repeated for other sub-blocks in the same CU. Let and be the horizontal and vertical offsets from the sample position to the center of the sub-block , may be derived by the following equations:
[0147] To maintain accuracy, the center of the sub-block is computed as ( ( W SB 1 ) / 2, ( H SB 1 ) / 2), where W SB and H SB are the sub-block width and height.
[0148] For the 4-parameter affine model,
[0149] For the 6-parameter affine model,
[0150] where , , are the top-left, top-right and bottom-left control point motion vectors, and are the width and height of the CU.
[0151] Step 4) Finally, the luma prediction refinement is added to the sub-block prediction . The final prediction I’ is generated as the following equation.
[0152]
[0153] For affine-coded CUs, PROF is not applied in two cases: 1) all control point MVs are the same, which indicates the CU has only translational motion; 2) the affine motion parameters are larger than a specified limit, because the sub-block based affine MC is downgraded to CU-based MC to avoid large memory access bandwidth requirement.
[0154] A fast coding method is applied to reduce the coding complexity of affine motion estimation with PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) if the CU is not a root block and its parent block does not select affine mode as its best mode, PROF is not applied because the probability that the current CU selects affine mode as the best mode is low; b) if the magnitudes of the four affine parameters (C, D, E, F) are all smaller than a pre-defined threshold and the current picture is not a low-delay picture, PROF is not applied because the improvement introduced by PROF is small in this case. In this way, the affine motion estimation with PROF can be accelerated.
[0155] 2.11. Subblock-based Temporal Motion Vector Prediction (SbTMVP) VVC supports a subblock-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in a collocated picture to improve the motion vector prediction for CUs and the Merge mode in the current picture. The same collocated picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; While TMVP obtains the temporal motion vector from a collocated block in the collocated picture (the collocated block is the right-bottom or center block relative to the current CU), SbTMVP applies a motion shift before obtaining the temporal motion information from the collocated picture, where the motion shift is obtained from a motion vector of one of the spatial neighboring blocks from the current CU.
[0156] The SbTVMP process is shown in Figure 18A and Figure 18B SbTMVP predicts the motion vector of a sub-CU within the current CU in two steps. In the first step, the spatial near-neighbor A1 in Figure 18A is checked. If A1 has a motion vector that uses the collocated picture as its reference picture, this motion vector is selected as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0, 0).
[0157] In the second step, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the motion information (motion vector and reference index) at the sub-CU level from the collocated picture as shown in Figure 18B Figure 18B The example in the figure assumes that the motion shift is set to the motion of block Al. Then, for each sub-CU, the motion information of its corresponding block in the collocated picture (covering the smallest motion grid of the center sample) is used to derive the motion information for the sub-CU. After the motion information of the collocated sub-CU is identified, it is converted to the motion vector and reference index of the current sub-CU in a similar way as the TMVP process of HEVC, where the temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0158] In VVC, a subblock-based combined Merge list containing both SbTMVP candidates and affine Merge candidates is used for the signaling of the subblock-based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry of the list of subblock-based Merge candidates, followed by the affine Merge candidates. The size of the subblock-based Merge list is signaled in the SPS, and the maximum allowed size of the subblock-based Merge list is 5 in VVC.
[0159] The sub-CU size used in SbTMVP is fixed to 8x8, and as with the affine Merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8.
[0160] The coding logic of the additional SbTMVP Merge candidate is the same as that of other Merge candidates, i.e., for each CU in a P slice or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.
[0161] 2.12. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of a CU and the predicted motion vector) is signaled in quarter-luma sample units. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be coded at different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows: - Normal AMVP mode: quarter-luma sample, half-luma sample, integer-luma sample, or double-luma sample.
[0162] - Affine AMVP mode: quarter-luma sample, integer-luma sample, or 1 / 16-luma sample.
[0163] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both horizontal and vertical MVDs for reference list L0 and reference list L1) are zero, the quarter luma sample MVD resolution is inferred.
[0164] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is needed and quarter luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half luma sample or other MVD precision (integer or quarter luma sample) is used for normal AMVP CUs. In the case of half luma sample, a 6-tap interpolation filter is used for the half luma sample position instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer luma sample MVD precision or quarter luma sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer luma sample MVD precision or 1 / 16 luma sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter luma sample, half luma sample, integer luma sample, or quarter luma sample), the motion vector predictor of the CU will be rounded to the same precision as the MVD before being added with the MVD. The motion vector predictor is rounded to approach zero (i.e., negative motion vector predictor is rounded to approach positive infinity and positive motion vector predictor is rounded to approach negative infinity).
[0165] The encoder uses RD check to determine the motion vector resolution for the current CU. To avoid performing CU-level RD check four times for each MVD resolution, in VTM11, the RD check for MVD precision other than quarter luma sample is only conditionally invoked. For normal AMVP mode, the RD cost of quarter luma sample MVD precision and integer luma sample MV precision are first calculated. Then, the RD cost of integer luma sample MVD precision is compared with that of quarter luma sample MVD precision to decide whether it is necessary to further check the RD cost of quarter luma sample MVD precision. When the RD cost of quarter luma sample MVD precision is much smaller than that of integer luma sample MVD precision, the RD check of quarter luma sample MVD precision is skipped. Then, the check of half luma sample MVD precision is skipped if the RD cost of integer luma sample MVD precision is significantly larger than the best RD cost of previously tested MVD precisions. For affine AMVP mode, if no affine inter mode is selected after checking the rate-distortion cost of affine Merge / skip mode, Merge / skip mode, quarter luma sample MVD precision normal AMVP mode and quarter luma sample MVD precision affine AMVP mode, the 1 / 16 luma sample MV precision and 1-pixel MV precision affine inter modes are not checked. In addition, in 1 / 16 luma sample and quarter luma sample MV precision affine inter modes, the affine parameters obtained in quarter luma sample MV precision affine inter mode are used as the starting search points.
[0166] 2.13. Bi-prediction with CU-level weights (BCW) In HEVC, bi-predicted signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, bi-prediction mode is extended beyond simple averaging to allow weighted averaging of two prediction signals.
[0167]
[0168] Five weights are allowed in weighted average bi-prediction, For each bi-predicted CU, the weight w is determined in one of two ways: 1) for non-Merge CU, the weight index is signaled after the motion vector difference; 2) for Merge CU, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w e {3, 4, 5}) are used.
[0169] - At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. For more details, the reader is referred to the VTM software. When combined with AMVR, if the current picture is a low-delay picture, the unequal weights are conditionally checked for 1-pel and 4-pel motion vector precision.
[0170] - When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode.
[0171] - When the two reference pictures in bi-prediction are the same, the unequal weights are conditionally checked.
[0172] - Depending on the POC distance between the current picture and its reference pictures, the coded QP and the temporal level, the unequal weights are not searched when certain conditions are met.
[0173] The BCW weight index is coded using one context-coded bin followed by bypass-coded bins. The first context-coded bin indicates whether equal weights are used; and if unequal weights are used, the additional bins are signaled using bypass coding to indicate the unequal weights used.
[0174] Weighted prediction (WP) is a coding tool supported by H.264 / AVC and HEVC standards for efficiently coding video content with fade-in and fade-out. The support of WP is also added to the VVC standard. WP allows the signaling of a weighting parameter (weight and offset) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weight(s) and offset(s) of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid the interaction between WP and BCW, which would complicate the VVC decoder design, if a CU uses WP, the BCW weight index is not signaled and w is assumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is assumed from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV.
[0175] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is encoded and decoded using CIIP mode, the BCW index of the current CU is set to 2, for example, with equal weights.
[0176] 2.14. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is an encoding / decoding tool that addresses the problem of local illumination variations between the current image and its temporal reference image. LIC is based on a linear model, where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples for the current block. Specifically, LIC can be mathematically modeled by the following equation:
[0177] in The current block is at coordinates Predicted signal at the location; It is composed of motion vectors The reference block it points to; and These are the corresponding scaling factors and offsets applied to the reference block. Figure 19 The LIC procedure is shown. Figure 19 In this context, when applying LIC to a block, the Minimum Mean Square Error (LMSE) method is employed to minimize the number of neighboring samples (i.e., ...) of the current block. Figure 19 templates in T ) and its corresponding reference sample in the time-domain reference image (i.e. Figure 19 In T0 or T1 The difference between ) is used to derive the LIC parameters (i.e. and The value of ). Furthermore, to reduce computational complexity, both the template samples and the reference template samples are downsampled (adaptive downsampling) to derive the LIC parameters; that is, only Figure 19 The shaded samples in the data were used for derivation. and .
[0178] To improve encoding and decoding performance, such as Figure 20 As shown, downsampling is not performed on the shorter side.
[0179] 2.15. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the motion vector refinement (MV) in the Merge mode, a decoder-side motion vector refinement based on bilateral matching (BM) is applied in the VVC. In the bidirectional prediction operation, a refined MV is searched around the initial MV in reference image lists L0 and L1. The BM method computes the distortion between two candidate blocks in reference image lists L0 and L1. Figure 21As shown, the SAD between two blocks based on each MV candidate (e.g., MV0' and MV1') around the initial MV is computed. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predicted signal.
[0180] In VVC, the application of DMVR is restricted and is only applied to CUs coded with the following modes and features: - CU-level Merge mode with bi-predicted MVs; - One reference picture is past and the other reference picture is future with respect to the current picture; - The distance (i.e., POC difference) from the two reference pictures to the current picture is the same; - Both reference pictures are short-term reference pictures; - The CU has more than 64 luma samples; - The CU height and CU width are both greater than or equal to 8 luma samples; - The BCW weight index indicates equal weights; - WP is not enabled for the current block; - CIIP mode is not used for the current block.
[0181] The refined MV derived by the DMVR process is used to generate the inter- predicted samples and is also used for temporal motion vector prediction for future picture coding. While the original MV is used in the deblocking process and is also used for spatial motion vector prediction for future CU coding.
[0182] Additional features of DMVR are mentioned in the following sub-items.
[0183] 2.15.1. Search scheme In DVMR, the search points are around the initial MV and the MV offsets obey the MV difference mirroring rule. In other words, any point (denoted by the candidate MV pair (MV0, MV1)) examined by DMVR obeys the following two equations:
[0184]
[0185] where denotes the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples away from the initial MV. The search consists of an integer sample offset search stage and a fractional sample refinement stage.
[0186] A 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainties, a bias towards the original MV is proposed during the DMVR process. The SAD value between reference blocks of the initial MV candidate references is reduced by 1 / 4.
[0187] The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the surface equation of the parameter error, rather than through an additional search with SAD comparison. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration with the minimum SAD at the center.
[0188] In subpixel offset estimation based on parametric error surfaces, the cost at the center location and the costs at the four nearest neighbor locations are used to fit a two-dimensional parabolic error surface equation of the following form:
[0189] in( This corresponds to the score position with the minimum cost, and C corresponds to the minimum cost. The above equation is solved by using the costs of the five search points. Calculated as:
[0190]
[0191] and The value is automatically constrained between -8 and 8 because all values are positive, and the minimum value is... This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated score ( It is added to integer distance thinning MV to obtain subpixel accurate thinning increment MV.
[0192] 2.15.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 of a lumen sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search point surrounds the initial fractional pixel MV with an integer sample offset, so samples at those fractional positions need to be interpolated for the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that, utilizing a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not needed by the interpolation process based on the original MV but are needed by the interpolation process based on the refined MV are filled from these available samples.
[0193] 2.15.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luminance samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luminance samples. The maximum cell size for the DMVR search process is limited to 16x16.
[0194] 2.16. Multi-pass decoder-side motion vector refinement In this contribution, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector predictions.
[0195] 2.16.1. First pass – Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying the BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.
[0196] The BM performs a local search to derive integer sample precision intDeltaMV and half-pel sample precision halfDeltaMv. The local search applies a 3x3 square search pattern that loops through the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum values of sHor and sVer are 8.
[0197] The bilateral matching cost is computed as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between the reference blocks. The intDeltaMV or halfDeltaMV local search is terminated when the bilCost at the center point of the 3x3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern and the search for the minimum cost continues until it reaches the end of the search range.
[0198] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: MV0_pass1 = MV0 + deltaMV MV1_pass1 = MV1 - deltaMV.
[0199] 2.16.2. Second pass - sub-block based bilateral matching MV refinement In the second pass, the refined MV is derived by applying BM to 16x16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass for the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.
[0200] For each sub-block, the BM performs a full search to derive integer sample precision intDeltaMV. The search range of the full search is [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum values of sHor and sVer are 8.
[0201] The bilateral matching cost is computed by applying a cost factor to the SATD cost between the two reference sub-blocks as follows: bilCost = satdCost * costFactor. The search region (2*sHor + 1) * (2*sVer + 1) is divided into Figure 22 up to 5 diamond search regions as shown. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed in order starting from the center of the search region. In each region, the search points are processed in a raster scan order from the top-left corner to the bottom-right corner of the region. When the minimum bilCost within the current search region is less than or equal to the threshold of sbW * sbH, the integer-pel full search is terminated, otherwise, the integer-pel full search continues to the next search region until all search points are checked.
[0202] The BM performs a local search to derive the half-pel precision halfDeltaMv. The search pattern and cost function are the same as defined in 2.9.1.
[0203] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). The refined MV at the second pass is then derived as: MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2) MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2).
[0204] 2.16.3. Third pass - subblock-based bilateral optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8x8 grid subblocks. For each 8x8 subblock, the BDOF refinement is applied to derive the unclipped scaled Vx and Vy from the refined MV of the parent block at the second pass, bioMv(Vx, Vy). The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0205] The refined MV at the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as: MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv MV1_pass3(sbldx3) = MV0_pass2(sbldx2) - bioMv.
[0206] 2.17. Sample-based BDOF In sample-based BDOF, instead of deriving motion refinements (Vx, Vy) on a block basis, it is performed on a sample-by-sample basis.
[0207] A coded block is divided into 8x8 sub-blocks. For each sub-block, whether to apply BDOF is determined by comparing the SAD between two reference sub-blocks with a threshold. If it is decided to apply BDOF to a sub-block, for each sample in the sub-block, a sliding 5x5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinements (Vx, Vy) are applied to adjust the bi-predicted sample value for the center sample of the window.
[0208] 2.18. Extended Merge prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: (1) Spatial MVP from spatial neighboring CUs (2) Temporal MVP from collocated CUs (3) History-based MVP from FIFO table (4) Pairwise average MVP (5) Zero MV.
[0209] The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coded in Merge mode, the index of the best Merge candidate is coded using truncated unary binarization (TU). The first bin of the Merge index is coded with a context, and bypass coding is used for the other bins.
[0210] The derivation processes of the various types of Merge candidates are provided in this section. As done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain size of region.
[0211] 2.18.1. Spatial candidate derivation Figure 23The positions of spatial Merge candidates are shown. The derivation of spatial Merge candidates in VVC is the same as in HEVC except that the positions of the first two Merge candidates are swapped. Out of the candidates at the depicted positions, at most four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only if one or more of the positions B0, A0, B1, A1 are not available (e.g., because it belongs to another slice or tile) or is intra coded. After the addition of the candidate at position A1, a redundancy check is performed for the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, improving coding efficiency. To reduce the computational complexity, not all possible pairs of candidates are considered in the mentioned redundancy check. Instead, only pairs of candidates that are linked with an arrow are considered and a candidate is added to the list only if the corresponding candidate for the redundancy check does not have the same motion information. Figure 24
[0212] 2.18.2. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on a collocated CU belonging to a collocated reference picture. The reference picture list to be used for the derivation of the collocated CU is explicitly signaled in the slice header. As depicted by the dashed line in Figure 25 , the scaled motion vector of the temporal Merge candidate is obtained by scaling the motion vector of the collocated CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set equal to 0.
[0213] As depicted in Figure 26 , the position for the temporal candidate is selected between the candidates C0 and C1. If the CU at position C0 is not available, intra coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal Merge candidate.
[0214] 2.18.3. History-based Merge candidate derivation A history-based MVP (HMVP) Merge candidate is added to the Merge list, after the spatial MVP and the TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as MVPs for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (emptied). As long as there is a non-subblock inter coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0215] The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if there is a same HMVP in the table. If found, the same HMVP is removed from the table, and all the HMVP candidates after it are moved forward, and the same HMVP is inserted to the last entry of the table.
[0216] The HMVP candidates can be used in the Merge candidate list construction process. The latest few HMVP candidates in the table are checked in order and inserted into the candidate list, after the TMVP candidates. For spatial or temporal Merge candidates, a redundancy check is applied to the HMVP candidates.
[0217] To reduce the number of redundancy check operations, the following simplifications are introduced: The number of HMPV candidates used for Merge list generation is set to (N<=4)?M:(8-N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table.
[0218] The Merge candidate list construction process from HMVP is terminated once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1.
[0219] 2.18.4. Pairwise average Merge candidate derivation The pair-wise average candidate is generated by averaging the pre-defined pairs of candidates in the existing Merge candidate list, and the pre-defined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If both motion vectors are available in one list, they are averaged even if they point to different reference pictures; if only one motion vector is available, it is used directly; if no motion vector is available, the list is kept invalid.
[0220] When the Merge list is not full after adding the pair-wise average Merge candidate, a zero-MVP is inserted at the end until the maximum number of Merge candidates is reached.
[0221] 2.18.5. Merge estimation region The Merge estimation region (MER) allows the Merge candidate list to be derived independently for CUs in the same Merge estimation region (MER). The candidate blocks that are within the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. In addition, the update process of the history-based motion vector predictor candidate list is only updated when ( xCb + cbWidth ) » Log2ParMrgLevel is greater than xCb » Log2ParMrgLevel and ( yCb + cbHeight ) » Log2ParMrgLevel is greater than ( yCb » Log2ParMrgLevel ), where ( xCb, yCb ) is the top-left luma sample position of the current CU in the picture and ( cbWidth, cbHeight ) is the CU size. The MER size is selected at the encoder side and is signaled in the sequence parameter set as log2_parallel_merge_level_minus2.
[0222] 2.19. New Merge candidate 2.19.1. Non-adjacent Merge candidate derivation In VVC, Figure 27 The five spatial neighboring blocks and one temporal neighbor shown are used to derive the Merge candidate.
[0223] It is proposed to use the same modes as in VVC to derive additional Merge candidates from positions that are not adjacent to the current block. To achieve this, for each search round i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated by the following equation: Offsetx =-i×gridX, Offsety = -i×gridY Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.
[0224] Second, the width and height of the virtual block are calculated using the following formula: newWidth = i×2×gridX+ currWidthnewHeight = i×2×gridY +currHeight.
[0225] Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.
[0226] gridX and gridY are currently set to currWidth and currHeight, respectively.
[0227] Figure 28 This illustrates the relationship between the virtual block and the current block. After the virtual block is generated, block A... i B i C i D i and E i These can be considered VVC spatial neighbor blocks, and their positions are obtained using the same pattern as the pattern in the VVC. Clearly, if the search round i is 0, the virtual block is the current block. In this case, block A... i B i C i D i and E i It is the spatial neighbor block used in VVCMerge mode.
[0228] When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbor blocks are utilized.
[0229] Non-adjacent spatial domain merge candidates are inserted into the merge list after the temporal domain merge candidates in the order B1->A1->C1->D1->E1.
[0230] 2.19.2.STMVP We propose using three spatial merge candidates and one temporal merge candidate to derive the average candidate as the STMVP candidate.
[0231] STMVP is inserted before the top-left spatial domain Merge candidate.
[0232] The STMVP candidate is de-duplicated together with all previous Merge candidates in the Merge list.
[0233] For spatial domain candidates, the first three candidates in the current Merge candidate list are used.
[0234] For temporal candidates, the same position as the VTM / HEVC collocated position is used.
[0235] For spatial domain candidates, the first, second and third candidates in the current Merge candidate list, which are inserted before the STMVP, are denoted as F, S and T.
[0236] The temporal candidate with the same position as the VTM / HEVC collocated position used in TMVP is denoted as Col.
[0237] The motion vector of the STMVP candidate in prediction direction X (denoted as mvLX) is derived as follows: 1) if the reference indices of the four Merge candidates are all valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_S + mvLX_T + mvLX_Col) » 2; 2) if the reference indices of three of the four Merge candidates are valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F x 3 + mvLX_S x 3 + mvLX_Col x 2) » 3 or mvLX = (mvLX_F x 3 + mvLX_T x 3 + mvLX_Col x 2) » 3 or mvLX = (mvLX_S x 3 + mvLX_T x 3 + mvLX_Col x 2) » 3; 3) if the reference indices of two of the four Merge candidates are valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_Col) » 1 or mvLX = (mvLX_S + mvLX_Col) » 1 or mvLX = (mvLX_T + mvLX_Col) » 1.
[0238] NOTE: If temporal candidates are not available, the STMVP mode is turned off.
[0239] 2.19.3. Merge list size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header and the maximum allowed size of the Merge list is 8.
[0240] 2.20. Geometric partition mode (GPM) In VVC, a geometric partition mode is supported for inter prediction. The geometric partition mode is signaled using a CU-level flag as a kind of Merge mode, other Merge modes include regular Merge mode, MMVD mode, CIIP mode and subblock Merge mode. For each possible CU size ( where 8x64 and 64x8 are not included), the geometric partition mode supports 64 partitions in total.
[0241] When this mode is used, the CU is partitioned into two parts by a geometrically positioned straight line Figure 29 . The position of the partition line is mathematically derived from the specific partition’s angle and offset parameters. Each part of the geometric partition in the CU is inter predicted using its own motion; only uni-prediction is allowed for each partition, i.e. each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure the same as regular bi-prediction, i.e. only two motion- compensated predictions are needed for each CU. The uni-prediction motion for each partition is derived using the process described in 2.20.1.
[0242] If the geometric partition mode is used for the current CU, it is further signaled to indicate the geometric partition index of the partition mode (angle and offset) and two Merge indices (one for each partition). The number of maximum GPM candidate size is explicitly signaled in the SPS and specifies the syntax binarization for GPM Merge indices. After predicting each part of the geometric partition, the sample values along the geometric partition edge are adjusted using a hybrid process with adaptive weights as in 2.20.2. This is the prediction signal for the whole CU and the transform and quantization processes will be applied to the whole CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partition mode is stored as described in 2.20.3.
[0243] 2.20.1. Uni-prediction candidate list construction The uni-prediction candidate list is derived directly from the Merge candidate list constructed according to the extended Merge prediction process in 2.18. Let n denote the index of the uni-prediction motion in the geometric uni-prediction candidate list. The LX motion vector of the n-th extended Merge candidate, where X equals the parity of n, is used as the n-th uni-prediction motion vector for the geometric partition mode. These motion vectors are marked with "x" in Figure 30 If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(l-X) motion vector of the same candidate is used as the uni-prediction motion vector for the geometric partition mode.
[0244] 2.20.2. Blending along the geometric partition edge After each part of the geometric partition is predicted using its own motion prediction, blending is applied to the two prediction signals to derive the samples around the geometric partition edge. The blending weight for each position of the CU is derived based on the distance between the single position and the partition edge.
[0245] Position The distance to the partition edge is derived as:
[0246]
[0247] where is the index for the angle and offset of the geometric partition, which depends on the signaled geometric partition index. and The sign of .
[0248] The weight of each part of the geometric partition is derived as follows:
[0249]
[0250] partldx depends on the angle index . One example of the weight is shown in Figure 31 .
[0251] 2.20.3. Motion field storage for geometric partition mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and a combined Mv of Mv1 and Mv2 are stored in the motion field of the geometric partition mode coded CU.
[0252] The motion vector type stored for each individual position in the motion field is determined as:
[0253] where motionldx is equal to which is re-computed from equation (2-18). partldx depends on the angle index .
[0254] If sType is equal to 0 or 1, Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, a combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), Mv1 and Mv2 are simply combined to form a bi-predictive motion vector.
[0255] Otherwise, if Mv1 and Mv2 are from the same list, only the uni-predictive motion Mv2 is stored.
[0256] 2.21. Multi-hypothesis prediction In multi-hypothesis prediction (MHP), up to two additional prediction values are signaled on top of the inter AMVP mode, regular Merge mode, affine Merge and MMVD modes. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0257]
[0258] The weighting factors α are specified according to the following Table 4.
[0259] Table 4 - Weighting factors for MHP
[0260] For the inter AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-predictive mode.
[0261] The additional hypotheses can be either Merge or AMVP mode. In the case of Merge mode, the motion information is indicated by the Merge index and the Merge candidate list is the same as in the geometric partition mode. In the case of AMVP mode, the reference index, MVP index and MVD are signaled.
[0262] 2.22. Non-adjacent spatial candidates Non-adjacent spatial Merge candidates are inserted in the regular Merge candidate list after TMVP. The pattern of spatial Merge candidates is shown in Figure 32 . The distance between the non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block.
[0263] 2.23. Template matching (TM) Template matching (TM) is a decoder-side MV derivation method to refine the motion information of a current CU by finding the closest match between a template (i.e., the top and / or left neighboring blocks of the current CU) in the current picture and a block (i.e., of the same template size) in the reference picture. As shown in Figure 33 , a better MV is searched around the initial motion of the current CU within a search range of [-8, +8] pixels. This template matching proposes two modifications: determining the search step size based on the AMVR mode, and TM can be cascaded with the bilateral matching process in Merge mode.
[0264] In AMVP mode, the MVP candidate is determined based on the template matching error, which is the one that minimizes the difference between the current block template and the reference block template, then TM only performs MV refinement for this specific MVP candidate. TM refines this MVP candidate by using an iterative diamond search, starting from the full-pel MVD precision (or 4-pixel for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined by using a cross search with full-pel MVD precision (or 4-pixel for 4-pixel AMVR mode), then half-pel and quarter-pel are used in turn according to the AMVR mode specified in Table 5. This search process ensures that the MVP candidate still maintains the same MV precision as indicated by the AMVR mode after the TM process.
[0265] Table 5 - Search pattern for AMVR and Merge mode with AMVR
[0266] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 5, TM can be performed up to 1 / 8-pel MVD precision, or skip those precisions beyond half-pel MVD precision, depending on whether the alternative interpolation filter is used according to the merged motion information (i.e. when AMVR is in half-pel mode). In addition, when TM mode is enabled, template matching can operate as an independent process, or as an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enabling condition check.
[0267] 2.24. Overlapped Block Motion Compensation (OBMC) Overlapped Block Motion Compensation (OBMC) has been previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, OBMC is performed for all motion compensation (MC) block boundaries except the right and bottom boundaries of a CU. In addition, it is applied to both luma and chroma components. In JEM, an MC block corresponds to a coding block. When a CU is coded with sub-CU mode (including sub-CU Merge, affine and FRUC modes), each sub-block of the CU is an MC block. To handle CU boundaries in a uniform way, OBMC is performed at sub-block level for all MC block boundaries, with the sub-block size set to be equal to 4x4, as shown in Figure 34 .
[0268] When OBMC is applied to a current sub-block, in addition to the current motion vector, the motion vectors of four neighboring sub-blocks that are connected to the current sub-block (if available and not identical to the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal of the current sub-block.
[0269] The prediction block based on the motion vector of a neighboring sub-block is denoted as P N , where N indicates the index of the neighboring above , below above , left and right sub-block, and the prediction block based on the motion vector of the current sub-block is denoted as P C When P N OBMC is not performed from P N based on the motion information of the neighboring sub-block containing the same motion information as the current sub-block. Otherwise,P N Each sample of P C is added to the same sample in P N Four rows / columns of P C are added to P N and weight factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . The exception is for small MC blocks (i.e., when the height or width of the coded block is equal to 4 or the CU is coded with sub-CU mode), for which only two rows / columns of P N are added to P C . In this case, weight factors {1 / 4, 1 / 8} are used for P N and weight factors {3 / 4, 7 / 8} are used for P C . For P N generated based on the motion vector of the vertically (horizontally) neighboring sub-blocks, P N the samples in the same row (column) of P C are added to
[0270] In JEM, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU for CUs with size smaller than or equal to 256 luma samples. For CUs with size larger than 256 luma samples or not coded with AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is considered during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top and left neighboring blocks is used to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.
[0271] 2.25. Multiple Transform Selection (MTS) for core transforms In addition to DCT-II already adopted in HEVC, a multiple transform selection (MTS) scheme is used for residual coding of both inter-coded blocks and intra-coded blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 6 shows the transform basis functions of the selected DST / DCT.
[0272] Table 6 - Transform basis functions of DCT-II / VIII and DST VII for N-point input
[0273] In order to maintain the orthogonality of the transform matrices, the transform matrices are quantized more precisely than in HEVC. In order to maintain the mid value of the transform coefficients within the 16-bit range, all coefficients must have 10 bits after the horizontal transform and after the vertical transform.
[0274] In order to control the MTS scheme, separate enabling flags are specified at the SPS level for intra and inter, respectively. When MTS is enabled at the SPS, CU level flags are signaled to indicate whether MTS is applied or not. Here, MTS is only applied to luma. MTS signaling is skipped when one of the following conditions is applied.
[0275] - The position of the last significant coefficient of the luma TB is less than 1 (i.e., DC only).
[0276] - The last significant coefficient of the luma TB is located in the MTS zero-out region.
[0277] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are additionally signaled to indicate the transform type for the horizontal direction and the vertical direction, respectively. The transform and signaling mapping table is shown in Table 7. By removing the intra mode and block shape dependency, the transform selection for ISP and implicit MTS is unified. If the current block is in ISP mode, or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are on, only DST7 is used for both the horizontal transform core and the vertical transform core. When it comes to transform matrix precision, the 8-bit primary transform core is used. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform cores (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8) use the 8-bit primary transform core.
[0278] Table 7 – Transform and signaling mapping table
[0279] To reduce the complexity of large size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with size (width or height, or both width and height) equal to 32, high frequency transform coefficients are zeroed. Only the coefficients within the 16x16 low frequency region are kept.
[0280] As in HEVC, the residual of a block can be coded with transform skip mode. To avoid the redundancy of syntax coding, the transform skip flag is not signaled when CU level MTS CU flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. The implicit MTS can still be enabled when MTS is enabled for inter coded blocks.
[0281] 2.26. Sub-block transform (SBT) In VTM, sub-block transform is introduced for inter predicted CUs. In this transform mode, only sub-parts of the residual block are coded for a CU. When a CU with cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the whole residual block or sub-parts of the residual block are coded. In the former case, inter MTS information is further parsed to determine the transform type of the CU. In the latter case, one part of the residual block is coded with the inferred adaptive transform and the other part of the residual block is zeroed.
[0282] When SBT is used for inter coded CUs, SBT type and SBT position information are signaled in the bitstream. As Figure 35 indicated, there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or ¼ of the CU width (or height), resulting in 2:2 partition or 1:3 / 3:1 partition. 2:2 partition is like binary tree (BT) partition, while 1:3 / 3:1 partition is like asymmetric binary tree (ABT) partition. In ABT partition, only the small region contains non-zero residual. If one dimension of the CU is 8 in luma samples, 1:3 / 3:1 partition along that dimension is forbidden. A CU has at most 8 SBT modes.
[0283] Position dependent transform core selection is applied to luma transform blocks in SBT-V and SBT-H (chroma TB always uses DCT-2). The two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal transform and the vertical transform of each SBT position are in Figure 35The SBT is specified jointly with the TU slice, cbf and the horizontal and vertical core transform types of the residual block. For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of a residual TU is larger than 32, both dimensions of the transform are set to DCT-2. Thus, the SBT jointly specifies the TU slice, cbf and the horizontal and vertical core transform types of the residual block.
[0284] The SBT is not applied to CUs coded with Inter-Intra Combined mode.
[0285] 2.27. Template matching based adaptive Merge candidate reordering To improve coding efficiency, after the construction of the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list according to the ascending order of the template matching cost. It is operated in the form of subgroups.
[0286] The template matching cost is measured by the SAD (Sum of Absolute Difference) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes motion information of bi-prediction, the corresponding reference samples are the average of the corresponding reference samples in reference list 0 and the corresponding reference samples in reference list 1, as shown in Figure 36 If the Merge candidate includes sub-CU level motion information, the corresponding reference samples consist of the neighboring samples of the corresponding reference sub-blocks, as shown in Figure 37
[0287] The ordering process is operated in the form of subgroups, as shown in Figure 38 The first three Merge candidates are ordered together. The last three Merge candidates are ordered together.
[0288] The template size (width of the left template or height of the top template) is 1. The subgroup size is 3.
[0289] 2.28. Adaptive Merge candidate list It is assumed that the number of Merge candidates is 8. It takes the first 5 Merge candidates as the first subgroup and the last 3 Merge candidates as the second subgroup (i.e. the last subgroup).
[0290] For the encoder, after the construction of the Merge candidate list, some Merge candidates are adaptively reordered in the ascending order of the Merge candidate cost, as shown in Figure 39
[0291] More specifically, the template matching cost is calculated for the merge candidates in all subgroups except the last one; then, the merge candidates in the subgroups of the merge candidates are reordered except for the last one; finally, the final merge candidate list is obtained.
[0292] For the decoder, after constructing the Merge candidate list, such as Figure 40 As shown, some / no merge candidates are adaptively reordered in ascending order at the merge candidate cost. Figure 40 In this context, the subgroup containing the selected (transmitted via signal) Merge candidate is called the selected subgroup.
[0293] More specifically, if the selected Merge candidate is located in the last subgroup, the Merge candidate list construction process is terminated after deriving the selected Merge candidate, no reordering is performed, and the Merge candidate list remains unchanged; otherwise, the execution process is as follows: After deriving all Merge candidates in the selected subgroup, the Merge candidate list construction process is terminated; the template matching cost for the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained.
[0294] For both the encoder and decoder, the template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template.
[0295] When deriving reference samples for the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.
[0296] The reference samples (RT) for the template used for bidirectional prediction are obtained by using the reference samples of the template in reference list 0 as follows ( ) and reference samples of the template in reference list 1 ( It is derived by weighted averaging.
[0297]
[0298] The weights (8-w) of the reference templates in reference list 0 and the weights (w) of the reference templates in reference list 1 are determined by the BCW indices of the Merge candidates. The BCW indices equal to {0,1,2,3,4} correspond to w equal to {-2,3,4,5,10}, respectively.
[0299] If the Local Illumination Compensation (LIC) flag of the Merge candidate is true, then the reference sample points of the template are derived using the LIC method.
[0300] The template matching cost is computed based on the sum of absolute difference (SAD) of T and RT.
[0301] The template size is 1. This means that the width of the left template and / or the height of the top template is 1.
[0302] If the coding mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered.
[0303] If the coding mode is GPM, the Merge candidates used to derive the uni-prediction candidate list are not reordered.
[0304] 2.29. IBC with extended reference region A design of IBC reference region that does not increase the current storage region required by ECM-3 is proposed and tested for performance.
[0305] Figure 41 The design is shown in the figure. In the figure, the blue square represents the current CTU, and the green square represents the CTU that can be used by IBC reference. Specifically, assuming W represents the maximum horizontal CTU index, and the current CTU index is (m, n), the CTUs with index (0, n)…(m, n) and (m-1, n)…(W, n) define the reference region that can be used by IBC for the coding unit in the current CTU.
[0306] One reason to have such a design is that in the current ECM, the left, top, and top-left CTUs are being used, so they need to be saved. To achieve this, all the CTUs to the right of the top CTU in the top CTU row (for the CTU to be coded in the current CTU row) and all the CTUs to the left of the current CTU in the current CTU row (for the CTU to be coded in the next CTU row) must be kept. This means that such a design does not increase the buffer size required by the current ECM.
[0307] 2.30. IBC with template matching It is proposed to use template matching with IBC for both IBC Merge mode and IBC AMVP mode.
[0308] The IBC-TM Merge list has been modified compared to the list used by the regular IBC Merge mode, the candidate selection employs a de-duplication method, and the motion distance between candidates is the same as for the regular TM Merge mode. The ending zero motion realization (which is meaningless with respect to intra coding) has been replaced by motion vectors to the left (-W, 0), above (0, -H) and top-left (-W, -H) CUs, and then, if necessary, the list is padded with the motion vector to the left without de-duplication.
[0309] In the IBC-TM Merge mode, the selected candidates are refined with a template matching method before the RDO or the decoding process. The IBC-TM Merge mode has been put in competition with the regular IBC Merge mode, and the TM-Merge flag is signaled.
[0310] In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC Merge list. Each of the 3 selected candidates is refined using a template matching method and ordered according to the resulting template matching cost. Then, usually only the first 2 are considered in the motion estimation process.
[0311] Since the IBC motion vector is constrained to be integer and within the reference region as shown in Figure 42 The template matching refinement is quite simple for both IBC-TM Merge and AMVP modes, since the IBC motion vector is constrained to be integer and within the reference region as shown in
[0312] 2.31. Reconstructed Re-ordered IBC (RR-IBC) Screen content coding tools such as Intra Block Copy (IBC) generate a prediction block by directly copying a previously coded reference region in the same picture. Symmetries are often observed in video content, especially in text character regions and computer generated graphics in screen content sequences, as shown in Figure 43 Therefore, a specific screen content coding tool that takes symmetries into account would efficiently compress such video content.
[0313] For screen content video coding, a Reconstructed Re-ordered IBC (RR-IBC) mode is proposed. When applied, the samples in the reconstructed block are flipped according to the flipping type of the current block. At the encoder side, the original block is flipped before the motion search and the residual calculation, while the prediction block is derived without flipping. At the decoder side, the reconstructed block is flipped back to recover the original block.
[0314] For blocks coded with RR-IBC, two flipping methods are supported, horizontal flipping and vertical flipping. First, for blocks coded with IBC AMVP, a syntax flag is signaled to indicate whether the reconstruction is flipped or not, and if it is flipped, another flag is further signaled to specify the flipping type. For IBC Merge, the flipping type is inherited from the neighboring block without syntax signaling. Considering the horizontal or vertical symmetry, the current block and the reference block are usually horizontally or vertically aligned. Therefore, when horizontal flipping is applied, the vertical component of the BV is not signaled and is assumed to be equal to 0. Similarly, when vertical flipping is applied, the horizontal component of the BV is not signaled and is assumed to be equal to 0.
[0315] To better exploit the symmetry property, a flipping-aware BV adjustment method is applied to refine the block vector candidates. For example, as shown in Figure 44A and Figure 44B , the coordinates of the center sample of the neighboring block and the current block are denoted by x nbr , y nbr and x cur , y cur respectively, and the BVs of the neighboring block and the current block are denoted by BV nbr and BV cur respectively. In the case that the neighboring block is coded with horizontal flipping, BV cur the horizontal component of the BV is not directly inherited from the neighboring block, but is calculated by adding a motion shift to the horizontal component of BV nbr , denoted as BV nbr h , i.e., BV cur h = 2( x nbr - x cur ) + BV nbr h . Similarly, in the case that the neighboring block is coded with vertical flipping, BV cur the vertical component of the BV is calculated by adding a motion shift to the vertical component of BV nbr , denoted as BV nbr v , i.e.BV cur v = 2 ( y nbr - y cur ) + BV nbr v .
[0316] 2.32. Intra Template Matching Prediction Intra Template Matching Prediction (Intra-TMP) is a special Intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the most similar template to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode and the same prediction operation is performed at the decoder side.
[0317] The prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in the predefined search area in Figure 45 , which consists of the following parts: R1 : the current CTU R2: the top-left CTU R3: the top CTU R4: the left CTU.
[0318] The SAD is used as the cost function.
[0319] Within each region, the decoder searches for the template with the smallest SAD with respect to the current template and uses its corresponding block as the prediction block.
[0320] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of per-pixel SAD comparisons. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH, where “a” is a constant that controls the gain / complexity trade-off. In practice, “a” is equal to 5.
[0321] The Intra Template Matching tool is enabled for CUs with width and height dimensions smaller than or equal to 64. This maximum CU size for Intra Template Matching is configurable.
[0322] When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level by a dedicated flag.
[0323] 2.33. Intra prediction blending The intra prediction blending method uses multiple prediction values generated from different modes / reference lines.
[0324] In sub-test a, multiple intra prediction values are generated and then blended by weighted averaging. The process of deriving the prediction values to be used in the blending process is described as follows: 1) For the angular intra prediction modes including TIMD and DIMD, the proposed method derives the intra prediction by weighting the intra predictions obtained from multiple reference lines denoted as where is the intra prediction from the default reference line, and is the prediction from the line above the default reference line. The weights are set as and .
[0325] 2) For the TIMD mode with mixing, is used for the 1st st mode ( ), and is used for the 2nd nd mode ( ).
[0326] 3) For the DIMD mode with mixing, the number of prediction values selected for the weighted averaging is increased from 3 to 6.
[0327] In sub-test b, the intra prediction blending is performed on the reference lines instead of the prediction block. Two reference lines (denoted as and ) are used for the intra prediction blending. The corresponding to the intra prediction angles are considered in the blending process. Each value in the blended reference line ( ) is derived from the following equation: .
[0328] The proposed intra prediction blending is applied to the luma block with MRL when the angular intra mode has non-integer slope (requiring reference sample interpolation) and the block size is larger than 16, and is not applied to the ISP coded block. In the method studied in sub-test a, PDPC is applied for the intra prediction modes using the reference line closest to the current block.
[0329] 2.34. Template-based multi-reference line intra prediction (TMRL) The proposed TMRL mode includes the following aspects: a) Extended reference line candidate list and intra prediction mode candidate list.
[0330] The extended reference line candidate list used in this proposal is {1, 3, 5, 7, 12}. The restriction on the top CTU line remains unchanged. The size of the intra prediction mode candidate list is 10. The construction of the intra prediction mode candidate list is similar to MPM. The difference is: The PLANAR mode is excluded from the proposed intra prediction mode candidate list.
[0331] If the DC mode is not already included, it is added after the 5 neighboring PU modes and the DIMD mode.
[0332] Angle modes with an incremental angle from to compared to the existing angle modes in the intra prediction mode candidate list are added.
[0333] b) Construction of the TMRL candidate list For a block, there are 5 x 10 = 50 combinations of the extended reference line and the allowed intra prediction modes. Since the extended reference line starts from reference line 1, the area covered by reference line 0 is used for template matching. On the template area (see Figure 46 ), the SAD cost is calculated between the prediction (generated by the 50 combinations) and the reconstruction. The 20 combinations with the smallest SAD cost are selected in ascending order to form the TMRL candidate list.
[0334] c) Signaling of the TMRL Instead of directly coding the reference line and the intra mode, the index of the TMRL candidate list is coded to indicate which combination of the reference line and the prediction mode is used to code the current block. In the proposed TMRL mode, the selected combination from the combination list is coded with truncated Golomb-Rice coding with a divisor of 4. The binarization process and the codewords are shown in Table 8.
[0335] Table 8 - Binarization process of TMRL index
[0336] d) Encoder side modification Encoder side modification is tested to further improve the coding efficiency. For intra blocks larger than 8x8, if no TMRL mode is selected by the SATD comparison, an additional TMRL RDO is added.
[0337] 2.35. Convolutional Cross-Component Model (CCCM) for intra prediction It is proposed to apply a Convolutional Cross-Component Model (CCCM) to predict chroma samples from reconstructed luma samples in a similar way as the current CCLM mode. As with CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid.
[0338] Furthermore, similar to CCLM, there is an option to use a single model or a multi-model variant of the CCCM. The multi-model variant uses two models, one for deriving samples above the average luma reference value and the other for the remaining samples (in a way following the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.
[0339] 2.35.1. Convolutional filter The proposed convolutional 7-tap filter consists of a 5-tap plus-shaped spatial component, a non-linear term and a bias term. The input of the spatial 5-tap component of the filter consists of the center (C) luma sample co-located with the chroma sample to be predicted and its above / north (N), below / south (S), left / west (W) and right / east (E) neighbors, as shown in Figure 47
[0340] The non-linear term P is expressed as the 2nd power of the center luma sample C and is scaled to the sample value range of the content: P = ( C*C + midVal ) » bitDepth.
[0341] I.e. for 10-bit content, it is computed as: P = ( C*C + 512 ) » 10.
[0342] The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the mid-chroma value (512 for 10-bit content).
[0343] The output of the filter is computed as the convolution between the filter coefficients c i and the input values, and is clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B.
[0344] 2.35.2. Calculation of filter coefficients The filter coefficients c i are computed by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region.Figure 48 The reference region is shown to consist of 6 rows of chroma samples above and to the left of the PU. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension of the region shown in blue is necessary to support the "edge samples" of the plus-shaped spatial filter and is padded when not available.
[0345] MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL-decomposed and the final filter coefficients are computed using back-substitution. The procedure roughly follows the computation of the ALF filter coefficients in the ECM, however, LDL-decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method only uses integer operations.
[0346] 2.35.3. Bitstream signalingThe use of the mode is signaled by a PU-level flag that is CABAC coded. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is only signaled when the intra prediction mode is either LM_CHROMA IDX (to enable single mode CCCM) or MM_LM_CHROMA IDX (to enable multi-model CCCM).
[0347] 2.36. Gradient linear model (GLM) In contrast to CCLM, GLM utilizes luma sample gradients to derive the linear model instead of down-sampled luma values. Specifically, when GLM is applied, the input of the CCLM process (i.e., the down-sampled luma samples ) are replaced by luma sample gradients . The other parts of CCLM (e.g., parameter derivation, prediction sample linear transformation) remain unchanged.
[0348]
[0349] For signaling, when the CCLM mode is enabled to the current CU, two flags are separately signaled for the Cb component and the Cr component to indicate whether GLM is enabled to each component; if GLM is enabled for one component, one syntax element is further signaled to select one of the 4 gradient filters for gradient computation. As Figure 49 shown, 4 gradient filters are enabled for GLM.
[0350] 2.37. Decoder-side intra mode derivation In JEM-2.0, the number of intra-frame modes has been expanded from 35 in HEVC to 67, and these modes are derived at the encoder and explicitly signaled to the decoder. In JEM-2.0, a significant amount of overhead is incurred in intra-frame mode encoding and decoding. For example, in a full intra-frame codec configuration, intra-frame mode signaling overhead can reach 5–10% of the total bit rate. This contribution proposes a decoder-side intra-frame mode derivation method to reduce intra-frame mode encoding and decoding overhead while maintaining prediction accuracy.
[0351] To reduce the overhead of intra-frame mode signaling, this contribution proposes a decoder-side intra-frame mode derivation (DIMD) method. In this method, instead of explicitly transmitting the intra-frame mode via signaling, the information is deduced at both the encoder and decoder from the reconstructed samples from the neighboring blocks of the current block. The intra-frame mode derived via DIMD is used in two ways: 1) For a 2N×2N CU, when the corresponding CU level DIMD flag is turned on, the DIMD mode is used as the intra-frame mode for intra-frame prediction. 2) For N×N CU, the DIMD mode is used to replace a candidate in the existing MPM list to improve the efficiency of intra-mode encoding and decoding.
[0352] 2.37.1. Template-based intra-frame mode derivation Figure 50 This shows the target sample, template sample, and reference sample of the template used in DIMD. For example... Figure 50 As shown, the target represents the current block (block size N) for which the intra-frame prediction mode is to be estimated. The template (made from...) Figure 50 The patterned region indicator (in the template) specifies a set of reconstructed samples used to derive the intra-frame mode. The template size is represented as the number of samples extending above and to the left of the target block within the template, i.e., L. In the current implementation, the template size is 2 (i.e., ) was used for 4×4 and 8×8 blocks, and the template size was 4 (i.e., ) was used for 16x16 and larger blocks. Template reference (by...) Figure 50 The dashed area (indicated by the region) refers to a set of neighboring samples from above and to the left of the template, as defined by JEM-2.0. Unlike template samples, which always come from reconstructed regions, the template's reference samples may not have been reconstructed when encoding / decoding the target block. In this case, JEM-2.0's existing reference sample replacement algorithm is used to replace unavailable reference samples with available ones.
[0353] For each intra prediction mode, DIMD calculates the sum of absolute difference (SAD) between the reconstructed template samples and the predicted samples obtained from the reference samples of the template. The intra prediction mode that produces the smallest SAD is selected as the final intra prediction mode for the target block.
[0354] 2.37.2. DIMD for Intra 2Nx2N CUs For intra 2Nx2N CUs, DIMD is used as an additional intra mode, which is adaptively selected by comparing the DIMD intra mode with the best normal intra mode (i.e., explicitly signaled in the bitstream). For each intra 2Nx2N CU, a flag is signaled to indicate the usage of DIMD. If the flag is 1, the CU is predicted using the intra mode derived by DIMD; otherwise, DIMD is not applied and the CU is predicted using the intra mode explicitly signaled in the bitstream. When DIMD is enabled, the chroma components always reuse the same intra mode as the one derived for the luma component, i.e., the DM mode.
[0355] In addition, for each DIMD-coded CU, the blocks in the CU can adaptively choose to derive their intra modes at the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is signaled to indicate the level at which DIMD is performed. If the flag is 0, it means that DIMD is performed at the PU level and all TUs in the PU use the same derived intra mode for their intra prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is performed at the TU level and each TU in the PU derives its own intra mode.
[0356] Furthermore, when DIMD is enabled, the number of angular directions is increased to 129, and the DC mode and the planar mode remain the same. To accommodate the increased granularity of angular intra modes, the precision of the intra interpolation filter for DIMD-coded CUs is increased from 1 / 32 pixel to 1 / 64 pixel. In addition, to use the derived intra mode of a DIMD-coded CU as an MPM candidate for neighboring intra blocks, the 129 directions of the DIMD-coded CU are converted to “normal” intra modes (i.e., 65 angular intra directions) before being used as MPMs.
[0357] 2.37.3. DIMD for Intra NxN CUs In the proposed method, the intra mode of an intra NxN CU is always signaled. However, to improve the efficiency of the intra mode coding, the intra mode derived from DIMD is used as an MPM candidate for predicting the intra modes of the four PUs in the CU. To not increase the overhead of the MPM index signaling, the DIMD candidate is always placed in the first position in the MPM list and the last existing MPM candidate is removed. In addition, a de-duplication operation is performed so that if the DIMD candidate is redundant, it will not be added to the MPM list.
[0358] 2.37.4. Intra mode search algorithm for DIMD To reduce the encoding / decoding complexity, a direct fast intra mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra mode search. Specifically, an initial candidate list is created by selecting N fixed modes from the allowed intra modes. Then, SAD is calculated for all candidate intra modes and one that minimizes SAD is selected as the starting intra mode. To achieve a good complexity / performance trade-off, the initial candidate list consists of 11 intra modes, including DC, Planar and every 4th mode of the 33 angular intra directions as defined in HEVC, i.e. intra modes 0, 1, 2, 6, 10...30, 34.
[0359] If the starting intra mode is DC or Planar, it is used as the DIMD mode. Otherwise, based on the starting intra mode, a refinement process is then applied where the best intra mode is identified through an iterative search. It works by comparing the SAD values of three intra modes separated by a given search interval at each iteration and maintaining the intra mode that minimizes SAD. The search interval is then reduced by half and the selected intra mode from the last iteration will serve as the center intra mode for the current iteration. For the current DIMD implementation with 129 angular intra directions, at most 4 iterations are used in the refinement process to find the best DIMD intra mode.
[0360] 2.38. Decoder-side intra mode derivation by computing gradients of neighboring samples Three angular modes are selected from a histogram of gradients (HoG) computed from the neighboring pixels of the current block. Once the three modes are selected, their prediction values are computed normally and then their weighted average is used as the final prediction value of the block. To determine the weights, the corresponding amplitudes in the HoG are used for each of the three modes. The DIMD mode is used as an alternative prediction mode and is always checked in the FullRD mode.
[0361] The current version of DIMD has modified some aspects of the signaling, HoG computation and prediction merging. The purpose of this modification is to improve the coding performance as well as to address the complexity issues raised during the last meeting (i.e. throughput for 4x4 blocks). The following sections describe the modifications for each aspect.
[0362] 2.38.1. Signaling Figure 51 The order of parsing flags / indices for integration of DIMD in VTM5 is shown.
[0363] As can be seen, the DIMD flag for a block is first parsed using a single CABAC context which is initialized to the default value 154.
[0364] If flag == 0, the parsing continues normally.
[0365] Otherwise (if flag == 1), only the ISP index is parsed and the following flags / indices are assumed to be zero: BDPCM flag, MIP flag, MRL index. In this case, the entire IPM parsing is also skipped.
[0366] During the parsing phase, when a regular non-DIMD block inquires the IPM of its DIMD neighbor, the mode PLANAR IDX is used as a virtual IPM for the DIMD block.
[0367] 2.38.2. Texture analysis Figure 52 The HoG computation from a template of width 3 pixels is shown. The texture analysis of DIMD includes a histogram of gradients (HoG) computation ( Figure 52 ). The HoG computation is performed by applying horizontal and vertical Sobel filters to the pixels in the 3-wide template around the block. In addition, if the top template pixels fall into a different CTU, they will not be used in the texture analysis.
[0368] Once computed, the IPM corresponding to the two highest histogram bins is selected for the block.
[0369] In the previous version, all pixels in the middle row of the template participated in the HoG computation. However, the current version improves the throughput of this process by applying the Sobel filters more sparsely to the 4x4 block. To this end, only one pixel from the left and one pixel from the top are used. This is shown in Figure 52 .
[0370] In addition to reducing the number of operations used for gradient computation, this property also simplifies the selection of the best 2 modes from the HoG, since the resulting HoG cannot have more than two non-zero amplitudes.
[0371] 2.38.3. Prediction fusion The current method uses a fusion of three prediction values for each block. However, the selection of the prediction mode is different and exploits the proposed combined hypothesis intra prediction method where the planar mode is considered in combination with other modes when computing the intra prediction candidates. In the current version, the two IPM corresponding to the two highest HoG stripes are combined with the planar mode.
[0372] The prediction fusion is applied as a weighted average of the above three prediction values. For this, the weight of the planar is fixed to 21 / 64 (~1 / 3). The remaining weight of 43 / 64 (~2 / 3) is then shared between the two HoG IPM, proportionally to the amplitude of their HoG stripes. Figure 53 This process is visualized. Figure 53 The prediction fusion by weighted average of the two HoG modes and the planar is shown.
[0373] 2.39. Template-based intra mode derivation (TIMD) This contribution proposes a template-based intra mode derivation (TIMD) method using MPMs where the TIMD mode is derived from the MPMs using a neighboring template. The TIMD mode is used as an additional intra prediction method for a CU.
[0374] 2.39.1. TIMD mode derivation For each intra prediction mode in the MPMs, the SATD between the prediction of the template and the reconstructed samples is computed. The intra prediction mode with the smallest SATD is selected as the TIMD mode and is used for the intra prediction of the current CU. The position-dependent intra prediction combination (PDPC) is included in the derivation of the TIMD mode.
[0375] 2.39.2. TIMD signaling A flag is signaled in the sequence parameter set (SPS) to enable / disable the proposed method. When this flag is true, a CU-level flag is signaled to indicate whether the proposed TIMD method is used or not. The TIMD flag is signaled immediately after the MIP flag. If the TIMD flag is equal to true, the remaining syntax elements related to the luma intra prediction mode (including the MRL, the ISP and the normal parsing stage of the luma intra prediction mode) are all skipped.
[0376] 2.39.3. Interaction with new coding tools The DIMD method with prediction fusion using the planar is integrated in EE2. When the EE2 DIMD flag is equal to true, the proposed TIMD flag is not signaled and is set equal to false.
[0377] Similar to PDPC, gradient PDPC is also included in the derivation of TIMD mode.
[0378] When secondary MPM is enabled, both primary and secondary MPMs are used for the derivation of TIMD mode. 6-tap interpolation filter is not used for the derivation of TIMD mode.
[0379] 2.39.4. Modification of MPM list construction in the derivation of TIMD mode During the construction of MPM list, the intra prediction mode of neighboring blocks is derived as planar when it is inter coded. To improve the accuracy of MPM list, when neighboring blocks are inter coded, the propagated intra prediction mode is derived using motion vector and reference picture, and is used in the construction of MPM list. This modification is only applied to the derivation of TIMD mode.
[0380] 2.39.5. TIMD with blending This contribution proposes that, instead of selecting only one mode with the smallest SATD cost, the top two modes with the smallest SATD cost are selected for the intra mode derived using TIMD method, then they are blended with weights, and such weighted intra prediction is used to code the current CU.
[0381] The cost of the two selected modes is compared with a threshold, in the test the cost factor 2 is applied as follows: costMode2 < 2'costMode1.
[0382] If the condition is true, blending is applied, otherwise only mode 1 is used.
[0383] The weights of the modes are calculated from their SATD cost as follows: weight1 = costMode2 / ( costMode1 + costMode2 ) weight2 = 1 - weight1 2.40. Spatial geometry partition mode (SGPM) SGPM is an inter coding tool similar to GPM, where two prediction parts are generated from the intra prediction process. In this mode, a candidate list is constructed, where each entry contains one partitioning and two intra prediction modes, as shown in Figure 54 Figure 54 Spatial GPM candidates are shown. 26 partition modes and 3 intra prediction modes are used to form combinations. The length of the candidate list is set to equal 16. The selected candidate index is signaled.
[0384] Figure 55 GPM templates are shown. The template list is reordered using the SAD between the prediction and reconstruction of the templates for ordering. Figure 55 The template size is fixed to 1.
[0385] For each partition mode, the same Intra-Inter GPM list derivation is used and the IPM list is derived for each part. The IPM list size is set to 3. In the list, the TIMD derived mode is replaced by 2 derived modes with horizontal and vertical directions.
[0386] The SGPM mode is applied with restricted block size: 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32.
[0387] A PPS flag is coded to indicate whether the mixing of two intra predictions is not allowed. Figure 56 GPM partition boundaries are shown. When this PPS flag is set to false, the following adaptive mixing is also used for spatial GPM, where Figure 56 The mixing depth τ shown in is derived as follows:
[0388] Otherwise (the PPS flag is set to true), 1 / 4τ is always used for spatial GPM coded blocks to ensure that no mixing is used when the SGPM block has a fully horizontal or vertical partition angle, and a much narrower mixing width is used when the SGPM block has other partition angles. It is noted that in the current common test conditions (CTC) for screen content video, this flag is set to true.
[0389] 2.41. DIMD Merge When DIMD Merge is used, the DIMD information extracted from neighboring blocks is used to compute the intra prediction of the current block. Specifically, a new gradient merge histogram (MHoG) is computed for the current block based on the HoG of neighboring blocks. Only neighboring blocks coded with DIMD or with DIMD Merge are considered.
[0390] When a single DIMD or DIMD Merge neighboring block is available, its gradient histogram is used to form the MHoG for the current block. If more than one DIMD or DIMD Merge neighboring block is available, the corresponding histograms are combined by amplitude averaging to derive the MHoG.
[0391] Finally, the MHoG is used to calculate the intra prediction mode and weights as in regular DIMD. The direction modes corresponding to the five highest amplitudes in the MHoG and their weights are selected and the corresponding prediction values are blended as in regular DIMD.
[0392] DIMD Merge is only available as an option if the current block has at least one neighbor coded with DIMD or DIMD Merge. Under these conditions, the use of DIMD Merge is then signaled through a CU level flag coded with CABAC. A new CABAC context is included to support the coding of the DIMD Merge flag. DIMD Merge is considered as a sub-mode of DIMD when it comes to signaling. That is, the DIMD Merge flag is only signaled if DIMD flag = 1 and there is a neighboring CU coded with DIMD or DIMD Merge.
[0393] 3. Problem In the current design of video coding standards or video codecs (e.g., VVC, ECM), the weighting parameters for two or more prediction signals of a coding tool (e.g., CIIP, MHP, BCW) are predefined. The content and texture of an image / video can vary a lot in different regions. Due to the predefined weighting parameters, the coding performance can be limited.
[0394] In the current design of DIMD Merge, the DIMD information of neighboring CUs can only be used if the CU is coded with DIMD or DIMD Merge. However, when a CU is not coded with DIMD or DIMD Merge, it can have DIMD information, such as TIMD or SGPM. 4. DETAILED DESCRIPTION The following specific solutions should be considered as examples to explain the general concept. The solutions should not be interpreted in a narrow way. Furthermore, the solutions can be combined in any way.
[0396] Adaptive derivation of weighting parameters 1. Propose to adaptively derive the weighting parameters for a video unit coded with a coding tool that generates a prediction / reconstruction with more than one prediction method or more than one prediction mode of a prediction method.
[0397] a. In one example, a video unit can refer to a color component / sub-picture / slice / tile / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transform unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / intra-subregion of a block / any other region containing more than one sample or pixel.
[0398] b. In one example, a weighting parameter can be derived using a template consisting of neighboring reconstructed samples.
[0399] i. In one example, neighboring reconstructed samples can refer to neighboring samples that are adjacent and / or non-adjacent.
[0400] ii. In one example, neighboring reconstructed samples can refer to left, and / or lower-left, and / or upper-left, and / or above, and / or upper-right neighboring samples.
[0401] iii. In one example, the number of neighboring reconstructed samples and / or the shape / dimension of the template consisting of neighboring reconstructed samples can be predefined or signaled or derived based on coding information.
[0402] 1) In one example, coding information can refer to block dimensions / dimensions.
[0403] iv. In one example, predicted samples of the template can be used to derive a weighting parameter.
[0404] 1) In one example, the same prediction method or prediction mode as the current video unit can be applied to the template to generate predicted samples of the template.
[0405] v. In one example, coordinate information can be used to derive a weighting parameter.
[0406] 1) In one example, coordinate information can refer to the relative position between a sample and a predefined position, such as the top-left position of the template.
[0407] c. In one example, a weighting parameter can be derived using a linear or non-linear equation. Let N, P T i , P T , R T denote the number of prediction methods or prediction modes, the i-th prediction signal of the template generated by the i-th prediction method or prediction mode, the prediction signal of the template, the reconstructed signal of the template, respectively.
[0408] i. In one example, a linear equation can refer to P T = w1* PT 1+ w2* P T 2+ … + w N * P T N + bias.
[0409] 1) In one example, bias equals 0.
[0410] 2) In one example, bias is derived together with w i .
[0411] 3) As a variant, P T = (w1* P T 1+ w2* P T 2+ … + w N * P T N + bias + offset)>>shift, where w i , bias, offset and shift are integers.
[0412] ii. In one example, a method can be applied to derive w T and / or bias by minimizing the difference between P T and R i on a set of training samples, such as minimizing SSD.
[0413] 1) A set of training samples can include all samples of a template.
[0414] 2) A set of training samples can include at least one sample of a template.
[0415] 3) In one example, a Least Mean Square (LMS) method can be used.
[0416] 4) In one example, an LDL method used in CCCM can be used.
[0417] 5) In one example, a Gaussian elimination method can be used.
[0418] 6) In one example, other methods to solve the weighting parameters by minimizing the difference between P T and R T can be used, such as a neural network.
[0419] d. In one example, one or more sets of weighting parameters can be derived.
[0420] i. In one example, multiple sets of weighting parameters can be derived using different templates.
[0421] ii. In one example, the set of weighting parameters used can be predefined or signaled or derived.
[0422] iii. In one example, the set of weighting parameters used can depend on the sample value.
[0423] 1) For example, if the sample value is greater than (or not less than) T, the first set of weighting values can be used.
[0424] 2) For example, if the sample value is less than (or not greater than) T, the second set of weighting values can be usede. In one example, the coding tool can refer to inter prediction.
[0425] i. In one example, the coding tool can refer to CIIP (e.g., CIIP-Planar, CIIP-TIMD, CIIP-TM), BCW (e.g., BCW index derived by TM), MMVD (e.g., MMVD or TM-based MMVD reordering), template matching (TM), IBC (e.g., IBC-TM, IBC with block vector difference, IBC with reconstruction reordering), affine (e.g., affine-MMVD, affine MMVD reordering based on TM), DMVR / multi-pass DMVR, PROF, BDOF or sample-based BDOF, adaptive decoder-side motion vector refinement (ADMVR), OBMC or TM-based OBMC, MHP, GPM (e.g., GPM, GPM-TM, GPM-MMVD, GPM-intra), LIC, bilateral / template matching AMVP-Merge mode or variants thereof, etc.
[0426] f. In one example, the coding tool can refer to IBC or palette or BDPCM or variations thereof.
[0427] i. In one example, the coding tool can refer to IBC-CIIP or IBC-GPM or IBC-LIC.
[0428] g. In one example, the coding tool can refer to intra prediction.
[0429] i. In one example, the coding tool can refer to regular intra prediction method, intra TMP, DIMD, TIMD, ISP, MIP, MRL, PDPC / gradient PDPC, intra prediction fusion, TMRL, cross-component prediction (CCLM), multi-model CCLM, left-side CCLM, above-side CCLM, CCCM, left-side / above-side CCCM, GLM, DIMD chroma, chroma fusion or variants thereof, etc.
[0430] h. In one example, the coding tool can refer to loop filter.
[0431] i. In one example, the coding tool can not be limited to the coding tool in the current ECM.
[0432] 2. In one example, the derived weighting parameters can be used to replace the existing weighting parameters, or used as one or more sets of additional weighting parameters.
[0433] a. In one example, the derived weighting parameters can be used for DIMD, TIMD, intra prediction fusion, TMRL, chroma fusion, CIIP (e.g., CIIP-Planar, CIIP-TIMD, CIIP-TM), BCW (e.g., BCW index derived by TM), bi-prediction, DMVR / multi-pass DMVR, MHP, OBMC or TM-based OBMC, IBC-CIIP.
[0434] 3. In one example, the above derivation method can be used to generate one or more sets of parameters for a coding tool.
[0435] a. In one example, the coding tool can refer to LIC / CCLM / MMLM / CCCM or variants thereof.
[0436] b. In one example, the coding tool can refer to a coding method that uses the current template and / or the reference of the current template to derive a set of parameters that are used for the prediction / reconstruction of the current unit.
[0437] 4. Whether and / or how to apply the derivation of weighting parameters for a video unit can depend on coding information, which can refer to: a. whether a particular coding method is allowed; b. block dimension and / or block size; c. block depth; d. slice / picture type and / or partition tree type (single tree, or dual tree, or local dual tree); e. temporal layer identification; f. block position; g. color format; h. color component i. In one example, the derivation of weighting parameters can be applicable to all color components.
[0438] ii. In one example, when the derivation of weighting parameters is applied to a chroma component, it can be different from the derivation for a luma component.
[0439] iii. In one example, whether and / or how the derivation of the weighting parameter is applied to the first component can depend on whether the derivation of the weighting parameter is applied to the second component.
[0440] 1) In one example, the first component can refer to a chroma component (e.g., Cb and / or Cr), and the second component can refer to a luma component (e.g., Y).
[0441] 2) In one example, the way the derivation of the weighting parameter is applied to the first component can be the same as the second component.
[0442] a) Alternatively, the way the derivation of the weighting parameter is applied to the first component can be different from the second component.
[0443] iv. In one example, the derivation of the weighting parameter can be applied to the luma component but not to the chroma component.
[0444] 1) In one example, the luma component can refer to Y in YCbCr color space or G in RGB color space.
[0445] 2) In one example, the chroma component can refer to Cb and / or Cr in YCbCr color space or R and / or B in RGB color space.
[0446] 5. In one example, the above-mentioned derivation of the weighting parameter can be used to replace the current derivation method of the weighting parameter for a coding tool, or as an additional method.
[0447] a. In one example, the coding tool can refer to LIC, CCLM, MMLM, CCCM, GLM, or variants thereof, etc.
[0448] Fusion of inter template matching 6. It is proposed to fuse more than one inter prediction signal to obtain a final prediction / reconstruction for a video unit, where the more than one inter prediction signal is generated using motion information that is generated / refined / modified using different templates.
[0449] a. In one example, the template can consist of neighboring (adjacent and / or non-adjacent) reconstructed samples of the video unit.
[0450] b. In one example, the template can be adjacent or non-adjacent to the video unit.
[0451] c. In one example, the template can be template-L, template-A, template-LA, template-LB, template-RA, or a combination thereof. Examples are shown in Figures 57A to 57I
[0452] i. Which template is used can depend on the current block's position / size / shape.
[0453] ii. Which template is used can be signaled.
[0454] iii. Which template is used can be implicitly derived.
[0455] d. In one example, the first prediction signal is generated using motion information refined by a first template and the second prediction signal is generated using motion information refined by a second template, wherein the first template is different from the second template.
[0456] i. In one example, the first template is a left template and the second template is an above template.
[0457] e. In one example, more than two inter prediction signals can be fused using different templates.
[0458] f. Whether and how to use the fusion method can be predefined or signaled or derived.
[0459] i. In one example, whether to use the fusion method can be signaled using one or more syntax elements.
[0460] ii. In one example, whether to use the fusion method can be conditionally determined.
[0461] 1) In one example, the condition can depend on coding information, such as: a) template matching cost, b) block size / dimension, c) template shape / size.
[0462] DIMD Merge using information of adjacent video units 7. In addition to neighboring video units coded with DIMD and DIMD Merge, DIMD information of one or more neighboring video units coded with other coding modes can be used for the current video unit.
[0463] a. In one example, the neighboring video units can refer to spatially neighboring video units, adjacent and / or non-adjacent.
[0464] b. In one example, the neighboring video units can refer to temporal video units.
[0465] c. In one example, the neighboring video units can refer to luma and / or chroma video units.
[0466] d. In one example, the DIMD information can refer to HoG and / or inferred intra prediction mode and / or blending weight.
[0467] e. In one example, the coding mode can refer to a specific coding mode, wherein the DIMD information is generated for the coding mode.
[0468] i. In one example, the specific coding mode can refer to TIMD and / or SGPM and / or DIMD chroma and / or GPM-intra and / or IBC-CIIP.
[0469] ii. In one example, when HoG has been computed, the gradient histogram (HoG) of neighboring video units can be used for the current video unit.
[0470] f. In one example, more than one DIMD information candidate can be used for the current video unit.
[0471] i. In one example, a DIMD information list can be constructed.
[0472] 1) In one example, DIMD information from spatial neighboring video units and / or temporal neighboring video units and / or default DIMD information can be used to construct the DIMD information list.
[0473] 2) In one example, one or more DIMD information can be used to infer intra prediction mode and blending weight for the current video unit.
[0474] ii. In one example, which DIMD information is used can be signaled or predefined.
[0475] 8. It is proposed that a history DIMD information table can be used.
[0476] a. In one example, DIMD information of video units coded with DIMD or DIMD Merge or other coding modes that generate DIMD information can be used to update the DIMD information table.
[0477] b. In one example, the maximum number of DIMD information table can be signaled or inferred or predefined.
[0478] c. In one example, the DIMD information table can be updated with a first-in-first-out rule.
[0479] d. In one example, deduplication can be used to update the DIMD information table.
[0480] e. In one example, the DIMD information in the history DIMD information table can be used to construct the list of DIMD information for the current video unit.
[0481] 9. In one example, the DIMD information can be derived from different regions.
[0482] a. In one example, a region can refer to a template including neighboring adjacent and / or non-adjacent reconstructed samples.
[0483] i. For example, left, and / or lower-left, and / or upper-left, and / or above, and / or upper-right templates.
[0484] b. In one example, a region can refer to different lines / rows / columns.
[0485] i. In one example, the DIMD information can be derived from different lines.
[0486] c. In one example, the regions used to derive the DIMD information can be signaled or derived or predefined.
[0487] 10. The DIMD information from neighboring video units can be derived on the fly, instead of using the DIMD information stored by the neighboring video units.
[0488] a. In one example, the DIMD information can be derived from one or neighboring video units.
[0489] i. In one example, the DIMD information can be derived from luma video units and used for chroma video units.
[0490] b. In one example, the DIMD information can refer to gradient histogram and / or derived intra prediction mode and / or blending weight.
[0491] 11. A prediction method with Merge mode is proposed, where the information of the neighboring video units coded with the prediction method can be used for the current video unit.
[0492] a. In one example, the prediction method can refer to an intra prediction method.
[0493] i. In one example, a TIMD with Merge mode is proposed.
[0494] 1) In one example, the intra prediction mode and / or blending weight and / or TM cost of the neighboring video units can be used for the current unit.
[0495] ii. In one example, a SGPM with Merge mode is proposed.
[0496] 1) In one example, the intra prediction mode and / or the partition mode and / or their combination of the neighboring video unit can be used for the current video unit.
[0497] iii. In one example, MIP with Merge mode is proposed.
[0498] iv. In one example, ISP with Merge mode is proposed.
[0499] b. In one example, the prediction method can refer to inter prediction method.
[0500] c. In one example, the prediction method can refer to IBC / Intra TMP.
[0501] d. In one example, the prediction method can refer to palette.
[0502] 12. In one example, the filtering process can be applied to the prediction signal of the video unit coded with IBC and / or Intra TMP.
[0503] a. In one example, the filtering process can refer to PDPC and / or gradient PDPC.
[0504] i. In one example, the intra prediction mode can be derived using DIMD and / or TIMD method for the video unit and be used for the filtering process.
[0505] ii. In one example, the intra prediction mode can be derived by the block vector and be used for the filtering process.
[0506] a) In one example, the PDPC / gradient PDPC can be different from the PDPC / gradient PDPC used for the regular intra prediction mode.
[0507] General aspects 13. In the above examples, the video unit can refer to color component / sub-picture / slice / tile / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transform unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.
[0508] 14. Whether and / or how to apply the above disclosed methods can be signaled at sequence level / group of pictures level / picture level / slice level / tile group level, such as in sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.
[0509] 15. Whether and / or how to apply the above methods can depend on the following information: a. messages signaled in DPS / SPS / VPS / PPS / APS / picture header / slice header / tile group header / largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit b. location of the CU / PU / TU / block / video coding unit c. block dimension of the current block and / or its neighboring blocks d. block shape of the current block and / or its neighboring blocks e. coding mode of the block, e.g. IBC or non-IBC inter mode or non-IBC subblock mode f. indication of color format, such as 4:2:0, 4:4:4 g. coding tree structure h. slice / tile group type and / or picture type i. color component (e.g. can be applied only to chroma component or luma component) j. temporal layer ID k. profile / level / tier of the standard.
[0510] 16. The above disclosed syntax elements can be binarized into a flag, a fixed length coding, an EG(x) coding, a unary coding, a truncated unary coding, a truncated binary coding, etc. It can be signed or unsigned.
[0511] 17. The above disclosed syntax elements can be coded with at least one context model. Or it can be bypass coded.
[0512] 18. The above disclosed syntax elements can be signaled in a conditional way.
[0513] a. SE is signaled only when the corresponding function is applicable.
[0514] b. SE is signaled only when the dimension (width and / or height) of the block meets the condition.
[0515] 19. The syntax elements disclosed above can be signaled at block level / sequence level / grouppicture level / picture level / slice level / tile group level, such as in the coding structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.
[0516] 20. The proposed method(s) can be combined with another coding tool, such as affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.
[0517] 21. The proposed method(s) can be exclusive with another coding tool, such as affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.
[0518] a. In one example, if the proposed method(s) is used, the excluded coding tool is implicitly disabled without signaling.
[0519] b. In one example, if the excluded coding tool is used, the proposed method(s) is implicitly disabled without signaling.
[0520] Figure 58 A flowchart of a method 5800 for video processing according to an embodiment of the disclosure is shown. The method 5800 is implemented during a conversion between a video unit of a video and a bitstream of the video.
[0521] At block 5810, for a conversion between a video unit of a video and a bitstream of the video, decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of the video unit is determined. The one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode. In some embodiments, the video unit comprises at least one of a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a tile, a subpicture, a block, a subregion within a block, or a region containing more than one sample or pixel.
[0522] At block 5820, conversion is performed based on the DIMD information. In some embodiments, the conversion includes encoding the video unit into a bitstream. In some other embodiments, the conversion includes decoding the video unit from a bitstream.
[0523] In some embodiments, the coding mode includes a target coding mode, wherein the DIMD information is generated for the coding mode. For example, the target coding mode includes at least one of a template-based intra mode derivation (TIMD) mode, a spatial geometry partitioning mode (SGPM), a DIMD chroma mode, a geometry partitioning mode (GPM)-intra mode, or an intra block copy (IBC) inter-intra combined prediction (CIIP) (IBC-CIIP) mode. In some other embodiments, if a histogram of gradients (HoG) of one or more neighboring video units is computed, the HoG of the one or more neighboring video units is used for the video unit.
[0524] In some embodiments, multiple candidates of the DIMD information are used for the video unit. In some embodiments, a DIMD information list is constructed. In some embodiments, DIMD information from at least one of the following is used to construct the DIMD information list: one or more spatial neighboring video units or one or more temporal neighboring video units. Alternatively or additionally, a default DIMD information is used to construct the DIMD information list. In some other embodiments, one or more DIMD information is used to derive an intra prediction mode and a blending weight for the video unit. In some embodiments, which DIMD information is used for the video unit is signaled or predefined.
[0525] In some embodiments, the one or more neighboring video units includes at least one of the following: one or more adjacent spatial neighboring video units or non-adjacent spatial neighboring video units. In some embodiments, the one or more neighboring video units includes one or more temporal neighboring video units.
[0526] In some embodiments, the one or more neighboring video units includes at least one of the following: one or more luma video units or one or more chroma video units. In some embodiments, the DIMD information includes at least one of the following: a HoG, a derived intra prediction mode, or a blending weight.
[0527] In some embodiments, one or more neighboring video units are coded with a prediction method with Merge mode and information of the one or more neighboring video units is used for the video unit. In some embodiments, the prediction method comprises an intra prediction method. Alternatively, the prediction method comprises an inter prediction method. In some other embodiments, the prediction method comprises IBC. In some further embodiments, the prediction method comprises Intra Template Matching Prediction (Intra TMP). In another embodiment, the prediction method comprises palette.
[0528] In some embodiments, the prediction method is TIMD with Merge mode. In some embodiments, at least one of the following of the one or more neighboring video units is used for the video unit: an intra prediction mode, a blending weight, or a template matching (TM) cost.
[0529] In some embodiments, the prediction method is SGPM with Merge mode. In some embodiments, at least one of the following of the one or more neighboring video units is used for the video unit: an intra prediction mode or a partition mode.
[0530] In some embodiments, DIMD information of one or more neighboring video units is stored. Alternatively, DIMD information of one or more neighboring video units is dynamically derived.
[0531] In some embodiments, DIMD information is derived from one or neighboring video units. In some other embodiments, DIMD information is derived from luma video units and used for chroma video units. In some embodiments, the DIMD information comprises at least one of the following: HoG, a derived intra prediction mode, or a blending weight.
[0532] In some embodiments, a filtering process is applied to a prediction signal of a video unit coded with at least one of IBC or Intra TMP. For example, the filtering process comprises at least one of the following: position dependent intra prediction combination (PDPC) or gradient PDPC.
[0533] In some embodiments, an intra prediction mode is derived using at least one of DIMD or TIMD methods for a video unit and the intra prediction mode is used for a filtering process. In some other embodiments, an intra prediction mode is derived by a block vector and used for a filtering process. In some embodiments, the PDPC or gradient PDPC is different from the PDPC or gradient PDPC used for a regular intra prediction mode.
[0534] In some embodiments, a history DIMD information table is used. In some embodiments, DIMD information of a target video unit is used to update the DIMD information table, the target video unit being coded with at least one of DIMD, DIMD Merge, or another coding mode that generates DIMD information. In some embodiments, a maximum number of DIMD information tables is signaled or derived or predefined.
[0535] In some embodiments, the DIMD information table is updated with a first-in-first-out rule. In some embodiments, de-duplication is used to update the DIMD information table. In some embodiments, DIMD information in a history DIMD information table is used to build a list of DIMD information for a video unit.
[0536] In some embodiments, DIMD information is derived from different regions. In some embodiments, a region is a template that includes at least one of neighboring adjacent reconstructed samples or neighboring non-adjacent reconstructed samples. In some other embodiments, a region includes at least one of a left template, a lower-left template, an upper-left template, an above template, or an upper-right template. In some further embodiments, a region includes a plurality of lines or rows or columns. In some embodiments, DIMD information is derived from different lines. In some embodiments, which region is used to derive DIMD information is signaled or derived or predefined.
[0537] In some embodiments, a video unit includes at least one of a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a tile, a subpicture, a block, a sub-region within a block, or a region containing more than one sample or pixel.
[0538] In some embodiments, an indication of whether and / or how DIMD information of one or more neighboring video units is determined is indicated at one of a sequence level, a picture group level, a picture level, a slice level, or a tile group level. In some other embodiments, an indication of whether and / or how DIMD information of one or more neighboring video units is determined is indicated in one of a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.
[0539] In some embodiments, the method 5800 further includes determining whether and / or how to determine DIMD information for one or more neighboring video units based on at least one of: a message indicating in one of: a DPS, a SPS, a VPS, a PPS, an APS, a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a LCU group, a TU, a PU block, a video coding unit, a location of one of: a CU, a PU, a TU, a block, a video coding unit, a block dimension of the current block and / or its neighboring blocks, a block shape of the current block and / or its neighboring blocks, a coding mode of the video unit, an indication of a color format, a coding tree structure, a slice type, a tile group type, a picture type, a color component, a temporal layer identification, a profile or level or layer of a standard.
[0540] In some embodiments, the SE is binarized into one of a flag, a fixed length coding, an EG(x) coding, a unary coding, a truncated unary coding, or a truncated binary coding. In some embodiments, the SE is signed or unsigned.
[0541] In some embodiments, the SE is coded with at least one context model. Alternatively, the SE is bypass coded.
[0542] In some embodiments, the SE is signaled in a conditional manner. In some embodiments, the SE is signaled only when the corresponding function is applicable. Alternatively, the SE is signaled only when a dimension of the video unit satisfies a condition.
[0543] In some embodiments, the SE is indicated at one of: a sequence level, a picture group level, a picture level, a slice level, or a tile group level. In some embodiments, the SE is indicated at one of: a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a coding tree block (CTB), or a coding tree unit (CTU).
[0544] In some embodiments, the derivation of the set of weighting parameters is combined with another coding tool. Alternatively, the derivation of the set of weighting parameters excludes another coding tool. In some embodiments, the other coding tool includes at least one of: affine, multiple transform selection (MTS), low frequency non-separable transform (LFNST), MMVD, MIP, ISP, CCLM, CCCM, symmetric motion vector difference (SMVD), BDOF, DMVR, history-based motion vector prediction (HMVP), template matching, IBC, or palette.
[0545] In some embodiments, if derivation of a set of weighting parameters is applied, the excluded coding tool is implicitly disabled without signaling. In some other embodiments, if the excluded coding tool is used, derivation of a set of weighting parameters is implicitly disabled without signaling.
[0546] According to yet some embodiments of the disclosure, a method for storing a bitstream of a video is provided. The method includes determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of a video unit of the video, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode; generating the bitstream based on the DIMD information; and storing the bitstream in a non-transitory computer-readable medium.
[0547] According to yet some embodiments of the disclosure, a method for storing a bitstream of a video is provided. The method includes determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of a video unit of the video, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode; generating the bitstream based on the DIMD information; and storing the bitstream in a non-transitory computer-readable medium.
[0548] Implementations of the present disclosure can be described according to the following clauses, which can be combined in any reasonable manner.
[0549] Clause 1. A method of video processing, comprising: for a conversion between a video unit of a video and a bitstream of the video, determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of the video unit, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode; and performing the conversion based on the DIMD information.
[0550] Clause 2. The method of clause 1, wherein the coding mode comprises a target coding mode, wherein the DIMD information is generated for the coding mode.
[0551] Clause 3. The method of clause 2, wherein the target coding mode comprises at least one of a template-based intra mode derivation (TIMD) mode, a spatial geometry partitioning mode (SGPM), a DIMD chroma mode, a geometry partitioning mode (GPM) - intra mode, or an intra block copy (IBC) cross intra inter prediction (CIIP) (IBC-CIIP) mode.
[0552] Item 4. The method according to item 2, wherein if a histogram of gradients (HoG) of the one or more neighboring video units is computed, the HoG of the one or more neighboring video units is used for the video unit.
[0553] Item 5. The method according to item 1, wherein multiple candidates of DIMD information are used for the video unit.
[0554] Item 6. The method according to item 5, wherein a list of DIMD information is constructed.
[0555] Item 7. The method according to item 5, wherein DIMD information from at least one of the following: one or more spatial neighboring video units or one or more temporal neighboring video units, and / or wherein default DIMD information is used to construct the list of DIMD information, is used to construct the list of DIMD information.
[0556] Item 8. The method according to item 5, wherein one or more DIMD information is used to derive intra prediction mode and blending weight for the video unit.
[0557] Item 9. The method according to item 5, wherein which DIMD information is used for the video unit is signaled or predefined.
[0558] Item 10. The method according to item 1, wherein the one or more neighboring video units comprise at least one of the following: one or more adjacent spatial neighboring video units or non-adjacent spatial neighboring video units.
[0559] Item 11. The method according to item 1, wherein the one or more neighboring video units comprise one or more temporal neighboring video units.
[0560] Item 12. The method according to item 1, wherein the one or more neighboring video units comprise at least one of the following: one or more luma video units or one or more chroma video units.
[0561] Item 13. The method according to item 1, wherein the DIMD information comprises at least one of the following: HoG, derived intra prediction mode or blending weight.
[0562] Item 14. The method according to item 1, wherein the one or more neighboring video units are coded with a prediction method having Merge mode, and information of the one or more neighboring video units is used for the video unit.
[0563] Item 15. The method of item 14, wherein the prediction method comprises an intra prediction method, or wherein the prediction method comprises an inter prediction method, or wherein the prediction method comprises IBC, or wherein the prediction method comprises Intra Template Matching Prediction (Intra TMP), or wherein the prediction method comprises a palette.
[0564] Item 16. The method of item 15, wherein the prediction method is TIMD with Merge mode.
[0565] Item 17. The method of item 16, wherein at least one of the following of the one or more neighboring video units is used for the video unit: an intra prediction mode, a blending weight, or a template matching (TM) cost.
[0566] Item 18. The method of item 15, wherein the prediction method is SGPM with Merge mode.
[0567] Item 19. The method of item 18, wherein at least one of the following of the one or more neighboring video units is used for the video unit: an intra prediction mode or a partition mode.
[0568] Item 20. The method of item 1, wherein the DIMD information of the one or more neighboring video units is stored, or wherein the DIMD information of the one or more neighboring video units is dynamically derived.
[0569] Item 21. The method of item 20, wherein the DIMD information is derived from one or neighboring video units.
[0570] Item 22. The method of item 21, wherein the DIMD information is derived from a luma video unit and used for a chroma video unit.
[0571] Item 23. The method of item 20, wherein the DIMD information comprises at least one of the following: HoG, a derived intra prediction mode, or a blending weight.
[0572] Item 24. The method of item 1, wherein a filtering process is applied to a prediction signal of the video unit coded with at least one of IBC or Intra TMP.
[0573] Item 25. The method of item 24, wherein the filtering process comprises at least one of the following: position dependent intra prediction combination (PDPC) or gradient PDPC.
[0574] Item 26. The method of item 25, wherein the intra prediction mode is derived using at least one of a DIMD or a TIMD method for the video unit, and the intra prediction mode is used for the filtering process.
[0575] Item 27. The method of item 25, wherein the intra prediction mode is derived by a block vector and used for the filtering process.
[0576] Item 28. The method of item 24, wherein the PDPC or gradient PDPC is different from a PDPC or gradient PDPC used for a regular intra prediction mode.
[0577] Item 29. The method of item 1, wherein a history DIMD information table is used.
[0578] Item 30. The method of item 29, wherein DIMD information of a target video unit coded with at least one of DIMD, DIMD Merge, or another coding mode that generates the DIMD information is used to update the DIMD information table.
[0579] Item 31. The method of item 29, wherein a maximum number of DIMD information tables is signaled or derived or predefined.
[0580] Item 32. The method of item 29, wherein the DIMD information table is updated with a first-in-first-out rule.
[0581] Item 33. The method of item 29, wherein deduplication is used to update the DIMD information table.
[0582] Item 34. The method of item 29, wherein DIMD information in the history DIMD information table is used to build a DIMD information list for the video unit.
[0583] Item 35. The method of item 1, wherein DIMD information is derived from different regions.
[0584] Item 36. The method of item 35, wherein a region is a template including at least one of a neighboring adjacent reconstructed sample or a neighboring non-adjacent reconstructed sample.
[0585] Item 37. The method of item 36, wherein the region includes at least one of a left template, a lower-left template, an upper-left template, an above template, or an upper-right template.
[0586] Item 38. The method of item 35, wherein the region comprises a plurality of rows or rows or columns.
[0587] Item 39. The method of item 38, wherein the DIMD information is derived from different rows.
[0588] Item 40. The method of item 35, wherein which region is used to derive the DIMD information is signaled or derived or predefined.
[0589] Item 41. The method of any of items 1 to 40, wherein the video unit comprises at least one of a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a tile, a subpicture, a block, a subregion within a block, or a region containing more than one sample or pixel.
[0590] Item 42. The method of any of items 1 to 41, wherein an indication of whether and / or how the DIMD information of the one or more neighboring video units is determined is indicated at one of a sequence level, a picture group level, a picture level, a slice level, or a tile group level.
[0591] Item 43. The method of any of items 1 to 41, wherein an indication of whether and / or how the DIMD information of the one or more neighboring video units is determined is indicated in one of a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.
[0592] Item 44. The method of any of items 1 to 41, further comprising determining whether and / or how the DIMD information of the one or more neighboring video units is determined based on at least one of a message indicated in one of a DPS, a SPS, a VPS, a PPS, an APS, a picture header, a slice header, a tile group header, a largest coding unit (LCU), a coding unit (CU), a LCU row, a LCU group, a TU, a PU block, a video coding unit, a location of one of a CU, a PU, a TU, a block, a video coding unit, a block dimension of a current block and / or its neighboring blocks, a block shape of a current block and / or its neighboring blocks, a coding mode of the video unit, an indication of a color format, a coding tree structure, a slice type, a tile group type, a picture type, a color component, a temporal layer identification, a profile or level or layer of a standard.
[0593] Item 45. The method of any of items 1 to 44, wherein the SE is binarized to one of a flag, a fixed length coding, an EG(x) coding, a unary coding, a truncated unary coding, or a truncated binary coding.
[0594] Item 46. The method of item 45, wherein the SE is signed or unsigned.
[0595] Item 47. The method of any of items 1 to 46, wherein the SE is coded with at least one context model, or wherein the SE is bypass coded.
[0596] Item 48. The method of any of items 1 to 46, wherein the SE is signaled in a conditional manner.
[0597] Item 49. The method of item 48, wherein the SE is signaled only when a corresponding function is applicable, or wherein the SE is signaled only when a dimension of the video unit satisfies a condition.
[0598] Item 50. The method of any of items 1 to 49, wherein the SE is indicated at one of a sequence level, a group of pictures level, a picture level, a slice level, or a tile group level.
[0599] Item 51. The method of any of items 1 to 49, wherein the SE is indicated at one of a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a coding tree block (CTB), or a coding tree unit (CTU).
[0600] Item 52. The method of any of items 1 to 51, wherein the derivation of the set of weighting parameters is combined with another coding tool, or wherein the derivation of the set of weighting parameters excludes another coding tool.
[0601] Item 53. The method of item 52, wherein the other coding tool comprises at least one of: affine, multiple transform selection (MTS), low frequency non-separable transform (LFNST), MMVD, MIP, ISP, CCLM, CCCM, symmetric motion vector difference (SMVD), BDOF, DMVR, history-based motion vector prediction (HMVP), template matching, IBC, or palette.
[0602] Item 54. The method of item 52, wherein the excluded coding tool is implicitly disabled without signaling if derivation of the set of weighting parameters is applied.
[0603] Item 55. The method of item 52, wherein the derivation of the set of weighting parameters is implicitly disabled without signaling if the excluded coding tool is used.
[0604] Item 56. The method of any of items 1 to 55, wherein the converting comprises encoding the video unit into the bitstream.
[0605] Item 57. The method of any of items 1 to 55, wherein the converting comprises decoding the video unit from the bitstream.
[0606] Item 58. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of items 1 to 57.
[0607] Item 59. A non-transitory computer-readable storage medium storing instructions to cause a processor to perform the method of any of items 1 to 57.
[0608] Item 60. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a video processing apparatus, wherein the method comprises determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of a video unit of the video, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode, and generating the bitstream based on the DIMD information.
[0609] Item 61. A method for storing a bitstream of a video, comprising determining decoder-side intra mode derivation (DIMD) information of one or more neighboring video units of a video unit of the video, wherein the one or more neighboring video units are coded in a coding mode different from a DIMD mode or a DIMD Merge mode, generating the bitstream based on the DIMD information, and storing the bitstream in a non-transitory computer-readable medium.
[0610] Example Device Figure 59A block diagram of a computing device 5900 in which various embodiments of the present disclosure can be implemented is shown. The computing device 5900 can be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or can be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).
[0611] It is to be understood that Figure 59 The computing device 5900 shown in FIG. 13 is for purposes of illustration and explanation only and is not intended as any limitation on the functionality and scope of the embodiments of the present disclosure.
[0612] As Figure 59 shown, the computing device 5900 includes a general-purpose computing device 5900. The computing device 5900 can include at least one or more processors or processing units 5910, a memory 5920, a storage unit 5930, one or more communication units 5940, one or more input devices 5950, and one or more output devices 5960.
[0613] In some embodiments, the computing device 5900 can be implemented as any user terminal or server terminal having computing capability. The server terminal can be a server provided by a service provider, a mainframe computing device, or the like. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal including a mobile telephone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, and includes accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 5900 can support any type of interface to the user (such as "wearable" circuitry, etc.).
[0614] The processing unit 5910 can be a physical processor or a virtual processor and can implement various processing based on programs stored in the memory 5920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 5900. The processing unit 5910 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0615] Computing device 5900 typically includes various computer storage media. Such media can be any media accessible by computing device 5900, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 5920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 5930 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 5900.
[0616] The computing device 5900 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 59 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0617] Communication unit 5940 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 5900 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 5900 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0618] Input device 5950 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 5960 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 5940, computing device 5900 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 5900 can also communicate with one or more devices that enable a user to interact with computing device 5900, or any device that enables computing device 5900 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).
[0619] In some embodiments, some or all of the components of computing device 5900 can not be integrated in a single device, but can also be arranged in a cloud computing architecture. In a cloud computing architecture, components can be provided remotely and work together to achieve the functionality described in this disclosure. In some embodiments, cloud computing provides computation, software, data access, and storage services that do not require end-user knowledge of the physical location or configuration of the system that delivers the services. In various embodiments, cloud computing delivers services via the internet using appropriate protocols. For example, a cloud computing provider provides applications through the internet that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, and corresponding data, can be stored on servers at remote locations. Computing resources in a cloud computing environment can be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure can provide services through shared data centers, although they appear as a single point of access for users. Thus, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, they can be provided from a conventional server or installed directly or otherwise on a client device.
[0620] In embodiments of the disclosure, computing device 5900 can be used to implement video encoding / decoding. Memory 5920 can include one or more video coding modules 5925 having one or more program instructions. These modules can be accessed and executed by processing unit 5910 to perform the functions of the various embodiments described herein.
[0621] In example embodiments that perform video encoding, input device 5950 can receive video data as input 5970 to be encoded. The video data can be processed by, for example, video coding module 5925 to generate an encoded bitstream. The encoded bitstream can be provided as output 5980 via output device 5960.
[0622] In example embodiments that perform video decoding, input device 5950 can receive an encoded bitstream as input 5970. The encoded bitstream can be processed by, for example, video coding module 5925 to generate decoded video data. The decoded video data can be provided as output 5980 via output device 5960.
[0623] While the present disclosure has been particularly shown and described with reference to the preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application as defined by the appended claims. Such changes are intended to fall within the scope of the application. Accordingly, the foregoing description of embodiments of the application are not intended to be limiting.
Claims
1. A video processing method, comprising: For the conversion between video units and the bitstream of the video, determine decoder-side intra-frame mode derivation (DIMD) information of one or more neighboring video units of the video unit, wherein the one or more neighboring video units are encoded and decoded in a codec mode different from the DIMD mode or the DIMD Merge mode; and The conversion is performed based on the DIMD information.
2. The method according to claim 1, wherein the encoding / decoding mode includes a target encoding / decoding mode, wherein DIMD information is generated for the encoding / decoding mode.
3. The method of claim 2, wherein the target encoding / decoding mode includes at least one of the following: Template-based intra-frame mode derivation (TIMD) mode, Spatial Geometric Partitioning Model (SGPM) DIMD chroma mode Geometric Segmentation Mode (GPM) - Intra-frame mode, or Intra-Block Copy (IBC) Inter-Intra-Frame Joint Prediction (CIIP) (IBC-CIIP) mode.
4. The method of claim 2, wherein if the gradient histogram (HoG) of the one or more neighboring video units is calculated, the HoG of the one or more neighboring video units is used for the video unit.
5. The method of claim 1, wherein a plurality of candidates of DIMD information are used for the video unit.
6. The method of claim 5, wherein a DIMD information list is constructed.
7. The method of claim 5, wherein DIMD information from at least one of the following is used to construct the DIMD information list: one or more spatially adjacent video units or one or more temporally adjacent video units, and / or The default DIMD information is used to construct the DIMD information list.
8. The method of claim 5, wherein one or more DIMD information are used to derive intra-predictive codes and mixing weights for the video unit.
9. The method of claim 5, wherein which DIMD information is used for the video unit is either transmitted via signal or predefined.
10. The method of claim 1, wherein the one or more neighboring video units comprise at least one of the following: one or more adjacent spatial neighboring video units or non-adjacent spatial neighboring video units.
11. The method of claim 1, wherein the one or more neighboring video units comprise one or more temporal neighboring video units.
12. The method of claim 1, wherein the one or more adjacent video units comprise at least one of: one or more luminance video units or one or more chroma video units.
13. The method of claim 1, wherein the DIMD information includes at least one of the following: HoG, derived intra-prediction mode, or hybrid weights.
14. The method of claim 1, wherein the one or more neighboring video units are encoded and decoded using a prediction method with a Merge mode, and information from the one or more neighboring video units is used for the video unit.
15. The method of claim 14, wherein the prediction method comprises an intra-frame prediction method, or The prediction method mentioned above includes inter-frame prediction methods, or The prediction method mentioned above includes IBC, or The prediction method mentioned above includes Intra-Template Matching Prediction (IntraTMP), or The prediction method mentioned above includes a color palette.
16. The method of claim 15, wherein the prediction method is TIMD with Merge mode.
17. The method of claim 16, wherein at least one of the following is used for the video unit in the one or more adjacent video units: Intra-frame prediction mode, Mixed weights, or Template matching (TM) cost.
18. The method of claim 15, wherein the prediction method is SGPM with Merge mode.
19. The method of claim 18, wherein at least one of the following is used in the one or more adjacent video units: Intra-prediction mode or Segmentation mode.
20. The method of claim 1, wherein the DIMD information of the one or more adjacent video units is stored, or The DIMD information of one or more neighboring video units is dynamically derived.
21. The method of claim 20, wherein the DIMD information is derived from one or a neighboring video unit.
22. The method of claim 21, wherein the DIMD information is derived from the luminance video unit and used in the chroma video unit.
23. The method of claim 20, wherein the DIMD information comprises at least one of the following: HoG, derived intra-prediction mode, or hybrid weights.
24. The method of claim 1, wherein the filtering process is applied to the prediction signal of the video unit encoded or decoded using at least one of IBC or IntraTMP.
25. The method of claim 24, wherein the filtering process comprises at least one of the following: position-dependent intra-prediction combination (PDPC) or gradient PDPC.
26. The method of claim 25, wherein the intra-frame prediction mode is derived using at least one of the DIMD or TIMD methods for the video unit, and the intra-frame prediction mode is used in the filtering process.
27. The method of claim 25, wherein the intra-frame prediction mode is derived by block vectors and used in the filtering process.
28. The method of claim 24, wherein the PDPC or gradient PDPC is different from the PDPC or gradient PDPC used for a conventional intra-frame prediction mode.
29. The method of claim 1, wherein a historical DIMD information table is used.
30. The method of claim 29, wherein the DIMD information of the target video unit is used to update the DIMD information table, and the target video unit is encoded or decoded using at least one of DIMD, DIMD Merge, or another encoding / decoding mode for generating the DIMD information.
31. The method of claim 29, wherein the maximum number of DIMD information tables is determined by signal transmission, derivation, or predefinition.
32. The method of claim 29, wherein the DIMD information table is updated using a first-in, first-out (FIFO) rule.
33. The method of claim 29, wherein deduplication is used to update the DIMD information table.
34. The method of claim 29, wherein the DIMD information in the historical DIMD information table is used to construct a DIMD information list for the video unit.
35. The method of claim 1, wherein the DIMD information is derived from different regions.
36. The method of claim 35, wherein the region is a template, the template comprising at least one of adjacent reconstructed samples or adjacent non-adjacent reconstructed samples.
37. The method of claim 36, wherein the region comprises at least one of the following: a left template, a lower left template, a upper left template, an upper top template, or a upper right template.
38. The method of claim 35, wherein the region comprises a plurality of lines, rows, or columns.
39. The method of claim 38, wherein the DIMD information is derived from different lines.
40. The method of claim 35, wherein the region used is used to deduce whether the DIMD information is transmitted via signal, deduced, or predefined.
41. The method according to any one of claims 1 to 40, wherein the video unit comprises at least one of the following: Color components, Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Code-decode tree block (CTB). Codec Unit (CU) Code-decode tree unit (CTU) CTU line, CTU group, strip, piece, Sub-images, piece, Sub-regions within a block, or A region containing more than one sample point or pixel.
42. The method according to any one of claims 1 to 41, wherein an indication of whether and / or how to determine the DIMD information of the one or more adjacent video units is indicated in one of the following places: sequence level, Image group level, Image quality, strip level, or Film series level.
43. The method according to any one of claims 1 to 41, wherein an indication of whether and / or how to determine the DIMD information of the one or more neighboring video units is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.
44. The method according to any one of claims 1 to 41, further comprising: The determination of whether and / or how to determine the DIMD information of the one or more neighboring video units is based on at least one of the following: The message is indicated by one of the following: DPS, SPS, VPS, PPS, APS, image header, strip header, slice header, maximum codec unit (LCU), codec unit (CU), LCU line, LCU group, TU, PU block, video codec unit. One of the following locations: CU, PU, TU, block, video codec unit. The block dimensions of the current block and / or its neighboring blocks. The block shape of the current block and / or its neighboring blocks. The encoding and decoding mode of the video unit, Indicators of color format, Encoder tree structure, Strip type, Film series type, Image type, Color components, Temporal layer identifier, Standard grade, level, or tier.
45. The method according to any one of claims 1 to 44, wherein the SE is binarized into one of a flag, a fixed-length codec, an EG(x) codec, a unary codec, a rounding unary codec, or a rounding binary codec.
46. The method of claim 45, wherein the SE is signed or unsigned.
47. The method according to any one of claims 1 to 46, wherein the SE is encoded and decoded using at least one context model, or The SE mentioned therein is bypassed encoding and decoding.
48. The method according to any one of claims 1 to 46, wherein the SE is conditionally transmitted via signal.
49. The method of claim 48, wherein the SE is transmitted via signal only when the corresponding function is applicable, or The SE is transmitted via signal only if the dimension of the video unit meets the condition.
50. The method according to any one of claims 1 to 49, wherein the SE is indicated in one of the following locations: sequence level, Image group level, Image quality, strip level, or Film series level.
51. The method according to any one of claims 1 to 49, wherein the SE is indicated in one of the following locations: Predicted blocks (PB). Transform block (TB) Code block (CB) Prediction Unit (PU) Transformer Unit (TU) Codec Unit (CU) Code-decode tree block (CTB), or Code-decode tree unit (CTU).
52. The method according to any one of claims 1 to 51, wherein the derivation of the set of weighting parameters is combined with another encoding / decoding tool, or The derivation of the set of weighted parameters excludes another encoding / decoding tool.
53. The method of claim 52, wherein the other encoding / decoding tool comprises at least one of the following: affine, multiple transform selection (MTS), low-frequency non-separable transform (LFNST), MMVD, MIP, ISP, CCLM, CCCM, symmetric motion vector difference (SMVD), BDOF, DMVR, history-based motion vector prediction (HMVP), template matching, IBC, or palette.
54. The method of claim 52, wherein if the derivation of the set of weighted parameters is applied, the excluded codec tools are implicitly disabled without signaling.
55. The method of claim 52, wherein if the excluded codec tool is used, the derivation of the set of weighted parameters is implicitly disabled without signaling.
56. The method according to any one of claims 1 to 55, wherein the conversion comprises encoding the video unit into the bitstream.
57. The method according to any one of claims 1 to 55, wherein the conversion comprises decoding the video unit from the bitstream.
58. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 57.
59. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 57.
60. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, the method comprising: Determine decoder-side intra-frame mode derivation (DIMD) information for one or more neighboring video units of the video unit, wherein the one or more neighboring video units are encoded and decoded in a codec mode different from the DIMD mode or the DIMD Merge mode; and The bit stream is generated based on the DIMD information.
61. A method for storing a bitstream of video, comprising: Determine decoder-side intra-frame mode derivation (DIMD) information of one or more neighboring video units of the video unit, wherein the one or more neighboring video units are encoded or decoded in a codec mode different from the DIMD mode or the DIMD Merge mode. The bit stream is generated based on the DIMD information; as well as The bitstream is stored in a non-transitory computer-readable medium.