Method and device for video processing and medium
By fusing multiple video block predictions based on different prediction schemes in video processing, generating target predictions, and converting video blocks based on target predictions, the problem of low fusion efficiency of multiple prediction schemes in the prior art is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380072685.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-13
- Filing Date
- 2023-10-12
- Publication Date
- 2025-05-23
Smart Images

Figure CN120035987A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to predictive fusion. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Versatile Video Codec (VVC) standard. However, it is generally expected to further improve the encoding and decoding quality and encoding and decoding efficiency of video encoding and decoding technologies. Summary of the invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: for conversion between a current video block of a video and a bitstream of the video, obtaining multiple predictions for the current video block, the multiple predictions being determined based on multiple different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; and performing conversion based on the target prediction.
[0005] According to the method of the first aspect of the present disclosure, multiple predictions for the current video block determined based on multiple different prediction schemes are fused to generate a fused prediction. Compared with conventional solutions that only use predictions determined based on specific prediction schemes, the proposed method can advantageously improve coding quality and coding efficiency.
[0006] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, and the instructions enable a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method performed by an apparatus for video processing. The method includes: obtaining multiple predictions for a current video block of the video, the multiple predictions being determined based on multiple different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; and generating a bitstream based on the target prediction.
[0009] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining multiple predictions for a current video block of the video, the multiple predictions being determined based on multiple different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; generating a bitstream based on the target prediction; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0013] Figure 2 A block diagram illustrating a first example video encoder is shown according to some embodiments of the present disclosure;
[0014] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0015] Figure 4 The nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in the picture are shown;
[0016] Figure 5 An example of an encoder block diagram is shown;
[0017] Figure 6 67 intra prediction modes are shown;
[0018] Figure 7 Reference samples for wide-angle intra prediction are shown;
[0019] Figure 8 The problem of discontinuity is shown in the case of orientations exceeding 45°;
[0020] Fig. 9 The locations of the sample points used to derive α and β are shown;
[0021] Fig. 10A Four Sobel-based gradient patterns for the gradient linear model (GLM) are shown;
[0022] Fig. 10B An example of classifying neighboring sample points into two groups is shown;
[0023] Fig.11A is a schematic diagram showing the definition of sample points used by PDPC applied to a diagonal upper right mode;
[0024] Fig. 11B is a schematic diagram showing the definition of sample points used by PDPC applied to a diagonal lower left mode;
[0025] Fig. 11C is a schematic diagram showing the definition of sample points used by PDPC applied to an adjacent diagonal upper right pattern;
[0026] Fig.11D is a schematic diagram showing the definition of sample points used by PDPC applied to an adjacent diagonal lower left pattern;
[0027] Fig.12 A gradient method for non-vertical / non-horizontal modes is shown;
[0028] Fig.13 The value of nScale is shown in relation to nTbH and the number of modes; for all cases where nScale<0, the gradient method is used;
[0029] Fig.14 Flowcharts are shown: current PDPC (left) and proposed PDPC (right);
[0030] Fig.15 Neighboring blocks (L, A, BL, AR, AL) used in the derivation of the generic MPM list are shown;
[0031] Fig.16 An example of the proposed intra reference mapping is shown;
[0032] Fig.17 An example of four reference rows adjacent to a prediction block is shown;
[0033] Fig.18A is a schematic diagram showing examples of sub-partitioning for 4×8 and 8×4 CUs;
[0034] Fig.18B is a diagram showing an example of sub-partitioning for a CU other than 4×8, 8×4, and 4×4;
[0035] Fig.19 The matrix-weighted intra prediction process is shown;
[0036] Fig. 20 Target sample points, template sample points, and reference sample points of the template used in DIMD are shown;
[0037] Fig.21 The proposed intra-block decoding process is shown;
[0038] Fig. 22 HoG computation from a template with a width of 3 pixels is shown;
[0039] Fig.23 shows the prediction fusion by weighted averaging of two HoG modes and planes;
[0040] Fig.24 shows neighboring reconstructed samples for DIMD chroma mode;
[0041] Fig.25 The MMVD search point is shown;
[0042] Fig.26 A diagram for a symmetric MVD mode is shown;
[0043] Fig. 27 shows the extended CU area used in BDOF;
[0044] Fig.28 An affine motion model based on control points is shown;
[0045] Fig.29 The affine MVF of each sub-block is shown;
[0046] Fig.30 The positions of the inherited affine motion prediction values are shown;
[0047] Fig.31 Control point motion vector inheritance is shown;
[0048] Fig.32 The positions of candidate positions for constructing the affine merge pattern are shown;
[0049] Fig.33 A diagram showing the use of motion vectors for the proposed combined approach;
[0050] Fig.34 Sub-block MV VSB and pixel Δv(i, j) are shown;
[0051] Fig.35A and Fig.35B The SbTMVP process in VVC is shown;
[0052] Fig.36 Local illumination compensation is shown;
[0053] Fig.37 It is shown that subsampling for the short edges is not performed;
[0054] Fig.38 Decoding side motion vector refinement is shown;
[0055] Fig.39 A diamond-shaped area in the search area is shown;
[0056] Fig.40 The location of the spatial Merge candidate is shown;
[0057] Fig.41 shows candidate pairs considered for redundancy check of spatial merge candidates;
[0058] Fig.42 A diagram showing motion vector scaling for temporal Merge candidates is shown;
[0059] Fig.43 The time domain Merge candidate C is shown 0 and C 1 Candidate location of
[0060] Fig.44 The VVC spatial domain neighboring blocks of the current block are shown;
[0061] Fig.45 shows a diagram of a virtual block in the i-th search round;
[0062] Fig.46 An example of GPM partitions grouped at the same angle is shown;
[0063] Fig.47 Unidirectional prediction MV selection for geometric partitioning mode is shown;
[0064] Fig.48 shows the blending weights w using the geometric partitioning mode 0 An exemplary generation of
[0065] Fig.49 The spatial neighboring blocks used to derive spatial Merge candidates are shown;
[0066] Fig.50 It shows that template matching is performed on the search area around the initial MV;
[0067] Fig.51 A diagram showing sub-blocks of an OBMC application is shown;
[0068] Fig.52 The SBT position, type and transformation type are shown;
[0069] Fig.53 The neighboring sample points used to calculate the SAD are shown;
[0070] Fig.54 shows neighboring samples used to calculate SAD for sub-CU level motion information;
[0071] Fig.55 The sorting process is shown;
[0072] Fig.56 shows the reordering process in the encoder;
[0073] Fig.57 shows the reordering process in the decoder;
[0074] Fig.58 The spatial domain portion of the convolution filter is shown;
[0075] Fig.59 The reference area (with its filling) used for deriving the filter coefficients is shown;
[0076] Fig.60 A flowchart showing a method for video processing according to an embodiment of the present disclosure; and
[0077] Fig.61 A block diagram of a computing device is shown in which various embodiments of the present disclosure may be implemented.
[0078] Same or similar reference numbers generally refer to same or similar elements throughout the drawings. DETAILED DESCRIPTION
[0079] The principle of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, without implying any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0080] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0081] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include the particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art to affect correlation with other embodiments.
[0082] It should be understood that, although the terms "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another element. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0083] The terms used herein are only used for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprises", "has", "has", "includes" and / or "comprising" are used herein to indicate the presence of the features, elements and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components and / or combinations thereof. Example Environment
[0084] Figure 1 1 is a block diagram illustrating an example video codec system 100 that may utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0085] The video source 112 may include a source such as a video acquisition device. Examples of a video acquisition device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0086] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a bit sequence that forms a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of pictures. The associated data may include sequence parameter sets, picture parameter sets, and other grammatical structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0087] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120, or may be outside the destination device 120, which is configured to be connected to an external display device interface.
[0088] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0089] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of a video encoder 114 in the system 100 is shown.
[0090] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared between the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0091] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0092] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is a picture in which the current video block is located.
[0093] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail below. Figure 2 are shown separately in the example.
[0094] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0095] The mode selection unit 203 may select one of a plurality of encoding modes (intra-frame encoding or inter-frame encoding), for example, based on the error result, and provide the generated intra-frame encoded block or inter-frame encoded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, the mode selection unit 203 may select a combined intra-frame and inter-frame prediction (CIIP) mode in which the prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 may also select a resolution for the motion vector (e.g., sub-pixel precision or integer pixel precision) for the block.
[0096] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0097] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, a "P slice" and a "B slice" may refer to a portion of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0098] In some examples, the motion estimation unit 204 may perform unidirectional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures of list 0 or list 1 to find the reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 containing the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0099] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indexes and a plurality of motion vectors, the plurality of reference indexes indicating a plurality of reference pictures in list 0 and list 1 containing a plurality of reference video blocks, and the plurality of motion vectors indicating a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0100] In some examples, motion estimation unit 204 may output a complete set of motion information for use in a decoding process by a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal motion information of a current video block with reference to motion information of another video block. For example, motion estimation unit 204 may determine that motion information of a current video block is sufficiently similar to motion information of a neighboring video block.
[0101] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0102] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0103] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0104] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a prediction video block and various syntax elements.
[0105] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of samples in the current video block.
[0106] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0107] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0108] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0109] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transformed coefficient video block respectively to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0110] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the block effect artifacts in the video block.
[0111] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate the entropy encoded data and output a bitstream including the entropy encoded data.
[0112] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0113] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in the present disclosure.
[0114] In Figure 3 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.
[0115] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy - encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy - encoded video data, and the motion compensation unit 302 may determine motion information from the entropy - decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction area in a B - slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.
[0116] The motion compensation unit 302 may generate a motion - compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub - pixel precision may be included in the syntax element.
[0117] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values for sub - integer pixels of the reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information, and the motion compensation unit 302 may use the interpolation filter to generate a prediction block.
[0118] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter - frame - encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.
[0119] The intra - prediction unit 303 may use, for example, the intra - prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse - quantizes (i.e., de - quantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0120] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter may also be applied to filter the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on a display device.
[0121] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the section titles used in this document are for ease of understanding, and the embodiments disclosed in the section are not limited to that section. In addition, although some embodiments are described with reference to multifunctional video codecs or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video coding technology. Specifically, the present disclosure relates to a coding tool that fuses multiple prediction values to generate a new hybrid prediction value in the prediction to obtain better coding efficiency. The present disclosure can be applied to existing video coding standards such as HEVC or Versatile Video Codec (VVC). The present disclosure can also be applied to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.264 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.264 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec structure that utilizes temporal prediction plus transform codec. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and put them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard with the goal of a 50% bitrate reduction compared to HEVC. 2.1 Color Space and Chroma Subsampling A color space (also called a color model (or color system)) is an abstract mathematical model that simply describes the range of colors as a tuple of numbers, usually 3 or 4 values or color components (such as RGB). Basically, a color space is a refinement of a coordinate system and subspace. For video compression, the most frequently used color spaces are YCbCr and RGB. YCbCr, Y'CbCr or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a family of color spaces used as part of the color image pipeline in video and digital photography systems. Y' is the luma component, CB and CR are the blue difference and red difference chroma components. Y' (with superscript notation) is different from Y, which is brightness, meaning that light intensity is encoded non-linearly based on the gamma-corrected RGB primaries. Chroma subsampling is the practice of encoding an image with less resolution for chroma information than for luma information, taking advantage of the human visual system's lower sharpness for color differences than for luma. 2.1.1. 4:4:4 Each of the three Y'CbCr components has the same sample rate, so there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2 The two chroma components are sampled at half the sampling rate of luma: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one third with almost no visual difference. Figure 4 Examples of nominal vertical and horizontal positions for a 4:2:2 color format are depicted in FIG. Figure 4 The nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in the picture are shown. 2.1.3. 4:2:0 In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on every alternate line in this scheme. Therefore, the data rate is the same. Cb and Cr are each subsampled horizontally and vertically at a factor of 2. There are three variants of the 4:2:0 scheme with different horizontal and vertical positioning. In MPEG-2, Cb and Cr are co-located horizontally. Cb and Cr are arranged between pixels in the vertical direction (arranged with gaps). In JPEG / JFIF, H261, and MPEG-1, Cb and Cr are set intermittently, halfway between alternate luma samples. In 4:2:0 DV, Cb and Cr are co-located horizontally and vertically on alternate lines. Table 1 SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag 2.2 Encoding and decoding process of typical video codecs Figure 5 An example of an encoder block diagram is shown. Figure 5 An example of an encoder block diagram for VVC is shown, which contains three loop filter blocks: deblocking filter (DF), sample adaptive offset (SAO) and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, where the coding side information transmits the offset and filter coefficients through the signal. ALF is located at the last processing stage of each picture and can be seen as a tool that tries to capture and repair artifacts produced by previous stages. 2.3 Intra-mode codec with 67 intra-prediction modes Figure 667 intra prediction modes are shown. To capture the arbitrary edge directions present in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65, as shown in Figure 6 As shown, planar and DC modes remain unchanged. These dense directional intra prediction modes are applicable to all block sizes and to both luma and chroma intra prediction. In HEVC, each intra-coded block has a square shape, and the length of each of its sides is a power of 2. Therefore, no division operation is required to generate intra prediction values using DC mode. In VVC, blocks can have a rectangular shape, which in general requires the use of a division operation for each block. To avoid division operations for DC prediction, only the longer sides are used to calculate the average value for non-square blocks. 2.3.1 Wide-angle intra prediction Although 67 modes are defined in VVC, the precise prediction direction for a given intra prediction mode index further depends on the block shape. The traditional angular intra prediction direction is defined as from 45 degrees to -135 degrees in a clockwise direction. In VVC, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced mode is signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e. 67, and the intra mode encoding and decoding method remains unchanged. Figure 7 The reference samples for wide-angle intra prediction are shown in Figure 1. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 7 shown. The number of replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2. Table 2 Intra-frame prediction modes replaced by wide-angle mode Figure 8 The problem of discontinuity is shown in the case of an orientation exceeding 45°. Figure 8As shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap Δpα. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, which are [-14, -12, -10, -6, 72, 76, 78, 80]. When predicting blocks through these modes, the samples in the reference cache are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it will align the design of non-fractional modes in traditional prediction mode and wide-angle mode. In VVC, 4:2:2 and 4:4:4 chroma formats and 4:2:0 chroma formats are supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, expanding the number of entries from 35 to 67 to align with the expansion of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2: chroma format is updated by replacing some values of the entries of the mapping table to more accurately convert the prediction angles of the chroma blocks. 2.4. Intra-frame prediction mode encoding and decoding for chroma components For the chroma component of the intra PU, the encoder selects the best chroma prediction mode from the five modes including planar, DC, horizontal, vertical and direct copy of the intra prediction mode for the luma component. The mapping between the intra prediction direction for chroma and the intra prediction mode number is shown in Table 3. When the intra prediction mode number for the chroma component is 4, the intra prediction direction for the luma component is used for intra prediction sample generation for the chroma component. When the intra prediction mode number for the chroma component is not 4 and is the same as the intra prediction mode number for the luma component, the intra prediction direction 66 is used for intra prediction sample generation for the chroma component. 2.5 Inter-frame prediction For each inter-frame prediction CU, the motion parameters include motion vectors, reference picture indices, and reference picture list usage indices, as well as additional information required for the new coding features of VVC that will be used for inter-frame prediction sample generation. Motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. A Merge mode is specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as additional scheduling introduced in VVC. Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to Merge mode is explicit transmission of motion parameters, where motion vectors, corresponding reference picture indices for each reference picture list, reference picture list usage flags, and other required information are explicitly transmitted through signals per CU. 2.6 Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level codec mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed within the current picture. The luminance block vector of the CU encoded and decoded by IBC has integer precision. The chrominance block vector is also rounded to integer precision. When used in conjunction with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The CU encoded and decoded by IBC is regarded as a third prediction mode in addition to the intra or inter prediction mode. The IBC mode is applicable to CUs with a width and height that are less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height no greater than 16 luma samples. For non-Merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In hash-based search, the hash key matching (32-bit CRC) between the current block and the reference blocks is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For current blocks of larger size, the hash key is determined to match the hash key of the reference block when all hash keys in all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matching reference is calculated and the one with the smallest cost is selected. In the block matching search, the search range is set to cover the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, which can be signaled as IBC AMVP mode or IBC Skip / Merge mode as shown below: -IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and pair candidates. - IBC AMVP mode: Block vector differences are encoded in the same way as motion vector differences. The block vector prediction method uses two candidates as predictors, one from the left neighbor and one from the top neighbor (if IBC coded). When either neighbor is not available, the default block vector is used as predictor. A flag is signaled to indicate the block vector predictor index. 2.7 Cross-component linear model prediction In order to reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in VVC. For this CCLM prediction mode, the chrominance samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model: pred C (i,j)=α·rec L ′(i,j)+β (2-1) Among them, pred C (i, j) represents the predicted chroma sample in the CU, and rec L (i, j) represents the downsampled reconstructed luma sample of the same CU. The CCLM parameters (α and β) are derived using up to four neighboring chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block size is W×H, W' and H' are set to: – When LM mode is applied, W'=W, H'=H; – When LM_T mode is applied, W'=W+H; – When LM_L mode is applied, H'=H+W. The upper neighboring position is denoted as S[0, -1] ... S[W'-1, -1], and the left neighboring position is denoted as S[-1, 0] ... S[-1, H'-1]. Then, four sample points are selected as: – When LM mode is applied and both the upper neighboring sample points and the left neighboring sample points are available, S[W' / 4,-1], S[3*W' / 4,-1], S[-1,H' / 4], S[-1,3*H' / 4]; – When the LMT mode is applied or only the upper neighboring samples are available, S[W' / 8,-1],S[3*W' / 8, -1],S[5*W' / 8,-1],S[7*W' / 8,-1]; – When LM-L mode is applied or only left neighbor samples are available, S[-1,H' / 8], S[-1,3*H' / 8],S[-1,5*H' / 8],S[-1,7*H' / 8]. The four adjacent brightness samples at the selected position are downsampled and compared four times to find the two larger values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B Their corresponding chrominance sample values are represented by y 0 A.y 1 A.y 0 B and y 1 B. Then x A 、x B ,y A and B is derived as: X 0 =(x 0 A+x 1 A +1)>>1;X b =(x 0 B +x 1 B +1)>>1;Y a =(y 0 A +y 1 A +1)>>1;Y b =(y 0 B +y 1 B +1)>>1(2-2) Finally, the parameters α and β of the linear model are obtained according to the following formula. β=Y b -α·X b (2-4) Fig. 9 An example of the positioning of the left sample points and the upper sample points involved in the CCLM mode and the samples of the current block is shown. Fig. 9The positioning of the sample points used to derive α and β is shown. The division operation is implemented using a lookup table to calculate the parameter. In order to reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are expressed as exponents. For example, diff is approximated by a 4-bit significant part and an exponent. Therefore, the table for 1 / diff is reduced to 16 elements for 16 values of the significant bit, as shown below: DivTable[ ]={0,7,6,5,5,4,4,3,3,2,2,1,1,0} (2-5) This will help reduce the complexity of the calculations and the memory size required to store the required tables. In addition to the upper and left templates being used together to calculate linear model coefficients, they can also be used alternately in two other LM modes, called LM_T and LM_L modes. In LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples. In LM mode, the left template and the upper template are used to calculate the linear model coefficients. To match the chroma sample positions for 4:2:0 video sequences, two types of downsampling filters are applied to the luma samples to achieve a 2 to 1 downsampling ratio in both the horizontal and vertical directions. The choice of downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "type-0" and "type-2" content, respectively. Note that when the upper reference line is at a CTU boundary, only one luma line (normal line buffer in intra prediction) is used for downsampling luma samples. This parameter calculation is performed as part of the decoding process, not just as an encoder search operation. Therefore, no syntax is used to communicate the α and β values to the decoder. For chroma intra mode coding and decoding, a total of 8 intra modes are allowed for chroma intra mode coding and decoding. These modes include five traditional intra modes and three cross-component linear model modes (LM, LM_T and LM_L). The chroma mode signaling and derivation process are shown in Table 3. Chroma mode coding and decoding depends directly on the intra prediction mode of the corresponding luminance block. Since separate block partitioning structures for luminance and chrominance components are enabled in the I stripe, one chroma block can correspond to multiple luminance blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luminance block covering the center position of the current chroma block is directly inherited. Table 3 Derivation of chrominance prediction mode from luma mode when CCLM is enabled As shown in Table 4, a single binarization table is used regardless of the value of sps_cclm_enabled_flag. Table 4 Unified binarization table for chroma prediction mode Value of intra_chroma_pred_mode Binary String 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111 In Table 4, the first binary bit indicates whether it is normal mode (0) or LM mode (1). If it is LM mode, the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 binary bit indicates whether it is LM_L (0) or LM_T (1). For this case, when sps_cclm_enabled_flag is 0, the first binary bit of the corresponding intra_chroma_pred_mode binarization table can be discarded before entropy coding and decoding. Or, in other words, the first binary bit is inferred to be 0 and therefore not coded and decoded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binary bits in Table 4 are context coded and decoded with their own context model, and the remaining binary bits are bypass coded and decoded. In addition, to reduce luma-chroma latency in dual trees, when the 64×64 luma codec tree node is split with Not Split (and ISP is not used for 64×64 CU) or QT, the chroma CU in the 32×32 / 32×16 chroma codec tree node is allowed to use CCLM in the following way: If a 32x32 chroma node is not partitioned or split QT partition, all chroma CUs in the 32x32 node can use CCLM. - If a 32x32 chroma node is split using horizontal BT and the 32x16 child node is not split or uses vertical BT split, all chroma CUs in the 32x16 chroma node can use CCLM. Under all other luma and chroma codec tree partitioning conditions, CCLM is not allowed for chroma CUs. 2.8 Gradient Linear Model (GLM) Compared with CCLM, GLM uses the gradient of luma samples to derive the linear model instead of the downsampled luma values. In other words, instead of low-pass filtered luma samples, the gradient G is replaced in the CCLM process. The other designs of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged. C=α·G+β Fig. 10A It is shown that it can be calculated by one of four Sobel-based gradient patterns. For signaling, when CCLM mode is enabled for the current CU, two flags are transmitted separately for the Cb / Cr components to indicate whether GLM is enabled for the component; if GLM is enabled for a component, a syntax element is further transmitted by signal to select Fig. 10A One of the four gradient patterns in is used for gradient calculation. 2.9 Multi-model Linear Model (MMLM) For MMLM, there can be more than one linear model between the luma samples and chroma samples in the CU. In this method, the neighboring luma samples and the neighboring chroma samples of the current block are classified into several groups, each of which is used as a training set to derive a linear model (i.e., specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified based on the same rules for the classification of neighboring luma samples. Neighboring samples can be classified into M groups, where M is 2 or 3. In the case of M=2 and M=3, the MMLM method is designed as two additional chroma prediction modes in addition to the original LM mode, called MMLM2 and MMLM3. The encoder selects the best mode in the RDO process and transmits it through the signal. When M is equal to 2, Fig. 10B An example of classifying neighboring samples into two groups is shown. The threshold is calculated as the average of neighboring reconstructed luminance samples. Neighboring samples with Rec'L[x, y] <= threshold are classified into group 1; and neighboring samples with Rec'L[x, y]> threshold are classified into group 2. Similar to CCLM, there are 3 modes in MMLM, namely MMLM, MMLM_T and MMLM_L. The two models are derived as follows. Fig. 10BAn example of classifying neighboring samples into two groups is shown. The threshold is the average of the neighboring samples of the luminance reconstruction. A linear model for each class is derived using the Least Mean Square (LMS) method (if enabled) or the Min / Max method of VVC. 2.10. Position-dependent intra prediction combination In VVC, the intra prediction results for DC, planar, and several angular modes are further modified by the Position Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that calls for a combination of boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. PDPC is not applied if the current block is in BDPCM mode or the MRL index is greater than 0. The prediction sample pred(x',y') is predicted using a linear combination of the reference samples and the intra prediction mode (DC, planar, angular) according to the following equation 2-8: pred(x',y')=Clip(0,(1< <BitDepth)-1,(wL×R -1,y’ +wT×R x-1 +(64-wL-wT)×pred(x',y')+32)>>6) (2-9) Where R x,-1 ,R -1,y Respectively represent the reference sample points located at the top and left boundaries of the current sample point (x, y). If PDPC is applied to DC, planar, horizontal and vertical intra modes, no additional boundary filtering is required, as is required in the case of HEVC DC mode boundary filtering or horizontal / vertical mode edge filtering. The PDPC process for DC mode and planar mode is the same. For angular modes, if the current angular mode is HOR_IDX or VER_IDX, the left or top reference samples are not used, respectively. The PDPC weights and scaling factors depend on the prediction mode and block size. PDPC is applied to blocks with width and height both greater than or equal to 4. Fig.11A and Figure 8 B shows the reference samples (R x,-1 and R -1,y ). The prediction sample pred(x',y') is located at (x',y') in the prediction block. For example, for the diagonal mode, the reference sample R x,-1 The coordinate x of is given by: x = x' + y' + 1, and the reference point R -1,yThe coordinate y of is similarly given by: y = x' + y' + 1. For other angular modes, the reference sample point R x,-1 and R -1,y Can be at a fractional sample position. In this case, the sample value at the nearest integer sample position will be used. FIG. 11A to FIG. 11D The definition of samples used by PDPC applied to diagonal and adjacent angle intra modes is shown. Fig.11A The diagonal upper right mode is shown. Fig. 11B The diagonal lower left corner mode is shown. Fig. 11C The adjacent diagonal upper right mode is shown. Fig.11D The adjacent diagonal lower left mode is shown. Gradient PDPC like Fig.12 As shown in Figure 2, the gradient-based approach is extended to non-vertical / non-horizontal modes. Here, the gradient is calculated as r(-1,y)–r(-1+d,-1), where d is the horizontal displacement depending on the angular direction. A few points need to be noted here. The gradient term r(-1,y)–r(-1+d,-1) needs to be calculated once per row because it does not depend on the x position. The calculation of d is already part of the original intra prediction process which can be reused, so there is no need to calculate d separately. Therefore, the accuracy of d is 1 / 32 pixel. When d is in fractional position, we already use two-tap (linear) filtering, i.e., if dPos is the displacement with 1 / 32 pixel precision, dInt is the (rounded down) integer part (dPos>>5), and dFract is the fractional part with 1 / 32 pixel precision (dPos&31), then r(-1+d) is calculated as: r(-1+d)=(32–dFrac)*r(-1+dInt)+dFrac*r(-1+dInt+1). As described in a, this 2-tap filtering is performed once per row (if necessary). Finally, the prediction signal is calculated: p(x,y)=Clip(((64–wL(x))*p(x,y)+wL(x)*(r(-1,y)-r(-1+d,-1))+32)>>6) Where wL(x)=32>>((x<<1)>>nScale2), and nScale2=(log2(nTbH)+log2(nTbW)–2)>>2, which are the same as the vertical / horizontal mode. In short, the same process is applied compared to the vertical / horizontal mode (in fact, d=0 indicates the vertical / horizontal mode). Fig. 9 The gradient method for non-vertical / non-horizontal patterns is shown. Secondly, for non-vertical / non-horizontal modes, when (nScale < 0) or when PDPC cannot be applied due to the unavailability of auxiliary reference samples, the gradient-based method is activated. Fig.13 The value of nScale related to TB size and angular mode is shown in Figure 2 to better visualize the use of the gradient method. In addition, the flowcharts for the existing PDPC and the proposed PDPC are shown. Fig.13 The value of nScale is shown relative to nTbH and the number of modes; for all cases where nScale<0, the gradient method is used. Fig.14 Flowcharts are shown: current PDPC (left) and proposed PDPC (right). 2.12 Auxiliary MPM A secondary MPM list is introduced. The existing primary MPM (PMPM) list includes 6 entries, while the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries in the general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the plane mode. Fig.15 As shown, the remaining entries consist of the intra modes of the left (L), above (A), lower left (BL), upper right (AR), and upper left (AL) neighboring blocks, the directional mode with the offset added from the first two available directional modes of the neighboring blocks, and the default mode. If the CU block is vertically oriented, the order of neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. Fig.15 Neighboring blocks (L, A, BL, AR, AL) used to derive the common MPM list are shown. First parse the PMPM flag, if it is equal to 1, parse the PMPM index to determine which entry of the PMPM list is selected, otherwise parse the SPMPM flag to determine whether to parse the SMPM index or the remaining modes. 2.13 6-tap intraframe interpolation filter In order to improve the prediction accuracy, a 6-tap interpolation filter is proposed to replace the 4-tap cubic interpolation filter. The filter coefficients are derived based on the same polynomial regression model, but the polynomial order is 6. The filter coefficients are as follows: {0,0,256,0,0,0}, / / 0 / 32 position {0,-4,253,9,-2,0}, / / 1 / 32 position {1,-7,249,17,-4,0}, / / 2 / 32 position {1,-10,245,25,-6,1}, / / 3 / 32 position {1,-13,241,34,-8,1}, / / 4 / 32 position {2,-16,235,44,-10,1}, / / 5 / 32 position {2,-18,229,53,-12,2}, / / 6 / 32 position {2,-20,223,63,-14,2}, / / 7 / 32 position {2,-22,217,72,-15,2}, / / 8 / 32 position {3,-23,209,82,-17,2}, / / 9 / 32 position {3,-24,202,92,-19,2}, / / 10 / 32 position {3,-25,194,101,-20,3}, / / 11 / 32 position {3,-25,185,111,-21,3}, / / 12 / 32 position {3,-26,178,121,-23,3}, / / 13 / 32 position {3,-25,168,131,-24,3}, / / 14 / 32 position {3,-25,159,141,-25,3}, / / 15 / 32 position {3,-25,150,150,-25,3}, / / half pixel position The reference samples used for interpolation come from the reconstructed samples or the padding samples in HEVC, so there is no need to perform conditional checks on the availability of reference samples. It is recommended to use a 4-tap cubic interpolation filter instead of using the nearest rounding operation to derive the extended intra-frame reference samples. Fig.16 As shown in the example in , a four-tap interpolation filter is used to derive the value of the reference sample point P, while in JEM-3.0 or HM, P is directly set to X1. Fig.16 An example of the proposed intra reference mapping is shown. 2.14. Multiple Reference Line (MRL) Intra Prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. Fig.17In Figure 1, an example of 4 reference lines is depicted, where the samples of segments A and F are not taken from the reconstructed neighboring samples, but are filled with the closest samples from segments B and E, respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines are used (reference line 1 and reference line 2). The index of the selected reference row (mrl_idx) is signaled and used to generate intra prediction values. For reference row indices greater than 0, only additional reference row modes are included in the MPM list, and only the MPM index is signaled without the remaining modes. The reference row index is signaled before the intra prediction mode, and in case a non-zero reference row index is signaled, the planar mode is excluded from the intra prediction mode. Fig.17 An example of four reference rows of neighboring prediction blocks is shown. MRL is disabled for the first row of a block within a CTU to prevent the use of extended reference samples outside the current CTU row. In addition, PDPC will be disabled when additional rows are used. For MRL mode, the derivation of the DC value in the DC intra prediction mode for non-zero reference row index is aligned with the derivation of reference row index 0. MRL requires 3 neighboring luma reference rows stored by the CTU to generate predictions. The Cross Component Linear Model (CCLM) tool also requires 3 neighboring luma reference rows for its downsampling filter. The definition of MRL using the same 3 rows is aligned with CCLM to reduce the storage requirements of the decoder. 2.15. Intra-frame sub-segmentation (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block into 2 or 4 sub-partitions vertically or horizontally depending on the block size. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block will be divided into 4 sub-partitions. It has been noted that M×128 (M≤64) and 128×N (N≤64) ISP blocks may generate potential problems with 64×64 VDPU. For example, an M×128 CU in a single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, then the luma TB will be split into 4 M×32TBs (only horizontal splitting is possible), each of which is smaller than a 64×64 block. However, in the current ISP design, the chroma blocks are not split. Therefore, the size of both chroma components will be larger than a 32×32 block. Similarly, using ISP with 128×NCU can create a similar situation. Therefore, these two situations are problems for a 64×64 decoder pipeline. Therefore, the CU size that can use ISP is limited to a maximum value of 64×64. Fig.18A and Fig.18B Examples of two possibilities are shown. All sub-divisions satisfy the condition of having at least 16 samples. In ISP, 1×N / 2×N sub-block predictions are not allowed to rely on the reconstructed values of previously decoded 1×N / 2×N sub-blocks of the codec block, so that the minimum prediction width of the sub-block becomes four samples. For example, an 8×N (N>4) codec block using ISP codec with vertical partitioning is divided into two prediction regions, each of size 4×N, and the size of the four transforms is 2×N. Similarly, a 4×N codec block using ISP codec with vertical partitioning is predicted using a full 4×N block; four transforms are used, each of 1×N. Although transform sizes of 1×N and 2×N are allowed, it can be asserted that the transforms of these blocks within the 4×N region can be performed in parallel. For example, when a 4×N prediction region contains four 1×N transforms, there is no transform in the horizontal direction; the transform in the vertical direction can be performed as a single 4×N transform in the vertical direction. Similarly, when a 4×N prediction region contains two 2×N transform blocks, the transform operations of two 2×N blocks in each direction (horizontally and vertically) can be performed in parallel. Therefore, there is no added latency when processing these smaller blocks compared to processing intra blocks for a 4x4 regular codec. Fig.18A and Fig.15 B shows the sub-division depending on the block size. Fig.18A Examples of sub-partitioning for 4x8 and 8x4 CUs are shown. Fig.18B Examples of sub-partitioning for CUs other than 4x8, 8x4, and 4x4 are shown. Table 5 Entropy coding and decoding coefficient group size Block size Coefficient group size 1×N,N≥16 1×16 N×1,N≥16 16×1 2×N,N≥8 2×8 N×2,N≥8 8×2 All other possible M×N situations 4×4 For each sub-division, the reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization, and inverse transformation. Therefore, the reconstructed sample values of each sub-division can be used to generate a prediction for the next sub-division, and each sub-division is processed repeatedly. In addition, the first sub-division to be processed is a word division that contains the upper left sample of the CU and then continues downward (horizontal division) or to the right (vertical division). Therefore, the reference samples used to generate the sub-division prediction signal are located only on the left and above the row. All sub-divisions share the same intra mode. The following is a summary of the interaction of ISP with other codec tools. – Multiple Reference Line (MRL): If the MRL index of a block is not 0, the ISP codec mode will be inferred to be 0, so the ISP mode information will not be sent to the decoder. – Entropy coding coefficient group size: As shown in Table 5, the size of the entropy coding sub-block has been modified to have 16 samples in all possible cases. It is worth noting that the new size only affects the ISP-generated blocks where one size is less than 4 samples. In all other cases, the coefficient group remains 4×4 in size. – CBF codec: It is assumed that at least one subpartition has a non-zero CBF. Thus, if n is the number of subpartitions, and the first n-1 subpartitions have yielded zero CBF, the CBF of the nth subpartition is inferred to be 1. – Transform size constraint: All ISP transforms with length greater than 16 points use DCT-II. –MTS flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0 and will not be sent to the decoder. Therefore, the encoder will not perform RD tests on the different available transforms for each result sub-split. The transform selection for ISP mode will instead be fixed and will be selected based on the intra mode used, the processing order and the block size. Therefore, no signaling is required. For example, let t H and t V are the horizontal and vertical transforms selected for the w×h sub-partition, respectively, where w is the width and h is the height. The transforms are then selected according to the following rules: - If w=1 or h=1, there is no horizontal transform or vertical transform, respectively. – If w ≥ 4 and w ≤ 16, t H =DST-VII, otherwise t H =DCT-II. – If h≥4 and h≤16, t V =DST-VII, otherwise t V =DCT-II. In ISP mode, all 67 intra prediction modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. In addition, the reference sample filtering process (reference smoothing) and the conditions for intra interpolation filtering selection no longer exist, and in ISP mode, the cubic (DCT-IF) filter is always applied for fractional position interpolation. 2.16. Matrix Weighted Intra Prediction (MIP) The matrix weighted intra prediction (MIP) method is a new intra prediction technique added to VVC. In order to predict the samples of a rectangular block of width W and height H, the matrix weighted intra prediction (MIP) takes a row of H reconstructed neighboring boundary samples on the left side of the block and a row of W reconstructed neighboring boundary samples above the block as input. If the reconstructed samples are not available, they are generated as in traditional intra prediction. Fig.19As shown, the generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication and linear interpolation. Fig.19 The matrix-weighted intra prediction process is shown. 2.16.1 Average neighboring sample points Among the boundary samples, four samples or eight samples are selected by averaging based on the block size and shape. Specifically, the boundary bdry is input by averaging the adjacent boundary samples according to a predefined rule depending on the size of the block. top and bdry left Reduced to a smaller boundary and Then, the two reduced boundaries and Splice to the reduced boundary vector bdry red , for blocks of 4×4 shape, its size is 4, and for blocks of all other shapes, its size is 8. If mode refers to a MIP mode, this stitching is defined as follows: 2.16.2 Matrix Multiplication Taking the averaged samples as input, we perform a matrix-vector multiplication and then add the offset. The result is a downscaled prediction signal over a subsampled set of samples in the original block. red In the above example, the reduced prediction signal pred is generated. red , the reduced prediction signal is of width W red and height H red Here, W red and H red is defined as: The reduced prediction signal pred is calculated by computing the matrix-vector product and adding an offset red : pred red =A.bdry red +b (2-13) Here, A is a matrix with W red ·H red rows, and if W=H=4, then it has 4 columns, in all other cases 8 columns. b is a red ·H red The matrix A and the offset vector b are taken from S 0 , S 1 , S 2 One of the sets. The index idx=idx(W,H) is defined as follows: Here, each coefficient of the matrix A is represented with 8 bits of precision. 0 By 16 matrices Each matrix has 16 rows and 4 columns, and 16 offset vectors Each offset vector has a size of 16. This set of matrices and offset vectors is for blocks of size 4×4. 1 Composed of 8 matrices Each matrix has 16 rows and 8 columns, and 8 offset vectors Each offset vector has a size of 16. The set S 2 The 6 matrices Each matrix has 64 rows and 8 columns, and 6 offset vectors Each offset vector has a size of 64. 2.16.3 Interpolation The prediction signals at the remaining positions are generated from the prediction signals on the subsampled set by linear interpolation, which is a single-step linear interpolation in each direction. The interpolation is performed first in the horizontal direction and then in the vertical direction, regardless of the shape of the block or the size of the block. 2.16.4.MIP mode signaling and coordination with other codec tools For each codec unit (CU) in intra mode, a flag indicating whether the MIP mode is to be applied is sent. If the MIP mode is to be applied, the MIP mode is signaled (predModeIntra). For the MIP mode, the transposed flag (isTransposed) (which determines whether the mode is transposed) and the MIP mode Id (modeId) (which determines which matrix to use for a given MIP mode) are derived as follows. isTransposed=predModeIntra&1 modeId=predModeIntra>>1 (2-15) The MIP codec mode is coordinated with other codecs by taking into account the following aspects: – Enable LFNST for MIPs on large blocks. Here the LFNST transform in planar mode is used. – The reference sample derivation for MIP is performed exactly as for the conventional intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled samples. – Perform clipping before upsampling instead of after upsampling. – Regardless of the maximum transform size, MIP is allowed to be up to 64×64. The number of MIP modes is 32 for sizeld=0, 16 for sizeld=1, and 12 for sizeld=2. 2.17. Decoder-side intra-mode derivation In JEM-2.0, intra modes are extended to 67 modes from 35 in HEVC, and they are derived at the encoder and explicitly signaled to the decoder. In JEM-2.0, a lot of overhead is spent on intra mode encoding and decoding. For example, in all intra codec configurations, the intra mode signaling overhead can be as high as 5% to 10% of the total bitrate. This paper proposes an intra mode derivation method on the decoder side to reduce the intra mode encoding and decoding overhead while maintaining prediction accuracy. In order to reduce the overhead of intra mode signaling, this paper proposes a decoder-side intra mode derivation (DIMD) method. In the proposed method, instead of explicitly signaling the intra mode, the information is derived from the neighboring reconstructed samples of the current block at the encoder and decoder. The intra mode derived by DIMD can be used in two ways: 1) For a 2N×2N CU, when the corresponding CU-level DIMD flag is turned on, the DIMD mode is used as the intra mode for intra prediction; 2) For an N×N CU, the DIMD mode is used to replace a candidate in the existing MPM list to improve the efficiency of intra-mode encoding and decoding. 2.17.1. Template-based intra-mode derivation like Fig. 20 As shown, the target represents the current block (block size is N) whose intra prediction mode is to be estimated. The template (composed of Fig. 20 The template size is expressed as the number of samples in the template that extend above and to the left of the target block, i.e., L. In the current implementation, the template size used for 4×4 and 8×8 blocks is 2 (i.e., L=2), and the template size used for 16×16 and larger blocks is 4 (i.e., L=4). As defined in JEM-2.0, the reference of the template (given by Fig. 20 The dotted area in the figure refers to the set of neighboring samples above and to the left of the template. Unlike the template samples that are always from the reconstructed area, the reference samples of the template may not have been reconstructed when encoding / decoding the target block. In this case, the existing reference sample replacement algorithm of JEM-2.0 is used to replace the unavailable reference samples with available reference samples. Fig. 20 Target points, template points, and reference points of a template used in DIMD are shown. For each intra prediction mode, DIMD calculates the absolute difference (SAD) between the reconstructed template samples and its predicted samples obtained from the reference samples of the template. The intra prediction mode that produces the smallest SAD is selected as the final intra prediction model for the target block. 2.17.2. DIMD for Intra 2N×2N CU For intra 2N×2N CUs, DIMD is used as an additional intra mode, which is adaptively selected by comparing the DIMD intra mode with the best normal intra mode (i.e., explicitly signaled). For each intra 2N×2N CU, a flag is signaled to indicate the use of DIMD. If the flag is 1, the intra mode derived by DIMD is used to predict the CU; otherwise, DIMD is not applied and the intra mode explicitly signaled in the bitstream is used to predict the CU. When DIMD is enabled, chroma components always reuse the same intra mode derived for luma components, i.e., DM mode. In addition, for each DIMD-encoded CU, blocks in the CU can adaptively choose to derive their intra modes at the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is signaled to indicate at which level DIMD is performed. If the flag is 0, it means that DIMD is performed at the PU level, and all TUs in the PU use the same derived intra mode for intra prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is performed at the TU level, and each TU in the PU derives its own intra mode. In addition, when DIMD is enabled, the number of angular directions increases to 129, and DC mode and planar mode remain unchanged. To accommodate the increased granularity of angular intra modes, the precision of intra interpolation filtering for DIMD-coded CUs is increased from 1 / 32 pixel to 1 / 64 pixel. In addition, in order to use the derived intra mode of a DIMD-coded CU as an MPM candidate for neighboring intra blocks, the 129 directions of the DIMD-coded CU are converted to "normal" intra mode (i.e., 65 angular intra directions) before being used as MPM. 2.17.3. DIMD for Intra N×N CU In the proposed method, the intra mode of an N×N CU is always signaled. However, to improve the efficiency of intra mode encoding and decoding, the intra mode derived from DIMD is used as the MPM candidate for predicting the intra mode of the four PUs in the CU. In order not to increase the overhead of MPM index signaling, the DIMD candidate is always placed first in the MPM list and the last existing MPM candidate mode is deleted. In addition, deduplication is performed so that if a DIMD candidate is redundant, it is not added to the MPM list. 2.17.4.DIMD Intra-frame Pattern Search Algorithm In order to reduce the complexity of encoding / decoding, a direct and fast intra mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra mode search. Specifically, an initial candidate list is created by selecting N fixed modes from the allowed intra modes. Then, the SAD is calculated for all candidate intra modes, and the intra mode that minimizes the SAD is selected as the starting intra mode. In order to achieve a good complexity / performance trade-off, the initial candidate list consists of 11 intra modes, including DC, planar, and every 4th mode in the 33 angular intra directions defined in HEVC, i.e., intra modes 0, 1, 2, 6, 10…30, 34. If the starting Intra mode is DC mode or Planar mode, it is used as the DIMD mode. Otherwise, based on the starting Intra mode, a refinement process is then applied where the best Intra mode is identified through an iterative search. It works by comparing the SAD values of three Intra modes separated by a given search interval at each iteration and keeping the Intra mode that minimizes the SAD. The search interval is then reduced to half and the Intra mode selected from the previous iteration is used as the center Intra mode for the current iteration. For the current DIMD implementation with 129 angular Intra directions, a maximum of 4 iterations are used in the refinement process to find the best DIMD Intra mode. 2.18. Decoder-side Intra-mode Derivation by Computing Gradients of Neighboring Samples Three angular modes are selected from the Histogram of Gradients (HoG) calculated from the neighboring pixels of the current block. Once these three modes are selected, their predictions are calculated normally and then their weighted average is used as the final prediction for the block. To determine the weights, for each of the three modes, the corresponding magnitude in the HoG is used. The DIMD mode is used as an alternative prediction mode, always checked in FullRD mode. The current version of DIMD has been modified in some aspects in signaling, HoG calculation and prediction fusion. The purpose of this modification is to improve the codec performance and address the complexity issues raised during the last meeting (i.e., throughput of 4×4 blocks). The following sections describe the modifications in each aspect. 2.18.1. Signaling Fig.21 The order of parsing flags / indexes integrated with the proposed DIMD in VTM5 is shown. It can be seen that the DIMD flag of the block is first parsed using a single CABAC context, which is initialized to a default value of 154. If flag == 0, parsing will continue normally. Otherwise (if flag == 1), only the ISP index is parsed, and the following flags / indexes are inferred to be zero: BDPCM flag, MIP flag, MRL index. In this case, the entire IPM parsing is also skipped. During the parsing phase, when a regular non-DIMD block queries the IPM of its DIMD neighbors, the pattern PLANAR_IDX is used as a virtual IPM for the DIMD block. Fig.21 The proposed intra-block decoding process is shown. 2.18.2. Texture analysis DIMD's texture analysis includes the calculation of the Histogram of Gradients (HoG) ( Fig. 22 ). Gradient histogram calculation is performed by applying horizontal and vertical Sobel filters to the pixels in a template of width 3 around the block. However, if the above template pixels belong to different CTUs, they will not be used for texture analysis. Once calculated, the IPMs corresponding to the two highest histogram bins are selected for the block. In previous versions, all pixels in the middle row of the template participated in the HoG calculation. However, the current version improves the throughput of this process by applying the Sobel filter more sparsely over 4x4 blocks. For this purpose, only one pixel from the left and one pixel from the top are used. Fig. 22 shown. Besides reducing the number of operations for gradient computation, this feature also simplifies the selection of the best 2 patterns from the HoG, since the resulting HoG cannot have more than two non-zero magnitudes. Fig. 22 The HoG computation from a template with a width of 3 pixels is shown. 2.18.3. Prediction Fusion The current version of the method also uses a fusion of three prediction values for each block. However, the choice of prediction mode is different and uses the combined hypothetical intra prediction method proposed in [2], where planar mode is considered to be combined with other modes when calculating intra prediction candidates. In the current version, the two IPMs corresponding to the two highest HoG stripes are combined with planar mode. Prediction fusion is applied as a weighted average of the three predictions above. For this purpose, the weight of the plane is fixed to 21 / 64 (~1 / 3). The remaining 43 / 64 (~2 / 3) of the weight is then shared between the two HoG IPMs in proportion to the amplitude of their HoG stripes. Fig.23 Shows this process. Fig.23 The prediction fusion by weighted averaging of two HoG modes and planes is shown. 2.19.DIMD Chroma Mode like Fig.24 As shown in , the DIMD chroma mode uses the DIMD derivation method to derive the chroma intra prediction mode of the current block based on the adjacent reconstructed Y, Cb and Cr samples in the second adjacent rows and columns. Specifically, for each co-located reconstructed luminance sample and reconstructed Cb and Cr samples of the current chroma block, the horizontal gradient and the vertical gradient are calculated to construct the HoG. Then, the intra prediction mode with the largest histogram amplitude value is used to perform chroma intra prediction of the current chroma block. Fig.24 Neighboring reconstructed samples for DIMD chroma mode are shown. When the intra prediction mode derived from the DIMD chroma mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram magnitude value is used as the DIMD chroma mode. A CU level flag is signaled to indicate whether the proposed DIMD chroma mode is applied. 2.20. Template-based Intra Mode Derivation (TIMD) This paper proposes a template-based intra mode derivation (TIMD) method using MPM, where the TIMD mode is derived from the MPM using neighboring templates. The TIMD mode is used as an additional intra prediction method for a CU. 2.20.1. TIMD mode derivation For each intra prediction mode in the MPM, the SATD between the prediction and reconstructed samples of the template is calculated. The intra prediction mode with the smallest SATD is selected as the TIMD mode and is used for intra prediction of the current CU. Position-dependent intra prediction combination (PDPC) is included in the derivation of the TIMD mode. 2.20.2.TIMD Signaling A flag is signaled in the sequence parameter set (SPS) to enable / disable the proposed method. When the flag is true, a CU level flag is signaled to indicate whether the proposed TIMD method is used. The TIMD flag is signaled after the MIP flag. If the TIMD flag is equal to true, the remaining syntax elements related to the luma intra prediction mode, including MRL, ISP, and the normal parsing stage for the luma intra prediction mode are skipped. 2.20.3. Interaction with new codec tools The DIMD method using planar prediction fusion is integrated into EE2. When the EE2 DIMD flag is true, the proposed TIMD flag is not signaled and is set to false. Similar to PDPC, gradient PDPC is also included in the derivation of TIMD mode. When the auxiliary MPM is enabled, both the primary MPM and the auxiliary MPM are used to derive the TIMD mode. The 6-tap interpolation filter is not used for the derivation of TIMD mode. 2.20.4. Modification of MPM list construction in TIMD mode derivation During the construction of the MPM list, the intra prediction mode of the neighboring blocks is derived as a plane when it is inter-coded. In order to improve the accuracy of the MPM list, when the neighboring blocks are inter-coded, the propagated intra prediction mode is derived using the motion vector and the reference picture and is used in the construction of the MPM list. This modification is only used for the derivation of the TIMD mode. 2.20.5. TIMD with Fusion This paper proposes that, for the intra-frame mode derived using the TIMD method, instead of selecting only one mode with the smallest SATD cost, the first two modes with the smallest SATD cost are selected, and then they are fused through weights, and this weighted intra-frame prediction is used to encode and decode the current CU. The costs of the two selected modes are compared to the threshold, applying a cost factor of 2 in the test as follows: costMode2<2′costMode1. If this condition is true, then fusion is applied, otherwise only mode1 is used. The weight of a mode is calculated based on its SATD cost as follows: weight1=costMode2 / (costMode1+costMode2c); weight2=1-weight1. 2.21. Merge Mode with MVD (MMVD) In addition to the Merge mode, a Merge mode with motion vector difference (MMVD) is introduced in VVC when implicitly derived motion information is directly used for prediction sample generation of the current CU. The MMVD flag is signaled immediately after the regular Merge flag is sent to specify whether the MMVD mode is used for the CU. In MMVD, after a Merge candidate is selected, it is further refined by the MVD information transmitted by the signal. Further information includes a Merge candidate flag, an index for specifying the magnitude of motion, and an index for indicating the direction of motion. In MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The MMVD candidate flag is transmitted by the signal to specify which one to use between the first Merge candidate and the second Merge candidate. Fig.25 MVD search points are shown. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Fig.25 As shown, the offset is added to the horizontal component or the vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 6. Table 6 Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown in Table 7. It should be noted that the meaning of the MVD symbol may vary depending on the information of the starting MV. When the starting MV is a non-predicted MV or a bidirectionally predicted MV with two lists pointing to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 7 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectionally predicted MV with two MVs to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), and the POC difference in list 0 is greater than the POC difference in list 1, the symbol in Table 7 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Otherwise, if the POC difference in list 1 is greater than the POC difference in list 0, the symbol in Table 7 specifies the sign of the MV offset added to the list 1 MV component of the starting MV, and the sign of the list 0 MV has the opposite value. The MVD is scaled based on the difference in POC in each direction. If the POC difference in both lists is the same, no scaling is required. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, then Fig.26 As described in , the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb. If the POC difference of L1 is greater than the POC difference of L0, the MVD of list 0 is scaled in the same way. If the starting MV is unidirectionally predicted, the MVD is added to the available MVs. Table 7 Signs of MV offsets specified by direction index Directions to IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.22. Symmetric MVD encoding and decoding In VVC, in addition to conventional unidirectional prediction mode MVD signaling and bidirectional prediction mode MVD signaling, symmetric MVD mode is applied for bidirectional prediction MVD signaling. In symmetric MVD mode, motion information including reference picture indices of both list 0 and list 1 and MVD of list 1 is not transmitted by signal but derived. The decoding process of the symmetric MVD mode is as follows: 1. At the stripe level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. – Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2. At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. The reference indexes of list 0 and list 1 are set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown in the following formula. Fig.26 Symmetrical MVD mode is shown. In the encoder, symmetric MVD motion estimation starts with initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The one with the lowest distortion cost is selected as the initial MV for symmetric MVD motion search. 2.23. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, formerly known as BIO, is included in JEM. Compared to the JEM version, BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if all of the following conditions are met: - The CU is encoded using "true" bi-prediction mode, ie, one of the two reference pictures precedes the current picture in display order, and the other of the two reference pictures follows the current picture in display order. – The distances from the two reference images to the current image (i.e., POC difference) are the same. – Both reference images are short-term reference images. –CU is not encoded or decoded using affine mode or SbTMVPMerge mode. –CU has more than 64 luma samples. – CU height and CU width are both greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. – The current CU does not have WP enabled. –CIIP mode is not used for the current CU. BDOF is applied only to the luma component. As its name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4×4 sub-block, motion refinement (v x ,v y ) is calculated by minimizing the difference between L0 prediction samples and L1 prediction samples. Then motion refinement is used to adjust the bi-prediction sample values in the 4x4 sub-block. The following steps are applied in the BDOF process. First, by directly calculating the difference between two adjacent samples, the horizontal and vertical gradients of the two predicted signals, and is calculated, that is, Among them I (k) (i, j) is the sample value at the prediction signal coordinate (i, j) in list k, k=0,1, and shift1 is calculated based on the luma bit depth bitDepth as shift1=max(6, bitDepth-6). Then, the gradient S 1 ,S 2 ,S 3 ,S 5 and S 6 The autocorrelation and cross-correlation are calculated as follows in where Ω is a 6×6 window around a 4×4 sub-block, and na and n b The values are set to min(1, bitDepth-11) and min(4, bitDepth-8) respectively. Then, using the cross-correlation and autocorrelation terms, motion refinement (v x ,v y ) is derived using the following method: in th′ BIO =2 max(5,BD-7) . is the floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as shown below: pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift (2-22) These values are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits. In order to derive the gradient value, some prediction samples I in the list k (k = 0, 1) outside the current CU boundary (k) (i,j) needs to be generated. Fig.25 As depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and a conventional 8-tap motion compensated interpolation filter is used to generate prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculations. For the rest of the steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e. repeated) from their nearest neighbors. Fig. 27 The extended CU area used in BDOF is shown. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are considered as CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set equal to (8*W*(H>>1), where W represents the sub-block width and H represents the sub-block height. In order to avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated in the DVMR process is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is 1 for either of the two reference pictures, BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded using symmetric MVD mode or CIIP mode. 2.24. Combined Inter and Intra Prediction (CIIP) 2.25. Affine Motion Compensated Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Fig.28 As shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). Fig.28 The affine motion model based on control points is shown. Fig.28 In , both a 4-parameter affine model and a 6-parameter affine model are shown. For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: Where (mv 0x ,mv 0y ) is the motion vector of the upper left control point, (mv 1x ,mv 1y ) is the motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the motion vector of the lower left control point. To simplify motion compensation prediction, block-based affine transformation prediction is applied. To derive the motion vector of each 4×4 luminance sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation (e.g. Fig.28 The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the 4 corresponding 4×4 luminance subblocks. Fig.29 The affine MVF of each sub-block is shown. Like translational motion inter prediction, there are two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.25.1. Affine Merge Prediction AF_MERGE mode can be applied to CUs with width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of the spatially adjacent CUs. There can be up to five CPMVP candidates, and the one to be used for the current CU is indicated by a signal transmission index. The following three types of CPVM candidates are used to form the affine Merge candidate list: – Inherited affine merge candidates inferred from the CPMV of neighboring CUs. – The constructed affine Merge candidate CPMVP derived using the translation MV of the neighboring CU. – Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Fig.30 As shown. For the prediction value on the left, the scanning order is A0->A1, and for the prediction value on the top, the scanning order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. As shown Fig.30 As shown, if the adjacent lower left block A is encoded and decoded in affine mode, the motion vectors v of the upper left corner, upper right corner, and lower left corner of the CU containing block A are obtained. 2 ,v 3 and v 4 When block A is encoded and decoded using a 4-parameter affine model, according to v 2 and v 3 Calculate the two CPMVs of the current CU. When block A is encoded and decoded using a 6-parameter affine model, according to v 2 ,v3 and v 4 Calculate the three CPMVs of the current CU. Fig.30 The positions of the inherited affine motion prediction values are shown. Fig.31 Control point motion vector inheritance is shown. The constructed affine candidate is constructed by combining the neighboring translation motion information of each control point. The motion information of the control point is obtained from Fig.31 The specified spatial and temporal neighbors shown in are derived, CPMV k (k=1,2,3,4) represents the kth control point. 1 , the B2->B3->A2 blocks are checked and the MV of the first available block is used. For CPMV 2 , the B1->B0 block is checked, and for CPMV 3 , the A1->A0 block is checked. TMVP is used as CPMV 4 (if available). After obtaining the MVs of the four control points, an affine Merge candidate is constructed based on the motion information. The following combinations of control point MVs are used to construct in order: {CPMV 1 , CPMV 2 , CPMV 3}, {CPMV 1 , CPMV 2 , CPMV 4}, {CPMV 1 , CPMV 3 , CPMV 4}, {CPMV 2 , CPMV 3 , CPMV 4}, {CPMV 1 , CPMV 2}, {CPMV 1 , CPMV 3}. The combination of 3 CPMVs constructs a 6-parameter affine Merge candidate, and the combination of 2 CPMVs constructs a 4-parameter affine Merge candidate. To avoid the motion scaling process, the relevant combination of control point MVs is discarded if the reference indexes of the control points are different. Fig.32 The positioning of candidate positions for constructing the affine merge pattern is shown. After checking the inherited affine merge candidates and the constructed affine merge candidates, if the list is still not full, a zero MV is inserted at the end of the list. 2.25.2. Affine AMVP Prediction Affine AMVP mode can be applied to CUs with width and height both greater than or equal to 16. A CU-level affine flag is signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2 and is generated by using the following four types of CPVM candidates in order: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs. – The constructed affine AMVP candidate CPMVP derived using the translated MV of the neighboring CU. – Translated MV from neighboring CU. – Zero MV. The order in which inherited affine AMVP candidates are checked is the same as that of inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When the inherited affine motion prediction value is inserted into the candidate list, the deduplication process is not applied. The AMVP candidates constructed from Fig.15 The specified spatial neighbors are derived as shown. The same check order is used as in the affine merge candidate construction. In addition, the reference picture index of the neighboring blocks is checked. The first block in the check order is used, which is inter-coded and has the same reference picture as the current CU. There is only one. When the current CU is coded using the 4-parameter affine mode and mv 0 and mv 1 When all three CPMVs are available, they are added as a candidate in the affine AMVP list. When the current CU is encoded and decoded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable. If the affine AMVP list of candidates is still less than 2 after inserting the valid inherited affine AMVP candidates and the constructed AMVP candidates, mv 0 ,mv 1 and mv 2 Will be added as translation MVs in order to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, the list is filled with zero MVs. 2.25.3 Affine Motion Information Storage In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is only used to generate the inherited CPMV in affine Merge mode and the inherited CPMV in affine AMVP mode for the most recently encoded CU. The sub-block MV derived from the CPMV is used for motion compensation, MV derivation of the Merge / AMVP list of translation MVs, and deblocking. To avoid picture row cache for additional CPMV, affine motion data inheritance from CUs above the CTU is handled differently than inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the row above the CTU, the bottom left and bottom right sub-block MVs in the row cache are used for affine MVP derivation instead of CPMVs. In this way, CPMVs are only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to a 4-parameter model. Fig.16 As shown, along the top boundary of the CTU, the bottom left and bottom right sub-block motion vectors of the CU are used for affine inheritance of the CU in the bottom of the CTU. Fig.33 A diagram showing the motion vector usage of the proposed combined approach. 2.25.4. Prediction refinement using optical flow for affine mode Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the expense of prediction accuracy. In order to achieve finer motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine sub-block-based affine motion compensation predictions without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the brightness prediction samples are refined by adding the differences derived from the optical flow equation. PROF is described as the following four steps: Step 1) Sub-block based affine motion compensation is performed to generate a sub-block prediction I(i,j). Step 2) Use a 3-tap filter [-1, 0, 1], the spatial gradient g of the sub-block prediction x (i,j) and g y (i,j) is calculated at each sample point. The gradient calculation is exactly the same as the gradient calculation in BDOF. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j)>>shift1) (2-25) g y (i,j)=(I(i,j+1)>>shift1)-(I(i,j-1)>>shift1) (2-26) Shift1 is used to control the accuracy of the gradient. The sub-block (i.e. 4x4) prediction is extended by one sample on each side of the gradient calculation. To avoid extra memory bandwidth and extra interpolation calculations, those extended samples on the extended boundary are copied from the nearest integer pixel position in the reference picture. Step 3) The brightness prediction refinement is calculated by following the optical flow equation. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i, j) (2-27) Among them Fig.34 As shown, Δv(i,j) is the difference between the sample MV calculated for the sample position (i,j), denoted as v(i,j), and the sub-block MV of the sub-block to which the sample (i,j) belongs. Δv(i,j) is quantized in units of 1 / 32 luma sample accuracy. Fig.34 The sub-block MV V is shown SB and pixel Δv(i,j) (indicated by the grey arrow). Since the affine model parameters and the sample position relative to the subblock center do not change from subblock to subblock, Δv(i,j) can be calculated for the first subblock and reused for other subblocks in the same CU. Let dx(i,j) and dy(i,j) be the distance from the sample position (i,j) to the subblock center (x SB ,y SB )’s horizontal and vertical offsets, Δv(x,y) can be derived by the following equations. To maintain accuracy, the sub-block (x SB ,y SB ) is calculated as ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the width and height of the sub-block respectively. For the 4-parameter affine model, For the 6-parameter affine model, Where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y) are the control point motion vectors of the upper left, upper right and lower left, and w and h are the width and height of the CU. Step 4) Finally, the luma prediction refinement ΔI(i,j) is added to the sub-block prediction I(i,j). The final prediction I' is generated as follows. I'(i,j)=I(i,j)+ΔI(i,j) (2-32) PROF is not applicable to affine codec CUs in two cases: 1) all control point MVs are the same, which indicates that the CU has only translational motion; 2) the affine motion parameters are larger than the specified limit, because the sub-block based affine MC is downgraded to CU based MC to avoid large memory access bandwidth requirements. A fast codec method is applied to reduce the codec complexity of affine motion estimation using PROF. PROF is not applied to the affine motion estimation stage in the following two cases: a) If the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, PROF is not applied because the possibility that the current CU selects the affine mode as the best mode is low; b) If the magnitudes of the four affine parameters (C, D, E, F) are all less than a predefined threshold and the current picture is not a low-latency picture, PROF is not applied because the improvement introduced by PROF for this case is small. In this way, affine motion estimation using PROF can be accelerated. 2.26. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode of the CU in the current picture. The same co-located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects: –TMVP predicts CU-level motion, but SbTMVP predicts CU-level motion; – While TMVP pre-fetches the temporal motion vector from the co-located block in the co-located picture (the co-located block is the lower right block or the center block relative to the current CU), SbTMVP applies motion offset before pre-fetching temporal motion information from the co-located picture. Where the motion offset is obtained from the motion vector of one of the spatially neighboring blocks of the current CU. The SbTVMP process Fig.18A SbTMVP predicts the motion vector of the sub-CU within the current CU in two steps. In the first step, check Fig.18AThe spatial neighbor A1 in it. If A1 has a motion vector using the co-located picture as its reference picture, then this motion vector is selected as the motion offset to be applied. If no such motion is identified, the motion offset is set to (0, 0). In the second step, the motion offset identified in step 1 is applied (i.e., added to the coordinates of the current picture) to obtain sub-CU level motion information (motion vector and reference index) from the co-located picture as Fig.18B shown. Fig.18B The example in it assumes that the motion offset is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample point) in the co-located picture is used to derive the motion information of the sub-CU. After identifying the motion information of the co-located sub-CU, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. Fig.35A and Fig.35B show the SbTMVP process in VVC. Fig.35A shows the spatial neighboring blocks used by ATVMP. Fig.35B shows the derivation of the sub-CU motion field by applying the motion offset from the spatial neighboring CU and scaling the motion information from the corresponding co-located sub-CU. In VVC, a combined sub-block based Merge list containing both SbTMVP candidates and affine Merge candidates is used for signaling of the sub-block based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry of the list of sub-block based Merge candidates, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8x8, and like the affine Merge mode, the SbTMVP mode only applies to CUs with both width and height greater than or equal to 8. The encoding / decoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P or B slice, additional RD checks are performed to decide whether to use the SbTMVP candidate. 2.27 Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the CU's motion vector and the predicted motion vector) is transmitted by signal in units of quarter-luminance samples. In VVC, the CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the CU's MVD to be encoded and decoded with different precisions. Depending on the current CU mode (normal AMVP mode or affine AVMP mode), the current CU's MVD can be adaptively selected as follows: – Normal AMVP mode: quarter brightness samples, half brightness samples, full brightness samples or four brightness samples. – Affine AMVP mode: quarter brightness samples, full brightness samples, or 1 / 16 brightness samples. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e. both horizontal and vertical MVD for reference list L0 and reference list L1) are zero, then quarter luma sample MVD resolution is inferred. For a CU with at least one non-zero MVD component, the first flag is signaled to indicate whether quarter luma sample MVD precision is used for the CU. If the first flag is 0, no further signal is required, and quarter luma sample MVD precision is used for the current CU. Otherwise, the second flag is signaled to indicate that half luma samples or other MVD precision (integer or four luma samples) are used for normal AMVP CUs. In the case of half luma samples, a 6-tap interpolation filter is used for the half luma sample position instead of the default 8-tap interpolation filter. Otherwise, the third flag is signaled to indicate whether whole luma samples or four luma sample MVD precision is used for normal AMVP CUs. In the case of an affine AMVP CU, the second flag is used to indicate whether whole luma sample MVD precision or 1 / 16 luma sample MVD precision is used. In order to ensure that the reconstructed MV has the expected precision (quarter luma sample, half luma sample, whole luma sample, or four luma samples), the motion vector prediction value of the CU will be rounded to the same precision as the MVD before being added to the MVD. Motion vector prediction values are rounded towards zero (that is, negative motion vector prediction values are rounded towards positive infinity, and positive motion vector prediction values are rounded towards negative infinity). The encoder uses RD check to determine the motion vector resolution of the current CU. In order to avoid always performing four CU-level RD checks for each MVD resolution, in VTM11, only the RD check of MVD precision other than quarter-luminance samples is conditionally called. For normal AVMP mode, the RD cost of quarter-luminance sample MVD precision and the RD cost of full-luminance sample MV precision are first calculated. Then, the RD cost of full-luminance sample MVD precision is compared with the RD cost of quarter-luminance sample MVD precision to determine whether it is necessary to further check the RD cost of four-luminance sample MVD precision. When the RD cost of quarter-luminance sample MVD precision is much smaller than the RD cost of full-luminance sample MVD precision, the RD check of four-luminance sample MVD precision is skipped. Then, if the RD cost of full-luminance sample MVD precision is significantly greater than the best RD cost of the previously tested MVD precision, the check of half-luminance sample MVD precision is skipped. For affine AMVP mode, if affine inter mode is not selected after checking the rate-distortion cost of affine Merge / Skip mode, Merge / Skip mode, quarter luma sample MVD accuracy normal AMVP mode, and quarter luma sample MVD accuracy affine AMVP mode, 1 / 16 luma sample MV accuracy and 1 pixel MV accuracy affine inter mode are not checked. In addition, in 1 / 16 luma sample and quarter luma sample MV accuracy affine inter mode, the affine parameters obtained in quarter luma sample MV accuracy affine inter mode are used as the starting search point. 2.28 Bidirectional Prediction with CU Level Weights (BCW) In HEVC, a bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow a weighted average of the two prediction signals. P bi-prcd ((8-w)*P 0 +w*P 1 +4)>>3 (2-33) Five weights are allowed in weighted average bidirectional prediction, w∈{-2,3,4,5,10}. For each bidirectional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CU, the weight index is transmitted by signal after the motion vector difference; 2) For Merge CU, the weight index is inferred from the neighboring block based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., the CU width multiplied by the CU height is greater than or equal to 256). For low-latency pictures, all 5 weights will be used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized below. When combined with AMVR, unequal weights for 1-pixel and 4-pixel motion vector precision are only conditionally checked if the current picture is a low-latency picture. - When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only checked conditionally. – Do not search for unequal weights when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using one context encoded bit followed by a bypass encoded bit. The first context encoded bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is signaled using the bypass codec to indicate which unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficient encoding and decoding of video content in fading conditions. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in list L0 and list L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. In order to avoid interaction between WP and BCW (which will complicate the VVC decoder design), if the CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For the constructed affine Merge mode, affine motion information is constructed based on motion information of up to 3 blocks. The BCW index of a CU using the constructed affine merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded and decoded using CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.29. Local Illumination Compensation (LIC) Local illumination compensation (LIC) is a codec tool used to address the problem of local illumination changes between a current picture and its temporal reference picture. LIC is based on a linear model where a scaling factor and an offset are applied to a reference sample to obtain a predicted sample for the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x,y)=α·P r (x+v x ,y+v y )+β Where P(x,y) is the prediction signal of the current block at the coordinate (x,y); P r (x+v x ,y+v y ) is the motion vector (v x ,v y ) points to the reference block; α and β are the corresponding scaling factor and offset applied to the reference block. Fig.19 The LIC process is shown in FIG. Fig.19 In , when LIC is applied to a block, the minimum mean square error (LMSE) method is used to minimize the neighboring samples of the current block (i.e., Fig.19 The template T in the temporal reference picture) and its corresponding reference sample point in the temporal reference picture (ie, Fig.19 In addition, in order to reduce the computational complexity, both the template samples and the reference template samples are subsampled (adaptive subsampling) to derive the LIC parameters, i.e., only Fig.19 α and β are derived from the shaded points in . In order to improve the encoding and decoding performance, such as Fig. 20 As shown, the short sides are not subsampled. Fig.36 Local illumination compensation is shown. Fig.37 It is shown that no subsampling is performed for the short edges. 2.30 Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of Merge mode MV, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bidirectional prediction operation, the refined MV is searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. Fig.21 As shown, the SAD between two blocks based on each MV candidate (eg, MV0' and MV1') around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate a bidirectional prediction signal. Fig.38Decoding side motion vector refinement is shown. In VVC, the application of DMVR is restricted and is only applied to CUs that are coded and decoded using the following modes and functions: – CU level Merge mode with bi-prediction MV. – Relative to the current picture, one reference picture is in the past and the other reference picture is in the future. – The distances from the two reference pictures to the current picture (i.e., POC difference) are the same. – Both reference images are short-term reference images. –CU has more than 64 luma samples. – CU height and CU width are both greater than or equal to 8 luma samples. – BCW weight index indicates equal weight. – The current block is not WP enabled. –CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used for temporal motion vector prediction for future picture encoding and decoding. The original MV is used for the deblocking process and is also used for spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-items. 2.30.1. Search Scheme In DVMR, the search point is around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point examined by DMVR represented by the candidate MV pair (MV0, MV1) follows the following two equations: MV0′=MV0+MV_offset (2-34) MV1′=MV1-MV_offset (2-35) Wherein, MV_offset represents the refinement offset between the initial MV and the refinement MV in one of the reference pictures. The refinement search range is two integer luminance samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. The integer sample offset search uses a 25-point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. In order to reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV in the DMVR process. The SAD between the reference blocks pointed to by the initial MV candidate reference is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation instead of performing an additional search using SAD comparisons. Fractional sample refinement is conditionally called based on the output of the integer sample search stage. Fractional sample refinement is further applied when the integer sample search stage ends with the center with the minimum SAD in the first or second iteration search. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y)=A(x min ) 2 +B(yy min ) 2 +C (2-36) Where (x min ,y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of the five search points, (x min ,y min ) is calculated as: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-37) y min =(E(0,·-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-38) x min and min The value of is automatically clamped between -8 and 8, since all cost values are positive, and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to obtain a sub-pixel accurate refinement delta MV. 2.30.2. Bilinear interpolation and sample filling In VVC, the resolution of MV is 1 / 16 luma sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so these fractional position samples need to be interpolated for the DMVR search process. In order to reduce the computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect is that by using a bilinear filter, DVMR does not access more reference samples than the normal motion compensation process within the 2 sample search range. After the refined MV is obtained through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. In order not to access more reference samples for the normal MC process, samples will be filled from those available samples, which are not required for the interpolation process based on the original MV, but are required for the interpolation process based on the refined MV. 2.30.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16. 2.31. Multi-pass decoder-side motion vector refinement In this contribution, multi-pass decoder side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the coding block. In the third pass, MVs in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction. 2.31.1. First pass - Block-based bilateral matching MV refinement In the first pass, refined MVs are derived by applying BM to the codec blocks. Similar to decoder-side motion vector refinement (DMVR), refined MVs are searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. Refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive the integer sample precision intDeltaMV and the half-pixel sample precision halfDeltaMV. The local search applies a 3×3 square search pattern to loop in the horizontal search range [–sHor, sHor] and the vertical search range [–sVer, sVer], where the values of sHor and sVer are determined by the block size and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between the reference blocks. When the bilCost of the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV or halfDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern and continues to search for the minimum cost until it reaches the end of the search range. The current fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as: MV0_pass1 = MV0 + deltaMV; ·MV1_pass1=MV1–deltaMV. 2.31.2 Second pass—sub-block based bilateral matching MV refinement In the second pass, refined MVs are derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the refined MVs are searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM performs a full search to derive the integer sample point precision intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and a search range of [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying the cost factor to the SATD cost between the two reference sub-blocks, such as: bilCost = satdCost * costFactor. The search area (2*sHor+1)*(2*sVer+1) is divided into 5 diamond search areas, such as Fig.28As shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order starting from the upper left corner of the area to the lower right corner. When the minimum bilCost in the current search area is less than a threshold equal to sbW*sbH, the integer pixel full search is terminated, otherwise, the integer pixel full search continues to the next search area until all search points are checked. Fig.39 The diamond regions in the search area are shown. BM performs a local search to derive the half-sample accuracy halfDeltaMv. The search pattern and cost function are the same as defined in Section 2.9.1. The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). Then, the refined MV of the second pass is derived as: ·MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2); ·MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2). 2.31.3. Third pass - sub-block-based bidirectional optical flow MV refinement In the third pass, refined MVs are derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied starting from the refined MV of the parent sub-block of the second pass to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32. The refined MVs of the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as: ·MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv; ·MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)–bioMv. 2.32. Sample-based BDOF In sample-based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample. The codec block is partitioned into 8×8 sub-blocks. For each sub-block, it is determined whether to apply BDOF by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to a sub-block, for each sample in the sub-block, a sliding 5x5 window is used and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the center sample of the window. 2.33. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: 1) Spatial MVP from spatially adjacent CUs. 2) Temporal MVP from the co-located CU. 3) History-based MVP from FIFO table. 4) Average MVP in pairs. 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU code in Merge mode, the index of the best Merge candidate is encoded using truncated unary binarization (TU). The first bin of the Merge index is coded using context, while bypass coding is used for other bins. The derivation process of Merge candidates for each category is provided in this section. As operated in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a region of a certain size. 2.33.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. At most four Merge candidates are selected from the candidates at the indicated positions. The derivation order is B 0 , A 0 , B 1 , A 1 and B 2 Only when position B 0 , A 0 , B 1 and A 1 Position B is considered only when one or more CUs are not available (for example, because they belong to another slice or slice) or are intra-coded. 2 . Added position A 1After the candidates at the end, the remaining candidates are added for redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving the encoding and decoding efficiency. In order to reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only Fig.24 The pairs are linked by arrows in , and candidates are added to the list only if the corresponding candidates used for redundancy checking do not have the same motion information. Fig.40 The positions of spatial merge candidates are shown. Fig.41 Candidate pairs considered for redundancy check of spatial merge candidates are shown. 2.33.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. In particular, in the derivation of the temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. Fig.25 As shown by the dotted line in , the scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set equal to zero. Fig.42 An illustration of motion vector scaling for temporal Merge candidates is shown. like Fig.26 As shown, the position of the time domain candidate is in the candidate C 0 With C 1 If position C 0 If the CU at position C is not available, is intra-coded, or is outside the current row of the CTU, then position C is used. 1 Otherwise, position C is used in the derivation of the time-domain Merge candidate 0 . Fig.43 The time domain Merge candidate C is shown 0 and C 1 candidate locations. 2.33.3. Merge candidate derivation based on history After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of the previous codec block is stored in a table and used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, and then all HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The most recent HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. For spatial or temporal Merge candidates, redundancy check is applied to HMVP candidates. To reduce the number of redundant checking operations, the following simplifications are introduced: The number of HMPV candidates for Merge list generation is set to (N<=4)?M:(8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list construction process from HMVP is terminated. 2.33.4. Pairwise Average Merge Candidate Derivation The pairwise average candidate is generated by averaging the predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the number represents the Merge index of the Merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are available in one list, they are averaged even if they point to different reference pictures; if only one motion vector is available, it is used directly; if no motion vector is available, this list is kept invalid. When the merge list is not full after adding pairwise average merge candidates, zero MVPs will be inserted at the end until the maximum number of merge candidates is encountered. 2.33.5.Merge estimated region The Merge Estimation Region (MER) allows the Merge candidate list to be independently derived for a CU in the same Merge Estimation Region (MER). For generating the Merge candidate list for the current CU, candidate blocks in the same MER as the current CU are not included. In addition, the update process for the history-based motion vector prediction candidate list is updated only when (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel and (yCb+cbHeight)>>Log2parMrglevel is greater than (yCb>>Log2ParMrgLevel), and (xCb, yCb) is the upper left luminance sample position of the current CU in the picture, and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder end and transmitted by signal in the form of log2_parallel_merge_level_minus2 in the sequence parameter set. 2.34. New Merge Candidates 2.34.1. Non-adjacent Merge Candidate Derivation In VVC, Fig. 27 The five spatial neighboring blocks and one temporal neighboring block shown are used to derive Merge candidates. It is proposed to derive additional Merge candidates from positions non-adjacent to the current block using the same style as in VVC. To achieve this, for each search round i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated using the following formula: Offsetx=-i×gridX, Offsety=-i×gridY Where Offsetx and Offsety represent the offset of the upper left corner of the virtual block relative to the upper left corner of the current block, and gridX and gridY are the width and height of the search grid. Second, the width and height of the virtual block are calculated using the following formula: newWidth=i×2×gridX+currWidth newHeight=i×2×gridY+currHeight. Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block. gridX and gridY are currently set to currWidth and currHeight respectively. Fig.44 Shows the VVC spatial neighboring blocks of the current block. Fig.28 Shows the relationship between the virtual block and the current block. Fig.45 Shows the illustration of the virtual block in the i-th search round. After generating the virtual block, block A i , B i , C i , D i and E i can be regarded as the VVC spatial neighboring blocks of the virtual block, and their positions are obtained using the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i , B i , C i , D i and E i are the spatial neighboring blocks used in the VVCMerge mode. When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighboring blocks are used. The non-adjacent spatial Merge candidates are inserted into the Merge list after the temporal Merge candidates in the order of B 1 ->A 1 ->C 1 ->D 1 ->E 1 . 2.34.2. STMVP Proposes to use three spatial Merge candidates and one temporal Merge candidate to derive the average candidate as the STMVP candidate. STMVP is inserted before the top-left spatial Merge candidate. The STMVP candidate is deduplicated together with all previous Merge candidates in the Merge list. For the spatial candidates, the first three candidates in the current Merge candidate list are used. For the temporal candidate, the same position as in VTM / HEVC at the same position is used. For the spatial candidates, the first, second, and third candidates inserted before STMVP in the current Merge candidate list are denoted as F, S, and T. The temporal candidate with the same position as in VTM / HEVC used in TMVP is denoted as Col. The motion vector of the STMVP candidate in prediction direction X (denoted as mvLX) is derived as follows: 1) If the reference indices of the four Merge candidates are all valid and equal to zero in the prediction direction X (X = 0 or 1), mvLX=(mvLX_F+mvLX_S+mvLX_T+mvLX_Col)>>2 2) If the reference indexes of three of the four Merge candidates are valid and equal to zero in the prediction direction X (X=0 or 1), mvLX=(mvLX_F×3+mvLX_S×3+mvLX_Col×2)>>3 or mvLX=(mvLX_F×3+mvLX_T×3+mvLX_Col×2)>>3 or mvLX=(mvLX_S×3+mvLX_T×3+mvLX_Col×2)>>3. 3) If the reference indexes of two of the four Merge candidates are valid and equal to zero in the prediction direction X (X=0 or 1), mvLX=(mvLX_F+mvLX_Col)>>1 or mvLX=(mvLX_S+mvLX_Col)>>1 or mvLX=(mvLX_T+mvLX_Col)>>1. NOTE: If time domain candidates are not available, STMVP mode is turned off. 2.34.3.Merge List Size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8. 2.35. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter-frame prediction. Use CU level flag as a Merge mode to transmit geometric partitioning mode through signaling. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode and sub-block Merge mode. For each possible CU size w×h=2 m ×2 n , where m,n∈{3…6} does not include 8x64 and 64x8, and the geometric segmentation mode supports a total of 64 segmentations. When using this mode, the CU is divided into two parts by a geometrically positioned straight line ( Fig.29). The position of the dividing line is mathematically derived from the angle and offset parameters of the specific partition. Each part in the geometric partition of the CU is inter-predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e. each part has a motion vector and a reference index. Unidirectional prediction motion constraints are applied to ensure that only two motion compensated predictions are required for each CU, which is the same as traditional bidirectional prediction. The unidirectional prediction motion for each partition is derived using the process described in 2.34.1. Fig.46 An example of GPM partitions grouped at the same angle is shown. If geometric partitioning mode is used for the current CU, a geometric partitioning index (angle and offset) indicating the partitioning mode of the geometric partitioning and two Merge indices (one for each partition) are further transmitted through a signal. The number of maximum GPM candidate sizes is explicitly transmitted through a signal in the SPS, and the syntax binary for the GPM Merge index is specified. After predicting each part of the geometric partitioning, as in 2.34.2, a hybrid process with adaptive weights is used to adjust the sample values along the edges of the geometric partitioning. This is the prediction signal for the entire CU, and the transformation process and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored, as described in 2.34.3. 2.35.1. One-way prediction candidate list construction In 2.32, the unidirectional prediction candidate list is derived directly from the Merge candidate list constructed according to the extended Merge prediction process. Denote n as the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (x equals the parity of n) of the nth extended Merge candidate is used as the nth unidirectional prediction motion vector of the geometric partition mode. These motion vectors are Fig.30 If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector of the geometric partition mode. Fig.47 Uni-directional prediction MV selection for geometric partitioning mode is shown. 2.35.2. Blending Along Geometric Partition Edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edges. The blending weights for each position of the CU are derived based on the distance between the independent position and the partition edge. The distance from position (x, y) to the segmentation edge is derived as: where i and j are the indices of the geometric segmentation angle and offset, which depend on the geometric segmentation index transmitted by the signal. ρ x,j and ρ y,j The sign of depends on the angle index i. The weight of each part of the geometric segmentation is derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) (2-43) w 1 (x,y) = 1 - w 0 (x,y) (2-45) partIdx depends on the angle index i. The weight w 0 An example of is shown in Fig.31 . Fig.48 Shows the use of the geometric segmentation pattern to generate the hybrid weight w 0 . 2.35.3. Motion Field Storage for Geometric Segmentation Mode Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded in the geometric segmentation mode. The type of motion vector stored for each independent position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx <= 0? (1 - partIdx) : partIdx) where motionIdx is equal to d(4x + 2, 4y + 2), which is recalculated from equation (2-18). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, the combined Mv from Mv1 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-predictive motion vector. Otherwise, if Mv1 and Mv2 are from the same list, only the uni-predictive motion Mv2 is stored. 2.36. Multiple Hypothesis Prediction In Multiple Hypothesis Prediction (MHP), up to two additional prediction values are signaled over Inter AMVP mode, Normal Merge mode, Affine Merge and MMVD mode. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 ) n +α n+1 h n+1 The weight factor α is specified according to Table 8 below. Table 8 Weight factors for MHP add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, MHP is applied only when unequal weights in BCW are selected in bi-prediction mode. Additional assumptions can be Merge mode or AMVP mode. In Merge mode, motion information is indicated by Merge index, and Merge candidate list is the same as in geometric partitioning mode. In AMVP mode, reference index, MVP index and MVD are transmitted by signal. 2.37. Non-adjacent airspace candidates Insert non-adjacent spatial merge candidates after TMVP in the regular merge candidate list. The pattern of spatial merge candidates is as follows: Fig.32 The distance between the non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. Fig.49 The spatial neighboring blocks used to derive spatial Merge candidates are shown. 2.38. Template Matching (TM) Template matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between the template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and the blocks in the reference picture (i.e., the same size as the template). Fig.33 As shown, in the [-8, +8] pixel search range, a better MV is searched around the initial motion of the current CU. Template matching is adopted in this paper with two modifications: the search step size is determined based on the AMVR mode, and the TM can be cascaded using the bilateral matching process in the Merge mode. Fig.50Template matching is shown to be performed on the search area around the initial MV. In AMVP mode, the MVP candidate is determined based on the template matching error by selecting the one that achieves the minimum difference between the current block template and the reference block template, and then TM performs MV refinement only on this specific MVP candidate. TM refines the MVP candidate by using an iterative diamond search, starting with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) within the [-8, +8] pixel search range. The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode), and then using half-pixels and quarter-pixels in turn according to the AMVR mode specified in Table 9. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. Table 9 AMVR search style and Merge mode using AMVR In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 9, TM can be performed all the way to 1 / 8 pixel MVD accuracy, or skip those accuracies exceeding half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used according to the merged motion information (i.e., used when AMVR is half-pixel mode). In addition, when TM mode is enabled, template matching can work as an independent process between a block-based bilateral matching (BM) method and a sub-block-based bilateral matching method, or as an additional MV refinement process, depending on whether BM can be enabled according to its enabling condition check. 2.39. Overlapped Block Motion Compensation (OBMC) Overlapped block motion compensation (OBMC) has been previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, OBMC is performed on all motion compensated (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it also applies to luminance and chrominance components. In JEM, an MC block corresponds to a codec block. When a CU is encoded and decoded with sub-CU modes (including Sub-CU Merge, Affine, and FRUC modes), each sub-block of the CU is an MC block. In order to handle CU boundaries in a unified manner, OBMC is performed on all MC block boundaries at the sub-block level, where the sub-block size is set equal to 4×4, as Fig.34 shown. When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of the four connected neighboring sub-blocks, if available and different from the current motion vector, are also used to derive the prediction block of the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal of the current sub-block. The prediction block based on the motion vector of the neighboring sub-block is represented as P N , where N indicates the index of the neighboring upper, lower, left, and right sub-blocks, and the prediction block based on the motion vector of the current sub-block is represented as P C When P N When based on the motion information of the neighboring sub-block containing the same motion information as the current sub-block, OBMC is not N Otherwise, P N Each sample point is added to P C The same point in P N Four rows / columns are added to P C The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . Except for small MC blocks (i.e., when the height or width of the codec block is equal to 4, or when the CU is encoded and decoded using sub-CU mode), only P N Two rows / columns are added to P C In this case, weighting factors {1 / 4, 1 / 8} are used for P N , and the weighting factors {3 / 4, 7 / 8} are used for P C For P generated based on the motion vectors of vertically (horizontally) adjacent sub-blocks N , P N The samples in the same row (column) of are added to P with the same weighting factor. C . Fig.51 An illustration of the sub-blocks to which OBMC is applied is shown. In JEM, for CUs with a size less than or equal to 256 luma samples, a CU-level flag is transmitted by signal to indicate whether OBMC is applied to the current CU. For CUs with a size greater than 256 luma samples or not encoded and decoded using AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account in the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top neighboring blocks and the left neighboring blocks is used to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied. 2.40. Multiple Transformation Selection (MTS) for Kernel Transformations In addition to DCT-II already adopted in HEVC, the multi-transform selection (MTS) scheme is also used for residual coding and decoding of blocks coded and decoded inter-frame and intra-frame. It uses multiple selected transforms in DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 10 shows the basis functions of the selected DST / DCT. Table 10 Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that in HEVC. To keep the intermediate values of the transform coefficients within 16 bits, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled at SPS, a CU level flag is signaled to indicate whether MTS is applied. Here, MTS applies only to luma. MTS signaling will be skipped when one of the following conditions applies: - The position of the last significant coefficient of the luma TB is less than 1 (i.e. DC only). - The last significant coefficient of the brightness TB is within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally signaled to indicate the transform types in the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 11. By eliminating intra-mode and block shape dependencies, a unified transform selection for ISP and implicit MTS is used. If the current block is in ISP mode or if the current block is an intra-block and both intra- and inter-frame explicit MTS are turned on, only DST7 is used for horizontal and vertical transform kernels. In terms of transform matrix accuracy, an 8-bit main transform kernel is used. Therefore, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8) all use 8-bit main transform kernels. Table 11 Transformation and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, high-frequency transform coefficients are cleared to zero for DST-7 and DCT-8 blocks with size (width or height, or width and height) equal to 32. Only coefficients in the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded or decoded using transform skip mode. To avoid syntax coding redundancy, the transform skip flag is not signaled when the CU level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, implicit MTS can still be enabled. 2.41. Sub-Block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, only a sub-part of the residual block is encoded and decoded for the CU. When cu_cbf is equal to 1 for an inter-predicted CU, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded and decoded. For the former case, the inter-MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded by inferring the adaptive transform, while the other part of the residual block is cleared. When SBT is used for an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, such as Fig.35A and Fig.35B As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is similar to the binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is similar to the asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 (in units of luminance samples), 1:3 / 3:1 partitioning along this dimension is not allowed. A CU has a maximum of 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). Different kernel transforms are associated with the two positions, SBT-H and SBT-V. More specifically, the horizontal and vertical transforms for each SBT position are in Fig.35A and Fig.35BFor example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transforms jointly specify the TU slice, cbf, and the horizontal and vertical kernel transform types for the residual block. SBT is not applied to CUs coded in combined inter-intra mode. Fig.52 The SBT location, type and transformation type are shown. 2.42. Adaptive Merge Candidate Reordering Based on Template Matching In order to improve the encoding and decoding efficiency, after constructing the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list according to the template matching cost in ascending order. It is operated in the form of subgroups. The template matching cost is measured by the SAD (sum of absolute differences) between the neighboring samples of the current CU and its corresponding reference samples. If the Merge candidate includes motion information of bidirectional prediction, then Fig.36 As shown, the corresponding reference sample is the average of the corresponding reference sample in reference list 0 and the corresponding reference sample in reference list 1. If the Merge candidate contains motion information at the sub-CU level, then Fig.37 As shown, the corresponding reference sample is composed of neighboring samples of the corresponding reference sub-block. like Fig.38 As shown, the sorting process is operated in the form of subgroups. The first three merge candidates are sorted together. The following three merge candidates are sorted together. The template size (width for the left template or height for the upper template) is 1. The subgroup size is 3. Fig.53 Neighboring sample points used to calculate the SAD are shown. Fig.54 Neighboring samples used to calculate SAD for sub-CU level motion information are shown. Fig.55 The classification process is shown. 2.43. Adaptive Merge Candidate List Assume that the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first subgroup, and the following 3 Merge candidates are taken as the second subgroup (ie, the last subgroup). For encoders, such as Fig.39 As shown, after the Merge candidate list is constructed, some Merge candidates are adaptively reordered in ascending order of Merge candidate cost. More specifically, the template matching costs of the Merge candidates in all subgroups except the last subgroup are calculated; then, the Merge candidates in their own subgroups except the last subgroup are reordered; finally, the final Merge candidate list will be obtained. Fig.56 The reordering process in the encoder is shown. For the decoder, after the Merge candidate list is constructed, such as Fig.40 As shown, some / no Merge candidates are adaptively reordered in ascending order of Merge candidate cost. Fig.40 In the example, the subgroup in which the selected (signaled) Merge candidate is located is called the selected subgroup. Fig.57 The reordering process in the decoder is shown. More specifically, if the selected Merge candidate is located in the last subgroup, the Merge candidate list construction process is terminated after the selected Merge candidate is derived, no reordering is performed and the Merge candidate list is not changed; otherwise, the execution process is as follows: After all Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; the template matching costs of the Merge candidates in the selected subgroup are calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list will be obtained. For both encoder and decoder: The template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template. When deriving the reference samples of the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision. The reference samples (RT) of the templates used for bidirectional prediction are obtained by subtracting the reference samples (RT) of the templates in reference list 0 as follows: 0 ) and the reference sample point (RY) of the template in reference list 1 1 ) is derived by taking the weighted average. RT=((8-w)*RT 0 +w*RT 1 +4)>>3 (2-47) The weight (8-w) of the reference template in reference list 0 and the weight (w) of the reference template in reference list 1 are determined by the BCW index of the Merge candidate. The BCW indexes equal to {0, 1, 2, 3, 4} correspond to w equal to {-2, 3, 4, 5, 10} respectively. If the local illumination compensation (LIC) flag of the Merge candidate is true, the LIC method is used to derive the reference samples of the template. The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT. The template size is 1. This means that the width of the left template and / or the height of the template above is 1. If the codec mode is MMVD, the Merge candidates used to derive the basic Merge candidate are not reordered. If the coding mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered. 2.44. Geometric prediction mode with motion vector difference In the geometric prediction mode with motion vector difference (GMVD), each geometric partition in GPM can decide whether to use GMVD. If GMVD is selected for a geometric region, the MV of the region is calculated as the sum of the MV and MVD of the Merge candidate. All other processing remains the same as in GPM. With GMVD, the MVD is signaled as a direction and distance pair. There are nine candidate distances (1 / 4-pixel, 1 / 2-pixel, 1-pixel, 2-pixel, 3-pixel, 4-pixel, 6-pixel, 8-pixel, 16-pixel), and eight candidate directions (four horizontal / vertical directions and four diagonal directions). In addition, when pic_fpel_mmvd_enabled_flag is equal to 1, the MVD in GMVD is also shifted left by 2 bits like in MMVD. 2.45. Affine MMVD In affine MMVD, affine merge candidates (called base affine merge candidates) are selected, and the MVs of control points are further refined by the signaled MVD information. The MVD information of the MVs of all control points is the same in one prediction direction. When the starting MV is a bidirectional prediction MV in which the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the MV offset added to the list 0MV component of the starting MV and the MV offset of the list 1MV have opposite values; otherwise, when the starting MV is a bidirectional prediction MV in which both lists point to the same side of the current picture (i.e., the POCs of the two references are both greater than the POC of the current picture, or both are less than the POC of the current picture), the MV offset added to the list 0MV component of the starting MV and the MV offset of the list 1MV are the same. 2.46. Adaptive Decoder-side Motion Vector Refinement (ADMVR) In ECM-2.0, if the selected Merge candidate meets the DMVR condition, the multi-pass decoder-side motion vector refinement (DMVR) method is applied in the conventional Merge mode. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the MV in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The adaptive decoder-side motion vector refinement method consists of two new Merge modes that are introduced to refine the MV in only one direction (L0 or L1) of bidirectional prediction for Merge candidates that meet the DMVR conditions. A multi-pass DMVR process is applied to refine the motion vector for the selected Merge candidate, however, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is set to zero. Similar to the conventional Merge mode, the Merge candidates for the proposed Merge mode are derived from spatially adjacent coded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The same Merge candidate list (i.e., ADMVR Merge list) is used by both proposed Merge modes, and the Merge index is encoded and decoded in the conventional Merge mode. 2.47. Convolutional Cross-Component Model (CCCM) for Intra Prediction It is proposed to apply a convolutional cross-component model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit as done by the current CCLM mode. As with CCLM, when chroma subsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. Again, similar to CCLM, there is an option to use a single model or a multi-model variant of CCCM. The multi-model variant uses two models, one derived for samples above the average luminance reference value and the other for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples. 2.47.1. Convolutional filters The proposed convolutional 7-tap filter consists of a 5-tap positive-signed spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of the center (C) luma sample as shown below, which is co-located with the chroma sample to be predicted and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors. Fig.58 The spatial domain portion of the convolution filter is shown. The nonlinear term P is expressed as a power of 2 in the center luma sample C and is scaled to the sample value range of the content: P=(C*C+midVal)>>bitDepth. That is, for 10 bits the content is calculated as: P=(C*C+512)>>10. The bias term B represents a scalar offset between input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content). The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples: predChromaVal=c 0 C+c 1 N+c 2 S+c 3 E+c 4 W+c 5 P+c 6 B. 2.47.2. Calculation of filter coefficients Filter coefficient c i It is calculated by minimizing the MSE between the predicted chrominance samples and the reconstructed chrominance samples in the reference region. Fig.59 A reference region consisting of 6 rows of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension of the region shown in blue is needed to support the "side samples" of the positive shape spatial domain filter and is filled in when in an unavailable region. Fig.59 The reference area (with its filling) used to derive the filter coefficients is shown. MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL decomposed and the final filter coefficients are calculated using back substitution. The process roughly follows the calculation of the ALF filter coefficients in ECM, however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer operations. 2.47.3. Bitstream Signaling The use of this mode is signaled using a PU level flag via CABAC codec. A new CABAC context is included to support this. When entering signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is signaled only if the intra prediction mode is LM_CHROMA_IDX (to enable single mode CCCM) or MMLM_CHROMA_IDX (to enable multi-model CCCM). 2.47.4. Encoder Operation The encoder performs two new RD checks in the chroma prediction mode loop, one for checking single-model CCCM mode and one for checking multi-model CCCM mode. 3. Question In the current design of video codec prediction, the codec selects only one type of prediction method and discards other prediction methods. However, this approach may limit the compression efficiency due to only one type of prediction signal. The benefits of mixing prediction values from multiple prediction methods are not considered. 4. Detailed solution The detailed solutions below should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. In the present disclosure, the term "block" may refer to a codec block (CB), a codec unit (CU), a prediction block (PB), a prediction unit (PU), a transform block (TB), a transform unit (TU), a codec tree block (CTB), a codec tree unit (CTU), or a rectangular area of samples / pixels. In the following discussion, Shift(x,n) is defined as Shift(x,n)=(x+offset0)>>n. In one example, offset0 and / or offset1 are set to (1 <<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0. In another example, offset0=offset1=((1<<n)>>1)-1 or ((1<<(n-1)))-1. In the following discussion, “fusion,” “hybrid,” and “integration” may have the same meaning. Prediction Fusion . 1. It is proposed that a fused prediction value is generated from a set of video coding prediction values in both the encoder and the decoder. Denote the prediction fusion as F(), such as B = F(p 0 , p 1 ,…,pn ), where p n is the nth prediction method. a. In one example, a prediction fusion may refer to combining a set of prediction values (p 0 , p 1 ,…,p n )mix. i. In one example, B = F (p i ), where p i is the predicted value of the ith prediction method, w i is the fusion weight corresponding to the predicted value. 1) In one example, the weight w i Is a fixed predefined value. 2) In one example, the weight w i is a varying value depending on the context of the codec. ii. In one example, B = F(p i ), where p i is the predicted value of the ith prediction method, w i is the fusion weight corresponding to the predicted value, b is the offset and n is the number of right shifts. b. In one example, a prediction fusion may refer to mixing a set of prediction values (p 0 , p 1 ,…,p n ) is a nonlinear method. i. In one example, the nonlinear fusion method can adopt a convolutional neural network. 1) In one example, C(p i )=conv2(w i ,p i )+b i (3) where p i is the predicted value of the ith prediction method, w i is the weight of the convolutional layer, b i is the bias of the convolutional layer. B(p i )=M(C(p 1 ),C(p 2 ),…,C(p i )) (4) Where M is the fusion method of i prediction values. R(p i ) is the rectified linear unit (ReLU), which is the activation function. c. In one example, the total number of predicted values n in (1) is equal to 2. i. In one example, these two predicted values (p 1 ,p 2 ) comes from angular prediction and MIP. F(p i )=w 0 p 0 +w 1 p 1 1) In one example, (w 0 ,w 1 ) are fixed predefined values. 2) In one example, (w 0 ,w 1 ) is a varying value depending on the context of the codec. d. In one example, the combination of video codec prediction values may include two types of intra-frame prediction values, namely, intra-frame angle prediction values with 65 angle modes and MIP prediction values with 32 modes. The total number of combinations is 65X32=2080. A 50-50 weight pair is applied to the combination of prediction values. The codec then uses these weights to mix the two prediction values using (1), where n is equal to 2. The codec calculates the rate-distortion (RD) cost for each combination and selects the one with the lowest cost. After the signaling of the template matching prediction, the flag of the hybrid method is transmitted to the decoder through a signal. If a hybrid method is selected, the index of the applied hybrid method will also be transmitted to the decoder through a signal after the flag of the hybrid method. The decoder uses the same combination of hybrid methods to predict the current CU. e. In one example, the combination of video codec prediction values may include two types of intra prediction values, namely, the intra angle prediction value with the best RD cost in 65 angle modes and the MIP prediction value with the best RD cost in 32 MIP modes. A 50-50 weight pair is applied to the combination of prediction values. The codec then uses these weights to mix the two RD best prediction values using (1), where n is equal to 2. The flag of the hybrid method is transmitted to the decoder through a signal after the signaling of the template matching prediction. If a hybrid method is selected, the index of the applied hybrid method will also be transmitted to the decoder through a signal after the flag of the hybrid method. The decoder uses the same combination of received hybrid methods to predict the current CU. f. In one example, the combination of video codec prediction values may include two types of intra-frame prediction values, namely, intra-frame angle prediction values from the MPM list and MIP prediction values with 32 MIP modes. The total number of combinations is 6X 32=192. A 50-50 weight pair is applied to the combination of prediction values. The codec then uses these weights to mix the two prediction values using (1), where n is equal to 2. The codec calculates the rate-distortion (RD) cost for each combination and selects the one with the lowest cost. The flag of the hybrid method is transmitted to the decoder through a signal after the signaling of the template matching prediction. If a hybrid method is selected, the index of the applied hybrid method will also be transmitted to the decoder through a signal after the flag of the hybrid method. The decoder uses the same combination of hybrid methods to predict the current CU. g. In one example, the combination of video codec prediction values may include two types of intra prediction values, namely, intra angle prediction values from the MPM list using MRL and MIP prediction values with 32 MIP modes. Specifically, the intra angle mode index is derived from the MPM list, and the reference row index is derived from the MRL candidate. The total number of combinations is 6X 32=192. A 50-50 weight pair is applied to the combination of prediction values. The codec then uses these weights to mix the two prediction values using (1), where n is equal to 2. The codec calculates the rate-distortion (RD) cost for each combination and selects the one with the lowest cost. The flag of the hybrid method is transmitted to the decoder through a signal after the signaling of the template matching prediction. If a hybrid method is selected, the index of the applied hybrid method will also be transmitted to the decoder through a signal after the flag of the hybrid method. The decoder uses the same combination of hybrid methods to predict the current CU. h. In one example, the combination of video codec prediction values may include two types of intra prediction values, namely, intra angle prediction values from the MPM list using MRL and MIP prediction values with 32 MIP modes. Specifically, the intra angle mode index is derived from the MPM list, and the reference line index is derived from the MRL candidate. The total number of combinations is 6X 32=192. A 50-50 weight pair is applied to the combination of prediction values. The codec then uses these weights to mix the two prediction values using (1), where n is equal to 2. The codec calculates the sum of absolute transform differences (SATD) cost for each combination and selects the one with the lowest cost. The selected combination is then placed in the RD search pipeline. The flag of the hybrid method is transmitted to the decoder through a signal after the signaling of the template matching prediction. If a hybrid method is selected, the index of the applied hybrid method will also be transmitted to the decoder through a signal after the flag of the hybrid method. The decoder uses the same combination of hybrid methods to predict the current CU. i. In one example, the combination of video codec prediction values may include two types of intra prediction values, namely, intra angle prediction values with fixed angle modes and MIP prediction values with fixed modes among 32 modes. A 50-50 weight pair is applied to the combination of prediction values. The codec then uses these weights to mix the two fixed prediction values using (1), where n is equal to 2. The flag of the mixing method is signaled to the decoder after the signaling of the template matching prediction. The decoder uses the same predefined combination of mixing methods to predict the current CU. j. The prediction value can be generated by any intra-frame prediction mode, such as planar, DC, other angular modes, MIP mode, DIMD, TIMD, MRL, ISP. i. The intra prediction mode used to generate the prediction value is transmitted by signal. ii. The intra prediction mode used to generate the prediction value can be derived at the decoder. iii. In one example, at least one of the prediction values is derived using MIP, and at least one of the prediction values is derived using an intra prediction method other than MIP (non-MIP). 1) In one example, the intra prediction method other than MIP may refer to a conventional intra prediction method and / or MRL and / or ISP and / or DIMD and / or TIMD. 2) In one example, the derivation of a prediction value using MIP may be predefined or derived or signaled in the bitstream. a) In one example, the MIP mode and / or the MIP transposition flag may be predefined or derived or signaled in the bitstream. 3) In one example, intra prediction methods other than MIP may be predefined or signaled in the bitstream. a) In one example, only conventional intra prediction methods or MRL or ISP, or DIMD or TIMD may be used. b) In one example, which intra prediction method is used in addition to MIP and how the intra prediction method is used can be signaled in the bitstream. c) In one example, one or more predefined intra prediction modes may be used to derive at least one prediction value using a non-MIP approach. i. In one example, one or more intra prediction modes are not allowed to be used. 1. In one example, planes and / or DC will not be allowed. d) In one example, one or more predefined indices of the MRL reference lines may be used to derive at least one prediction value in the angle prediction method. i. In one example, one or more intra prediction modes are not allowed to be used. 1. In one example, rows 3, 4, and 5 The row is not allowed to be used. k. In one example, the weight value for a sample may depend on the sample position in the block. Whether and how to apply fusion of predictions. 2. In one example, whether and / or how to apply prediction fusion depends on codec information. a. In one example, the codec information may refer to picture / slice type, temporal layer, QP, color format, color component, codec mode, etc. i. In one example, prediction fusion is applied to I-slices only. 3. In one example, whether and / or how to apply prediction fusion may be signaled using at least one syntax element (SE). a. In one example, the syntax elements may be signaled in SPS / PPS / APS / picture header / slice header / CTU / CU, etc. b. In one example, a first syntax element may be signaled to indicate whether prediction fusion is applied. c. In one example, a second syntax element may be signaled to indicate which mode index of the selected prediction method is applied or not applied. i. In one example, the second syntax element may be signaled only when the first syntax element indicates that prediction fusion is applied. d. In one example, a single SE may be signaled that indicates whether prediction fusion is used and which mode index of the selected method is used. e. In one example, SE may be signaled to indicate or derive weight values. f. SE(s) may be coded using fixed length codec, EG codec, truncated (unary) codec, etc. g. SE(s) may be coded or decoded using at least one context in arithmetic coding or decoding. h. SE(s) can be bypassed for encoding and decoding. i. SE(s) may be signaled only when prediction fusion is allowed to be used. j. SE(s) may be signaled in SEI or VUI messages. Determination of prediction fusion 4. In one example, one or more codec tools may be conditionally used during the determination of the use of predictive fusion. a. In one example, a codec tool may refer to a segmentation method. i. In one example, BT and / or TT may not be used. ii. In one example, a QT with a specific size (S) may be used. 1) In one example, S=128, or S=64, or S=32, or S=16, or S=8. b. In one example, a codec may refer to a color component. i. In one example, the chroma components may not be used. c. In one example, a codec tool may refer to a transform tool. i. In one example, MTS and / or LFNST may not be used. d. In one example, codec tools may refer to intra prediction and / or inter prediction methods. i. In one example, MIP / MRL / ISP / DIMD / TIMD / CCLM / MMLM / CCCM / Chroma DIMD may not be used. ii. In one example, all other intra prediction methods except regular intra prediction may not be used. General 5. Whether and / or how to apply the method disclosed above can be transmitted by signal at the sequence level / picture group level / picture level / slice level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / In the stripe header / slice group header. 6. Whether and / or how to apply the above disclosed methods may depend on the coded information, such as block size, color format, single / dual tree partitioning, color component, slice / picture type. 7. The proposed method disclosed in this document can be used in other codec tools that require prediction fusion.
[0122] More details of embodiments of the present disclosure related to the fusion method in codec prediction will be described below. The embodiments of the present disclosure should be considered as examples to explain general concepts and should not be interpreted in a narrow sense. In addition, these embodiments can be applied individually or in combination in any way.
[0123] As used herein, the term "block" may refer to a color component, a sub-picture, a picture, a slice, a codec tree unit (CTU), a CTU row, a CTU group, a codec unit (CU), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec block (CB), a prediction block (PB), a transform block (TB), a sub-block of a video block, a sub-region within a video block, a video processing unit including a plurality of samples / pixels, etc. A block may be rectangular or non-rectangular.
[0124] Fig.60 6000 is a flowchart of a method 6000 for video processing according to some embodiments of the present disclosure. The method 6000 may be implemented during a conversion between a current video block of a video and a bitstream of the video. Fig.60 As shown, method 6000 begins at 6002, where multiple predictions for a current video block are obtained. The multiple predictions are determined based on multiple different prediction schemes.
[0125] At 6004, a target prediction for the current video block is generated by fusing the multiple predictions. As used herein, the term "fusing" may refer to mixing or combining more than one prediction to obtain a single prediction, and "fusing" may be used interchangeably with the terms "mixing" and "integrating".
[0126] In some embodiments, a target prediction may be generated based on a weighted sum of multiple predictions. In some further embodiments, multiple predictions may be fused based on a nonlinear fusion scheme. As an example and not limitation, a convolutional neural network may be used in a nonlinear fusion scheme. This will be described in detail below.
[0127] At 6006, a conversion is performed based on the target prediction. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream. It should be understood that the above illustrations and / or examples are described for descriptive purposes only. The scope of the present disclosure is not limited in this regard.
[0128] In view of the above, multiple predictions for the current video block determined based on multiple different prediction schemes are fused to generate a fused prediction. Compared with conventional solutions that only use predictions determined based on a specific prediction scheme, the proposed method can advantageously improve encoding and decoding quality and encoding and decoding efficiency.
[0129] In some embodiments, the target prediction may be generated as follows: Where B represents the target prediction, p i represents the i-th prediction among multiple predictions, w irepresents the weight for the i-th prediction, and n represents the number of predictions in the plurality of predictions minus one. As an example, the weight w i The value of can be predefined. Alternatively, the weight w i The value of may depend on the context of the codec used to encode and decode the current video block.
[0130] In some alternative embodiments, the target prediction may be generated as follows: Where B represents the target prediction, p i represents the i-th prediction among multiple predictions, w i represents the weight for the i-th prediction, n represents the number of predictions in the plurality of predictions minus one, b represents the offset, and k represents the number of right shifts. As an example, the weight w i The value of can be predefined. Alternatively, the weight w i The value of may depend on the context of the codec used to encode and decode the current video block.
[0131] In some further embodiments, the target prediction may be generated as follows: C(p i )=conv2(w i ,p i )+b i ,i=0,1,…,n B=M(C(p 1 ),C(p 2 ),…,C(p n )) Where B represents the target prediction, M( ) represents the fusion scheme, and p i represents the i-th prediction among multiple predictions, C(p i ) represents the intermediate result corresponding to the i-th prediction, conv2( ) represents the two-dimensional convolution, w i represents the weight of the convolutional layer for the i-th prediction, b i is the bias of the convolutional layer, and n represents the number of predictions in the multiple predictions minus one.
[0132] In some embodiments, a rectified linear unit (ReLU) may be used as an activation function for a convolutional neural network. By way of example and not limitation, a rectified linear unit may be as follows: Where R(p i ) represents the rectified linear unit corresponding to the i-th prediction.
[0133] In some embodiments, the number of predictions in the plurality of predictions may be 2. It should be understood that the specific values recited herein are intended to be exemplary and not limiting of the scope of the present disclosure. As an example and not limitation, one of the plurality of predictions may be determined based on an angular prediction scheme, and another of the plurality of predictions may be determined based on a matrix weighted intra prediction (MIP).
[0134] In some embodiments, a first prediction of the plurality of predictions may be determined based on an intra angular prediction scheme having a fixed angular mode. Additionally, a second prediction of the plurality of predictions may be determined based on a fixed MIP mode.
[0135] In some alternative embodiments, a first prediction among the plurality of predictions may be determined from a first set of predictions generated based on a first prediction scheme among a plurality of different prediction schemes. Additionally, a second prediction among the plurality of predictions may be determined from a second set of predictions generated based on a second prediction scheme among a plurality of different prediction schemes. As an example and not limitation, the first prediction scheme may include a matrix-weighted intra prediction (MIP), and the second prediction scheme may include an intra-angle prediction scheme, an intra-angle prediction mode from a most probable mode (MPM) list, an intra-angle prediction mode from an MPM list using multiple reference lines (MRLs), and the like.
[0136] In some embodiments, multiple combinations of predictions may be generated based on the first set of predictions and the second set of predictions. Each of the multiple combinations may include a prediction from the first set of predictions and a prediction from the second set of predictions. A rate-distortion (RD) cost may be determined for each of the multiple combinations. In addition, the prediction in one of the multiple combinations having the lowest RD cost may be determined as the first prediction and the second prediction.
[0137] In some embodiments, an RD cost may be determined for each prediction in the first set of predictions, and one prediction in the first set of predictions having the lowest RD cost may be determined as the first prediction. In addition, an RD cost may be determined for each prediction in the second set of predictions, and one prediction in the second set of predictions having the lowest RD cost may be determined as the second prediction.
[0138] In some embodiments, multiple combinations of predictions can be generated based on a first set of predictions and a second set of predictions. Each combination in the multiple combinations may include a prediction from the first set of predictions and a prediction from the second set of predictions. A sum of absolute transform differences (SATD) cost can be determined for each combination in the multiple combinations. Then, the first prediction and the second prediction can be determined from the multiple combinations based on the SATD cost. In one example, the prediction in one combination with the lowest SATD cost in the multiple combinations can be determined as the first prediction and the second prediction.
[0139] Alternatively, a predicted set of combinations can be selected from a plurality of predicted combinations based on SATD cost. For example, a predicted set of combinations with the lowest N SATD cost can be selected, where N is an integer. Then, the RD cost can be determined for each combination of the set of combinations, and the prediction in one combination of the set of combinations with the lowest RD cost can be determined as the first prediction and the second prediction.
[0140] In some embodiments, the target prediction can be generated based on a weighted sum of the first prediction and the second prediction. In one example, the weight for the first prediction can be the same as the weight for the second prediction, such as a 50-50 weight pair. Alternatively, the weight for the first prediction can be different from the weight for the second prediction.
[0141] In some embodiments, a first indication (such as a flag, etc.) can be included in the bitstream, which indicates whether the target prediction is generated by fusing multiple predictions. Additionally, the first indication can be signaled after the indication associated with the template-matched prediction.
[0142] In some embodiments, a second indication can be included in the bitstream, which indicates how to fuse multiple predictions. Additionally, the second indication can be signaled after the first indication. By way of example and not limitation, the second indication can be an index of the fusion method.
[0143] In some embodiments, the multiple different prediction schemes can include at least one intra prediction mode. As an example, at least one intra prediction mode can include planar mode, DC mode, angular mode, MIP mode, decoder-side intra mode derivation (DIMD) mode, template-based intra mode derivation (TIMD) mode, intra subpartition (ISP) mode. It should be understood that the above examples are described only for the purpose of description. The scope of the present disclosure is not limited in this regard.
[0144] In some embodiments, at least one intra prediction mode can be indicated in the bitstream. Alternatively, at least one intra prediction mode can be determined at the decoder.
[0145] In some embodiments, at least one of the multiple predictions can be determined based on MIP. Additionally, at least one of the multiple predictions can be determined based on an intra prediction scheme different from MIP. By way of example and not limitation, the intra prediction scheme can include intra angular prediction scheme, MRL, ISP, DIMD, TIMD, etc.
[0146] In some embodiments, the determination of at least one prediction based on MIP may be predefined. Alternatively, the determination of at least one prediction based on MIP may be determined at the decoder. In some additional embodiments, the determination of at least one prediction based on MIP may be indicated in the bitstream.
[0147] In some embodiments, at least one of the MIP mode or the MIP transpose flag may be predefined. Alternatively, at least one of the MIP mode or the MIP transpose flag may be determined at the decoder. In some additional embodiments, at least one of the MIP mode or the MIP transpose flag may be indicated in the bitstream.
[0148] In some embodiments, the intra prediction scheme may be indicated in the bitstream or be predefined. Additionally or alternatively, information on how to use the intra prediction scheme may be indicated in the bitstream.
[0149] In some embodiments, at least one intra prediction mode may not be allowed to be used for determining multiple predictions. By way of example and not limitation, at least one intra prediction mode may include a planar mode or a DC mode, etc.
[0150] In some embodiments, multiple predictions may be determined based on at least one predefined MRL reference row. Additionally, one or more MRL reference rows may not be allowed to be used for determining multiple predictions. For example, one or more MRL reference rows may be predefined, such as row 3, row 4, and row 5.
[0151] In some embodiments, multiple predictions may be fused based on weight values for samples in the current video block. In this case, the weight values for the samples may depend on the positions of the samples. That is, prediction fusion may be performed at the sample level.
[0152] In some embodiments, at least one of the following may depend on the codec information of the current video block or the codec information of neighboring video blocks of the current video block: whether to generate a target prediction by fusing multiple predictions, or how to fuse multiple predictions. For example, the codec information may include picture type, slice type, temporal layer, quantization parameter (QP), color format, color component, and / or codec mode. It should be understood that the above examples are only described for the purpose of description. The scope of the present disclosure is not limited in this regard. In some embodiments, prediction fusion is specifically applied to intra-coded slices (I slices). In this case, the slice including the current video block is an I slice.
[0153] Additionally or alternatively, at least one of the following may be indicated by at least one syntax element in the bitstream: whether to generate a target prediction by fusing multiple predictions, or how to fuse multiple predictions. For example, at least one syntax element may be included in a sequence parameter set (SPS), a picture parameter set (PPS), an adaptation parameter set (APS), a picture header, a slice header, a codec tree unit (CTU), a codec unit (CU), etc.
[0154] In some embodiments, the at least one syntax element may include a first syntax element indicating whether to generate the target prediction by fusing the multiple predictions. Additionally or alternatively, the at least one syntax element may include a second syntax element indicating how to fuse the multiple predictions.
[0155] In some embodiments, if the first syntax element indicates that the target prediction is generated by fusing a plurality of predictions, the at least one syntax element may further include a second syntax element indicating how to fuse the plurality of predictions.
[0156] In some alternative embodiments, the at least one syntax element may include a single syntax element indicating whether to generate the target prediction by fusing multiple predictions and how to fusing the multiple predictions.
[0157] In some embodiments, the bitstream may include a syntax element indicating a weight value for fusing the multiple predictions. Alternatively, the bitstream may include a syntax element for determining the weight value.
[0158] In some embodiments, at least one syntax element may be encoded and decoded using one of the following: fixed length codec, exponential Golomb (EG) codec, truncation unary codec, or unary codec. In some embodiments, at least one syntax element may be encoded and decoded using at least one context in arithmetic codec. Alternatively, at least one syntax element may be bypassed. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this respect.
[0159] In some embodiments, if multiple predictions are allowed to be fused to generate a target prediction, at least one of the following is indicated by at least one syntax element in the bitstream: whether to generate the target prediction by fusion of multiple predictions, or how to fusion of multiple predictions.
[0160] In some embodiments, the at least one syntax element may be included in a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) message.
[0161] In some embodiments, during determining information about whether to generate a target prediction by fusing a plurality of predictions, at least one codec tool may be conditionally used.
[0162] In some embodiments, at least one codec tool may include a segmentation scheme. For example, a binary tree (BT) segmentation scheme and / or a ternary tree (TT) segmentation scheme may not be used. Additionally or alternatively, a quadtree segmentation scheme with a predetermined size may be used. As an example and not limitation, the predetermined size may include: 128 pixels × 128 pixels, 64 pixels × 64 pixels, 32 pixels × 32 pixels, 16 pixels × 16 pixels, or 8 pixels × 8 pixels.
[0163] In some embodiments, at least one codec tool may include a color component. For example, a chroma component of the current video block may not be used. In some embodiments, at least one codec tool may include a transform scheme. For example, multiple transform selection (MTS) and / or low frequency non-separable transform (LFNST) may not be used.
[0164] In some embodiments, at least one codec tool may include an intra prediction scheme and / or an inter prediction scheme. For example, at least one of the following may not be used: MIP, MRL, ISP, DIMD, TIMD, cross-component linear model (CCLM), multi-model linear model (MMLM), convolutional cross-component model (CCCM), or chroma decoder side intra mode derivation (DIMD). Additionally or alternatively, only at least one predetermined intra prediction scheme may be allowed to be used. As an example and not limitation, all other intra prediction methods except conventional intra prediction (such as intra angle prediction, etc.) may not be used.
[0165] In some embodiments, whether and / or how to apply a method may be indicated at one of: sequence level, group of pictures level, picture level, slice level, or slice group level.
[0166] In some embodiments, whether and / or how to apply the method may be indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0167] In some embodiments, whether and / or how to apply the method may depend on the coded information of the current video block. As an example and not limitation, the coded information may include block size, color format, single tree partition, dual tree partition, color component, slice type and / or picture type.
[0168] In some embodiments, the method may be applicable to codec tools that require prediction fusion. As an example, the proposed method may be applied to both intra-frame prediction and inter-frame prediction. The scope of the present disclosure is not limited in this respect.
[0169] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable medium stores a bitstream of a video generated by a method performed by a device for video processing. In the method, multiple predictions for a current video block of the video can be obtained. The multiple predictions are determined based on multiple different prediction schemes. By fusing the multiple predictions, a target prediction for the current video block is generated. In addition, a bitstream is generated based on the target prediction.
[0170] According to another implementation of the present disclosure, a method for storing a bitstream of a video is provided. In the method, multiple predictions for a current video block of the video are obtained. The multiple predictions are determined based on multiple different prediction schemes. By fusing the multiple predictions, a target prediction for the current video block is generated. In addition, a bitstream is generated based on the target prediction, and the bitstream is stored in a non-transitory computer-readable recording medium.
[0171] Implementations of the present disclosure may be described according to the following items, features of which may be combined in any reasonable manner.
[0172] Item 1. A method for video processing, comprising: for conversion between a current video block of a video and a bitstream of the video, obtaining multiple predictions for the current video block, the multiple predictions being determined based on multiple different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; and performing the conversion based on the target prediction.
[0173] Item 2. A method according to item 1, wherein the target prediction is generated based on a weighted sum of the multiple predictions.
[0174] Clause 3. The method of clause 2, wherein the target prediction is generated as follows: Where B represents the target prediction, p i represents the i-th prediction among the multiple predictions, w i represents a weight for the i-th prediction, and n represents the number of predictions in the plurality of predictions minus one.
[0175] Clause 4. The method of clause 2, wherein the target prediction is generated as follows: Where B represents the target prediction, p i represents the i-th prediction among the multiple predictions, w i represents a weight for the i-th prediction, n represents the number of predictions in the plurality of predictions minus one, b represents an offset, and k represents the number of right shifts.
[0176] Item 5. The method according to any one of Items 3 to 4, wherein the weight w i has a predefined value.
[0177] Item 6. The method according to any one of Items 3 to 4, wherein the weight w i has a value that depends on the context of the codec used for encoding and decoding the current video block.
[0178] Item 7. The method according to Item 1, wherein the plurality of predictions are fused based on a non-linear fusion scheme.
[0179] Item 8. The method according to Item 7, wherein a convolutional neural network is used in the non-linear fusion scheme.
[0180] Item 9. The method according to Item 8, wherein the target prediction is generated as follows: C(p i ) = conv2(w i , p i ) + b i , i = 0, 1, …, n B = M(C(p 1 ), C(p 2 ), …, C(p n )) where B represents the target prediction, M( ) represents the fusion scheme, p i represents the i-th prediction among the plurality of predictions, C(p i ) represents the intermediate result corresponding to the i-th prediction, conv2( ) represents two-dimensional convolution, w i represents the weight of the convolutional layer for the i-th prediction, b i is the bias of the convolutional layer, and n represents one less than the number of predictions among the plurality of predictions.
[0181] Item 10. The method according to Item 9, wherein a rectified linear unit (ReLU) is used as the activation function for the convolutional neural network.
[0182] Item 11. The method according to Item 10, wherein the rectified linear unit is as follows:
[0183] where R(p i ) represents the rectified linear unit corresponding to the i-th prediction.
[0184] Clause 12. A method according to any one of Clauses 1 to 11, wherein the number of predictions in the plurality of predictions is two.
[0185] Item 13. The method of Item 12, wherein a prediction in the plurality of predictions is determined based on an angular prediction scheme and another prediction in the plurality of predictions is determined based on a matrix-weighted intra prediction (MIP).
[0186] Item 14. A method according to Item 12, wherein a first prediction among the multiple predictions is determined from a first set of predictions generated based on a first prediction scheme among the multiple different prediction schemes, and a second prediction among the multiple predictions is determined from a second set of predictions generated based on a second prediction scheme among the multiple different prediction schemes.
[0187] Item 15. A method according to Item 14, wherein multiple combinations of predictions are generated based on the first group of predictions and the second group of predictions, each of the multiple combinations includes a prediction from the first group of predictions and a prediction from the second group of predictions, a rate-distortion (RD) cost is determined for each of the multiple combinations, and a prediction in one of the multiple combinations having a lowest RD cost is determined as the first prediction and the second prediction.
[0188] Item 16. A method according to Item 14, wherein an RD cost is determined for each prediction in the first group of predictions, and one prediction in the first group of predictions with the lowest RD cost is determined as the first prediction, and an RD cost is determined for each prediction in the second group of predictions, and one prediction in the second group of predictions with the lowest RD cost is determined as the second prediction.
[0189] Item 17. A method according to Item 14, wherein multiple combinations of predictions are generated based on the first group of predictions and the second group of predictions, each of the multiple combinations includes a prediction from the first group of predictions and a prediction from the second group of predictions, a sum of absolute transform differences (SATD) cost is determined for each of the multiple combinations, and the first prediction and the second prediction are determined from the multiple combinations based on the SATD cost.
[0190] Clause 18. The method of clause 17, wherein a prediction in one of the plurality of combinations having a lowest SATD cost is determined as the first prediction and the second prediction.
[0191] Item 19. A method according to Item 17, wherein a set of combinations of predictions are selected from the multiple combinations of predictions based on the SATD cost, an RD cost is determined for each combination of the set of combinations, and a prediction in a combination in the set of combinations having a lowest RD cost is determined as the first prediction and the second prediction.
[0192] Item 20. A method according to any one of Items 14 to 19, wherein the first prediction scheme comprises matrix-weighted intra prediction (MIP), and the second prediction scheme comprises one of: an intra angle prediction scheme, an intra angle prediction mode from a most probable mode (MPM) list, or an intra angle prediction mode from an MPM list using multiple reference lines (MRL).
[0193] Item 21. The method of Item 12, wherein a first prediction of the plurality of predictions is determined based on an intra-frame angular prediction scheme having a fixed angular mode, and a second prediction of the plurality of predictions is determined based on a fixed MIP mode.
[0194] Item 22. A method according to any one of Items 14 to 21, wherein the target prediction is generated based on a weighted sum of the first prediction and the second prediction, and the weight for the first prediction is the same as the weight for the second prediction.
[0195] Item 23. A method according to any one of items 14 to 22, wherein a first indication is included in the bitstream, the first indication indicating whether the target prediction is generated by fusing the multiple predictions.
[0196] Clause 24. The method of clause 23, wherein the first indication is signaled after an indication associated with a template matching prediction.
[0197] Item 25. A method according to any one of Items 23 to 24, wherein the first indication comprises a flag.
[0198] Item 26. A method according to any one of items 23 to 25, wherein a second indication is included in the bitstream, the second indication indicating how to fuse the multiple predictions.
[0199] Item 27. A method according to Item 26, wherein the second indication is transmitted via a signal after the first indication.
[0200] Item 28. A method according to any one of Items 26 to 27, wherein the second indication comprises an index.
[0201] Item 29. A method according to any one of Items 1 to 28, wherein the plurality of different prediction schemes includes at least one intra-frame prediction mode.
[0202] Item 30. A method according to Item 29, wherein the at least one intra-frame prediction mode includes at least one of the following: planar mode, DC mode, angular mode, MIP mode, decoder-side intra-frame mode derivation (DIMD) mode, template-based intra-frame mode derivation (TIMD) mode, or intra-frame sub-partitioning (ISP) mode.
[0203] Item 31. A method according to any one of Items 29 to 30, wherein the at least one intra-frame prediction mode is indicated in the bitstream.
[0204] Item 32. A method according to any one of Items 29 to 30, wherein the at least one intra-frame prediction mode is determined at a decoder.
[0205] Item 33. A method according to any one of Items 1 to 32, wherein at least one of the plurality of predictions is determined based on a MIP, and at least one of the plurality of predictions is determined based on an intra-frame prediction scheme different from the MIP.
[0206] Item 34. The method according to Item 33, wherein the intra-frame prediction scheme includes one of the following: an intra-frame angle prediction scheme, MRL, ISP, DIMD, or TIMD.
[0207] Item 35. A method according to any one of items 33 to 34, wherein the determination of the at least one prediction based on the MIP is predefined, or the determination of the at least one prediction based on the MIP is determined at the decoder, or the determination of the at least one prediction based on the MIP is indicated in the bitstream.
[0208] Item 36. A method according to any one of items 33 to 35, wherein at least one of the MIP mode or the MIP transposition flag is predefined, or at least one of the MIP mode or the MIP transposition flag is determined at the decoder, or at least one of the MIP mode or the MIP transposition flag is indicated in the bitstream.
[0209] Item 37. A method according to any one of items 33 to 36, wherein the intra-frame prediction scheme is indicated or predefined in the bitstream.
[0210] Item 38. A method according to any one of items 33 to 37, wherein information about how to use the intra-frame prediction scheme is indicated in the bitstream.
[0211] Item 39. A method according to any of Items 33 to 37, wherein at least one intra-frame prediction mode is not allowed to be used to determine the multiple predictions.
[0212] Item 40. The method of Item 39, wherein the at least one intra-prediction mode comprises at least one of a planar mode or a DC mode.
[0213] Item 41. A method according to any one of Items 1 to 40, wherein the plurality of predictions are determined based on at least one predefined MRL reference row.
[0214] Item 42. A method according to any one of Items 1 to 41, wherein one or more MRL reference rows are not allowed to be used in determining the plurality of predictions.
[0215] Item 43. A method according to Item 42, wherein the one or more MRL reference lines are predefined.
[0216] Item 44. A method according to any one of Items 1 to 43, wherein the multiple predictions are fused based on weight values for samples in the current video block, and the weight values for the samples depend on the positions of the samples.
[0217] Item 45. A method according to any one of Items 1 to 44, wherein at least one of the following depends on the codec information of the current video block or the codec information of the neighboring video blocks of the current video block: whether to generate the target prediction by fusing the multiple predictions; or how to fuse the multiple predictions.
[0218] Item 46. A method according to Item 45, wherein the codec information includes at least one of the following: picture type, slice type, temporal layer, quantization parameter (QP), color format, color component, or codec mode.
[0219] Item 47. The method of any one of Items 1 to 46, wherein the slice comprising the current video block is an intra-coded slice (I slice).
[0220] Item 48. A method according to any one of items 1 to 44, wherein at least one of the following is indicated by at least one syntax element in the bitstream: whether to generate the target prediction by fusing the multiple predictions, or how to fusing the multiple predictions.
[0221] Item 49. The method according to Item 48, wherein the at least one syntax element is included in one of the following: sequence parameter set (SPS), picture parameter set (PPS), adaptive parameter set (APS), picture header, slice header, coding tree unit (CTU), or coding unit (CU).
[0222] Item 50. The method according to any one of Items 48 to 49, wherein the at least one syntax element includes a first syntax element indicating whether the target prediction is generated by fusing the plurality of predictions.
[0223] Item 51. The method according to any one of Items 48 to 50, wherein the at least one syntax element includes a second syntax element indicating how to fuse the plurality of predictions.
[0224] Item 52. The method according to Item 50, wherein if the first syntax element indicates that the target prediction is generated by fusing the plurality of predictions, the at least one syntax element further includes a second syntax element indicating how to fuse the plurality of predictions.
[0225] Item 53. The method according to any one of Items 48 to 49, wherein the at least one syntax element includes a single syntax element indicating whether the target prediction is generated by fusing the plurality of predictions and how to fuse the plurality of predictions.
[0226] Item 54. The method according to any one of Items 48 to 53, wherein the bitstream includes a syntax element indicating a weight value for fusing the plurality of predictions or a syntax element for determining the weight value.
[0227] Item 55. The method according to any one of Items 48 to 54, wherein the at least one syntax element is coded using one of the following: fixed-length coding, exponential Golomb (EG) coding, truncated unary coding, or unary coding.
[0228] Item 56. The method according to any one of Items 48 to 55, wherein the at least one syntax element is coded using at least one context in arithmetic coding.
[0229] Item 57. The method according to any one of Items 48 to 55, wherein the at least one syntax element is bypass-coded.
[0230] Item 58. The method according to any one of Items 1 to 47, wherein if the plurality of predictions are allowed to be fused to generate the target prediction, at least one of the following is indicated by at least one syntax element in the bitstream: whether the target prediction is generated by fusing the plurality of predictions, or how to fuse the plurality of predictions.
[0231] Item 59. A method according to any one of items 48 to 58, wherein the at least one syntax element is included in a supplemental enhancement information (SEI) message or a video usability information (VUI) message.
[0232] Item 60. A method according to any one of Items 1 to 59, wherein during the determination of the information regarding whether to generate the target prediction by fusing the multiple predictions, at least one codec tool is conditionally used.
[0233] Item 61. A method according to Item 60, wherein at least one of the encoding and decoding tools includes a segmentation scheme.
[0234] Item 62. The method according to Item 61, wherein at least one of the following is not used: a binary tree (BT) partitioning scheme, or a ternary tree (TT) partitioning scheme.
[0235] Item 63. A method according to any one of Items 60 to 62, wherein a quadtree partitioning scheme having a predetermined size is used.
[0236] Item 64. The method of Item 63, wherein the predetermined size comprises one of: 128 pixels x 128 pixels, 64 pixels x 64 pixels, 32 pixels x 32 pixels, 16 pixels x 16 pixels, or 8 pixels x 8 pixels.
[0237] Item 65. A method according to any one of Items 60 to 64, wherein the at least one codec comprises a color component.
[0238] Item 66. The method of Item 65, wherein chroma components of the current video block are not used.
[0239] Item 67. A method according to any one of Items 60 to 66, wherein the at least one encoding and decoding tool comprises a transformation scheme.
[0240] Item 68. The method of Item 67, wherein at least one of the following is not used: multiple transform selection (MTS), or low frequency non-separable transform (LFNST).
[0241] Item 69. A method according to any one of items 60 to 68, wherein the at least one coding tool includes at least one of the following: an intra-frame prediction scheme, or an inter-frame prediction scheme.
[0242] Item 70. A method according to Item 69, wherein at least one of the following is not used: MIP, MRL, ISP, DIMD, TIMD, cross-component linear model (CCLM), multi-model linear model (MMLM), convolutional cross-component model (CCCM), or chroma decoder side intra-frame mode derivation (DIMD).
[0243] Item 71. A method according to any one of items 60 to 70, wherein only at least one predetermined intra-frame prediction scheme is allowed to be used.
[0244] Item 72. A method according to any one of items 1 to 71, wherein whether and / or how the method is applied is indicated at one of the following: sequence level, picture group level, picture level, slice level, or slice group level.
[0245] Item 73. A method according to any one of Items 1 to 71, wherein whether and / or how the method is applied is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0246] Item 74. A method according to any one of Items 1 to 73, wherein whether and / or how to apply the method depends on encoded information of the current video block.
[0247] Item 75. The method of Item 74, wherein the encoded information comprises at least one of: block size, color format, single tree partitioning, dual tree partitioning, color component, slice type, or picture type.
[0248] Item 76. A method according to any one of items 1 to 75, wherein the method is applicable to a codec tool requiring predictive fusion.
[0249] Item 77. A method according to any one of Items 1 to 76, wherein the converting comprises encoding the current video block into the bitstream.
[0250] Item 78. A method according to any one of Items 1 to 76, wherein the converting comprises decoding the current video block from the bitstream.
[0251] Item 79. An apparatus for video processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1 to 78.
[0252] Item 80. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method according to any one of Items 1 to 78.
[0253] Item 81. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by a device for video processing, wherein the method comprises: obtaining multiple predictions for a current video block of the video, the multiple predictions being determined based on multiple different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; and generating the bitstream based on the target prediction.
[0254] Item 82. A method for storing a bitstream of a video, comprising: obtaining multiple predictions for a current video block of the video, the multiple predictions being determined based on multiple different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; generating the bitstream based on the target prediction; and storing the bitstream in a non-transitory computer-readable recording medium. Example Device
[0255] Fig.61 A block diagram of a computing device 6100 in which various embodiments of the present disclosure may be implemented is shown. The computing device 6100 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0256] It should be understood that Fig.61 The computing device 6100 shown in the figure is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0257] like Fig.61 As shown, computing device 6100 includes a general computing device 6100. Computing device 6100 may include at least one or more processors or processing units 6110, memory 6120, storage unit 6130, one or more communication units 6140, one or more input devices 6150, and one or more output devices 6160.
[0258] In some embodiments, the computing device 6100 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 6100 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0259] The processing unit 6110 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 6120. In a multi-processor system, multiple processing units execute computer executable instructions in parallel to improve the parallel processing capability of the computing device 6100. The processing unit 6110 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0260] The computing device 6100 typically includes various computer storage media. Such media can be any media accessible by the computing device 6100, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 6120 can be a volatile memory (e.g., a register, a cache, a random access memory (RAM)), a non-volatile memory (such as a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM) or flash memory) or any combination thereof. The storage unit 6130 can be any removable or non-removable medium, and can include machine-readable media, such as a memory, a flash drive, a disk or other media that can be used to store information and / or data and can be accessed in the computing device 6100.
[0261] The computing device 6100 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Fig.61 Although not shown in the figure, a disk drive for reading from and / or writing to a removable nonvolatile disk and an optical drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to the bus (not shown) via one or more data medium interfaces.
[0262] The communication unit 6140 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 6100 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Therefore, the computing device 6100 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.
[0263] The input device 6150 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. The output device 6160 may be one or more of various output devices, such as a display, a speaker, a printer, and the like. With the aid of the communication unit 6140, the computing device 6100 may also communicate with one or more external devices (not shown), such as storage devices and display devices, and the computing device 6100 may also communicate with one or more devices that enable a user to interact with the computing device 6100, or, if necessary, the computing device 6100 may also communicate with any device (e.g., a network card, a modem, etc.) that enables the computing device 6100 to communicate with one or more other computing devices. Such communication may be performed via an input / output (I / O) interface (not shown).
[0264] In some embodiments, some or all components of the computing device 6100 may also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access and storage services, which will not require the end user to know the physical location or configuration of the system or hardware that provides these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using a suitable protocol. For example, a cloud computing provider provides an application via a wide area network, which can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on a server at a remote location. The computing resources in a cloud computing environment may be merged or distributed at the location of a remote data center. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single access point to the user. Therefore, a cloud computing architecture may be used to provide components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server, or may be installed on a client device directly or otherwise.
[0265] In an embodiment of the present disclosure, the computing device 6100 may be used to implement video encoding / decoding. The memory 6120 may include one or more video encoding / decoding modules 6125 having one or more program instructions. These modules are accessible and executable by the processing unit 6110 to perform the functions of the various embodiments described herein.
[0266] In an example embodiment performing video encoding, the input device 6150 may receive video data as input to be encoded 6170. The video data may be processed, for example, by the video codec module 6125 to generate an encoded bitstream. The encoded bitstream may be provided as output 6180 via the output device 6160.
[0267] In an example embodiment performing video decoding, the input device 6150 may receive an encoded bitstream as input 6170. The encoded bitstream may be processed, for example, by the video codec module 6125 to generate decoded video data. The decoded video data may be provided as output 6180 via the output device 6160.
[0268] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be appreciated by those skilled in the art that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These modifications are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, include: For conversion between a current video block of a video and a bitstream of the video, obtain a plurality of predictions for the current video block, the plurality of predictions being determined based on a plurality of different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; as well as The transforming is performed based on the target prediction.
2. The method of claim 1, wherein the target prediction is generated based on a weighted sum of the multiple predictions.
3. The method of claim 2, wherein the target prediction is generated as follows: Where B represents the target prediction, p i represents the i-th prediction among the multiple predictions, w i represents a weight for the i-th prediction, and n represents the number of predictions in the plurality of predictions minus one.
4. The method of claim 2, wherein the target prediction is generated as follows: Where B represents the target prediction, p i represents the i-th prediction among the multiple predictions, w i represents a weight for the i-th prediction, n represents the number of predictions in the plurality of predictions minus one, b represents an offset, and k represents the number of right shifts.
5. The method according to any one of claims 3 to 4, wherein the weight w i The values are predefined.
6. The method according to any one of claims 3 to 4, wherein the weight w i The value of depends on the context of the codec used to encode and decode the current video block. The method of claim 1 , wherein the plurality of predictions are fused based on a non-linear fusion scheme.
8. The method of claim 7, wherein a convolutional neural network is used in the nonlinear fusion scheme.
9. The method of claim 8, wherein the target prediction is generated as follows: C(p i )=conv2(w i ,p i )+b i ,i=0,1,…,n B=M(C(p 1 ),C(p 2 ),…,C(pn)) Where B represents the target prediction, M() represents the fusion scheme, and p i represents the i-th prediction among the multiple predictions, C(p i ) represents the intermediate result corresponding to the i-th prediction, conv2() represents a two-dimensional convolution, and w i represents the weight of the convolutional layer for the i-th prediction, b i is the bias of the convolutional layer, and n represents the number of predictions in the plurality of predictions minus one.
10. The method of claim 9, wherein a rectified linear unit (ReLU) is used as an activation function for the convolutional neural network.
11. The method according to claim 10, wherein the rectified linear unit is as follows: Where R(p i ) represents the revised linear unit corresponding to the i-th prediction.
12. The method according to any one of claims 1 to 11, wherein the number of predictions in the plurality of predictions is 2.
13. The method of claim 12, wherein a prediction of the plurality of predictions is determined based on an angular prediction scheme and another prediction of the plurality of predictions is determined based on a matrix-weighted intra prediction (MIP).
14. The method of claim 12, wherein a first prediction among the plurality of predictions is determined from a first set of predictions generated based on a first prediction scheme among the plurality of different prediction schemes, and a second prediction among the plurality of predictions is determined from a second set of predictions generated based on a second prediction scheme among the plurality of different prediction schemes.
15. The method of claim 14, wherein multiple combinations of predictions are generated based on the first set of predictions and the second set of predictions, each of the multiple combinations includes a prediction from the first set of predictions and a prediction from the second set of predictions, a rate-distortion (RD) cost is determined for each of the multiple combinations, and a prediction in one of the multiple combinations having a lowest RD cost is determined as the first prediction and the second prediction.
16. The method of claim 14, wherein an RD cost is determined for each prediction in the first set of predictions, one prediction in the first set of predictions having a lowest RD cost is determined as the first prediction, and an RD cost is determined for each prediction in the second set of predictions, one prediction in the second set of predictions having a lowest RD cost is determined as the second prediction.
17. A method according to claim 14, wherein multiple combinations of predictions are generated based on the first group of predictions and the second group of predictions, each of the multiple combinations includes a prediction from the first group of predictions and a prediction from the second group of predictions, a sum of absolute transform differences (SATD) cost is determined for each of the multiple combinations, and the first prediction and the second prediction are determined from the multiple combinations based on the SATD cost. 18 . The method of claim 17 , wherein a prediction in one of the plurality of combinations having a lowest SATD cost is determined as the first prediction and the second prediction.
19. The method of claim 17, wherein a set of combinations of predictions are selected from the plurality of combinations of predictions based on the SATD cost, an RD cost is determined for each combination of the set of combinations, and a prediction in one of the set of combinations having a lowest RD cost is determined as the first prediction and the second prediction.
20. The method of any one of claims 14 to 19, wherein the first prediction scheme comprises matrix-weighted intra prediction (MIP), and the second prediction scheme comprises one of: Intra-frame angle prediction scheme, An intra angular prediction mode from the Most Probable Mode (MPM) list, or Intra angular prediction modes from MPM lists using multiple reference lines (MRL).
21. The method of claim 12, wherein a first prediction of the plurality of predictions is determined based on an intra angular prediction scheme having a fixed angular mode, and a second prediction of the plurality of predictions is determined based on a fixed MIP mode.
22. The method of any one of claims 14 to 21, wherein the target prediction is generated based on a weighted sum of the first prediction and the second prediction, and the weight for the first prediction is the same as the weight for the second prediction.
23. The method according to any one of claims 14 to 22, wherein a first indication is included in the bitstream, the first indication indicating whether the target prediction is generated by fusing the plurality of predictions.
24. The method of claim 23, wherein the first indication is signaled after an indication associated with a template matching prediction.
25. A method according to any one of claims 23 to 24, wherein the first indication comprises a flag.
26. The method according to any one of claims 23 to 25, wherein a second indication is included in the bitstream, the second indication indicating how to fuse the multiple predictions.
27. The method of claim 26, wherein the second indication is signaled after the first indication.
28. The method of any one of claims 26 to 27, wherein the second indication comprises an index.
29. The method of any one of claims 1 to 28, wherein the plurality of different prediction schemes comprises at least one intra prediction mode.
30. The method of claim 29, wherein the at least one intra prediction mode comprises at least one of: Plane mode, DC mode, Angle mode, MIP mode, Decoder-side intra mode derivation (DIMD) mode, Template-based Intra Mode Derivation (TIMD) mode, or Intra-frame sub-segmentation (ISP) mode.
31. The method according to any one of claims 29 to 30, wherein the at least one intra prediction mode is indicated in the bitstream.
32. The method of any one of claims 29 to 30, wherein the at least one intra prediction mode is determined at a decoder.
33. The method of any one of claims 1 to 32, wherein at least one of the plurality of predictions is determined based on a MIP, and at least one of the plurality of predictions is determined based on an intra prediction scheme different from the MIP.
34. The method of claim 33, wherein the intra prediction scheme comprises one of: Intra-frame angle prediction scheme, MRL, ISP, DIMD, or TIMD.
35. The method according to any one of claims 33 to 34, wherein the determination of the at least one prediction based on the MIP is predefined, or The determination of the at least one prediction based on the MIP is determined at a decoder, or The determination of the at least one prediction based on the MIP is indicated in the bitstream.
36. The method according to any one of claims 33 to 35, wherein at least one of the MIP mode or the MIP transposition flag is predefined, or At least one of the MIP mode or the MIP transposition flag is determined at a decoder, or At least one of the MIP mode or the MIP transposition flag is indicated in the bitstream.
37. The method according to any one of claims 33 to 36, wherein the intra prediction scheme is indicated or predefined in the bitstream.
38. The method according to any one of claims 33 to 37, wherein information on how to use the intra prediction scheme is indicated in the bitstream.
39. The method of any one of claims 33 to 37, wherein at least one intra prediction mode is not allowed for determining the plurality of predictions.
40. The method of claim 39, wherein the at least one intra-prediction mode comprises at least one of a planar mode or a DC mode.
41. A method according to any one of claims 1 to 40, wherein the plurality of predictions are determined based on at least one predefined MRL reference row.
42. A method according to any one of claims 1 to 41, wherein one or more MRL reference rows are not permitted to be used in determining the plurality of predictions.
43. The method of claim 42, wherein the one or more MRL reference lines are predefined.
44. The method according to any one of claims 1 to 43, wherein the multiple predictions are fused based on weight values for samples in the current video block, and the weight values for samples depend on the positions of the samples.
45. The method according to any one of claims 1 to 44, wherein at least one of the following depends on codec information of the current video block or codec information of a neighboring video block of the current video block: whether to generate the target prediction by fusing the multiple predictions; or How to fuse the multiple predictions.
46. The method according to claim 45, wherein the codec information comprises at least one of the following: Image type, Strip type, Time domain layer, Quantization Parameter (QP), Color format, Color components, or Codec mode.
47. The method of any one of claims 1 to 46, wherein the slice comprising the current video block is an intra-coded slice (I slice).
48. The method according to any one of claims 1 to 44, wherein at least one of the following is indicated by at least one syntax element in the bitstream: whether to generate the target prediction by fusing the multiple predictions, or How to fuse the multiple predictions.
49. The method of claim 48, wherein the at least one syntax element is included in one of: Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture header, Strip head, Codec Tree Unit (CTU), or Codec Unit (CU).
50. The method of any one of claims 48 to 49, wherein the at least one syntax element comprises a first syntax element indicating whether the target prediction is generated by fusing the multiple predictions.
51. The method of any one of claims 48 to 50, wherein the at least one syntax element comprises a second syntax element indicating how to fuse the multiple predictions.
52. The method of claim 50, wherein if the first syntax element indicates that the target prediction is generated by fusing the multiple predictions, the at least one syntax element further includes a second syntax element indicating how the multiple predictions are fused.
53. The method of any one of claims 48 to 49, wherein the at least one syntax element comprises a single syntax element indicating whether the target prediction is generated by fusing the multiple predictions and how the multiple predictions are fused.
54. The method according to any one of claims 48 to 53, wherein the bitstream comprises a syntax element indicating a weight value for fusing the plurality of predictions or a syntax element for determining the weight value.
55. The method of any one of claims 48 to 54, wherein the at least one syntax element is encoded using one of: Fixed length codec, Exponential Columbus (EG) codec, Truncated unary codec, or Unary codec.
56. The method according to any one of claims 48 to 55, wherein the at least one syntax element is encoded in arithmetic coding using at least one context.
57. The method according to any one of claims 48 to 55, wherein the at least one syntax element is bypass coded.
58. The method of any one of claims 1 to 47, wherein if the multiple predictions are allowed to be fused to generate the target prediction, at least one of the following is indicated by at least one syntax element in the bitstream: whether to generate the target prediction by fusing the multiple predictions, or How to fuse the multiple predictions.
59. The method of any one of claims 48 to 58, wherein the at least one syntax element is included in a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) message.
60. The method according to any one of claims 1 to 59, wherein during determining the information about whether to generate the target prediction by fusing the multiple predictions, at least one codec tool is conditionally used.
61. The method of claim 60, wherein the at least one codec tool comprises a segmentation scheme.
62. The method of claim 61, wherein at least one of the following is not used: Binary Tree (BT) partitioning scheme, or Ternary Tree (TT) partitioning scheme.
63. A method according to any one of claims 60 to 62, wherein a quadtree partitioning scheme with a predetermined size is used.
64. The method of claim 63, wherein the predetermined size comprises one of: 128 pixels × 128 pixels, 64 pixels × 64 pixels, 32 pixels × 32 pixels, 16 pixels x 16 pixels, or 8 pixels x 8 pixels.
65. The method of any one of claims 60 to 64, wherein the at least one codec comprises a color component.
66. The method of claim 65, wherein chroma components of the current video block are not used.
67. A method according to any one of claims 60 to 66, wherein the at least one codec tool comprises a transform scheme.
68. The method of claim 67, wherein at least one of the following is not used: Multiple Transformation Select (MTS), or Low frequency non-separable transform (LFNST).
69. The method of any one of claims 60 to 68, wherein the at least one codec tool comprises at least one of: Intra prediction scheme, or Inter-frame prediction scheme.
70. The method of claim 69, wherein at least one of the following is not used: MIP, MRL, ISP, DIMD, TIMD, Cross-Component Linear Model (CCLM), Multiple Model Linear Model (MMLM), Convolutional Cross-Component Model (CCCM), or Chroma decoder side intra mode derivation (DIMD).
71. The method according to any one of claims 60 to 70, wherein only at least one predetermined intra prediction scheme is allowed to be used.
72. A method according to any one of claims 1 to 71, wherein whether and / or how to apply the method is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.
73. A method according to any one of claims 1 to 71, wherein whether and / or how to apply the method is indicated in one of: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
74. The method according to any one of claims 1 to 73, wherein whether and / or how to apply the method depends on coded information of the current video block.
75. The method of claim 74, wherein the encoded information comprises at least one of: Block size, Color format, Single tree partitioning, Dual tree partitioning, Color component, Strip type, or Image type.
76. The method of any one of claims 1 to 75, wherein the method is applicable to a codec requiring prediction fusion.
77. The method of any one of claims 1 to 76, wherein the converting comprises encoding the current video block into the bitstream.
78. The method of any one of claims 1 to 76, wherein the converting comprises decoding the current video block from the bitstream.
79. An apparatus for video processing, comprising a processor and a non-volatile memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of claims 1 to 78.
80. A non-transitory computer readable storage medium storing instructions, the instructions causing a processor to perform the method according to any one of claims 1 to 78.
81. A non-transitory computer-readable recording medium storing a bit stream of a video generated by a method performed by an apparatus for video processing, wherein the method include: Obtaining a plurality of predictions for a current video block of the video, the plurality of predictions being determined based on a plurality of different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; as well as The bitstream is generated based on the target prediction.
82. A method for storing a bit stream of a video, include: Obtaining a plurality of predictions for a current video block of the video, the plurality of predictions being determined based on a plurality of different prediction schemes; generating a target prediction for the current video block by fusing the multiple predictions; generating the bitstream based on the target prediction; as well as The bit stream is stored in a non-transitory computer-readable recording medium.