Method and device for video processing and medium

By acquiring multiple sets of weights and determining target predictions, the efficiency problem of existing video encoding and decoding technologies when processing different signals and textures is solved, achieving more efficient encoding and decoding effects.

CN120345246APending Publication Date: 2025-07-18DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085361.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-15
Filing Date
2023-12-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially when processing video blocks or frames with different signals and texture distributions, traditional methods are difficult to achieve the best results.

Method used

By acquiring multiple sets of weights and determining the target prediction based on the multiple candidate predictions of the first color component of the current video block, the conversion is performed to improve the codec quality.

Benefits of technology

Improve the quality of video encoding and decoding, especially for blocks or frames with different signals and texture distributions, and improve the encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120345246A_ABST
    Figure CN120345246A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises the following steps: acquiring a plurality of groups of weights for conversion between a current video block of a video and a bit stream of the video; determining a target prediction for the first color component based on the plurality of sets of weights and the plurality of candidate predictions for the first color component of the current video block; and performing a conversion based on the target prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to component fusion for video encoding and decoding. Background Art

[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: obtaining multiple sets of weights for the conversion between the current video block of a video and the bitstream of the video; determining a target prediction for a first color component based on the multiple sets of weights and multiple candidate predictions for the first color component of the current video block; and performing the conversion based on the target prediction.

[0005] According to the method of the first aspect of the present disclosure, the prediction fusion for color components is performed based on multiple sets of weights. Compared with traditional solutions that only use a single set of predetermined weights, the proposed method can advantageously improve the encoding and decoding quality, especially for blocks or frames with different signal and texture distributions.

[0006] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a video processing device for a video. The method includes: obtaining multiple sets of weights; determining a target prediction for a first color component based on the multiple sets of weights and multiple candidate predictions for the first color component of the current video block of the video; and generating a bitstream based on the target prediction.

[0009] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining multiple sets of weights; determining a target prediction for a first color component based on the multiple sets of weights and multiple candidate predictions for the current video block of the video for the first color component; generating a bitstream based on the target prediction; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The present disclosure is provided to introduce a selection of concepts further described below in the detailed description in a simplified form. The present disclosure is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The above and other objects, features, and advantages of the example embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the example embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram showing an example video codec system according to some embodiments of the present disclosure is shown;

[0013] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure is shown;

[0014] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure is shown;

[0015] Figure 4 The nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture are shown;

[0016] Figure 5 An example of an encoder block diagram is shown;

[0017] Figure 6 Sixty-seven intra prediction modes are shown;

[0018] Figure 7 Reference samples for wide-angle intra prediction are shown;

[0019] Figure 8 The problem of discontinuity in the case of directions exceeding 45° is shown;

[0020] Figure 9 The positions of the samples for the derivation of α and β are shown;

[0021] Figure 10A Four Sobel-based gradient patterns for a gradient linear model (GLM) are shown;

[0022] Figure 10B An example of classifying neighboring samples into two groups is shown;

[0023] Figure 11A It is a schematic diagram showing the definition of samples used by the PDPC applied to the upper-right diagonal pattern;

[0024] Figure 11B It is a schematic diagram showing the definition of samples used by the PDPC applied to the lower-left diagonal pattern;

[0025] Figure 11C It is a schematic diagram showing the definition of samples used by the PDPC applied to the adjacent upper-right diagonal pattern;

[0026] Figure 11D It is a schematic diagram showing the definition of samples used by the PDPC applied to the adjacent lower-left diagonal pattern;

[0027] Figure 12 A gradient method for non-vertical / non-horizontal patterns is shown;

[0028] Figure 13 The nScale values with respect to nTbH and the pattern number are shown; for all cases where nScale < 0, the gradient method is used;

[0029] Figure 14 A flowchart is shown: the current PDPC (left) and the proposed PDPC (right);

[0030] Figure 15 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown;

[0031] Figure 16 An example of the proposed intra reference mapping is shown;

[0032] Figure 17 An example of four reference lines adjacent to the prediction block is shown;

[0033] Figure 18A It is a schematic diagram showing an example of the sub-division for 4×8 and 8×4 CUs;

[0034] Figure 18B It is a schematic diagram showing an example of the sub-division for CUs other than 4×8, 8×4, and 4×4;

[0035] Figure 19 A matrix weighted intra prediction process is shown;

[0036] Figure 20Shows the target sample points, template sample points, and reference sample points of the template used in DIMD;

[0037] Figure 21 Shows the proposed intra-block decoding process;

[0038] Figure 22 Shows the HoG calculation from a template with a width of 3 pixels;

[0039] Figure 23 Shows the prediction fusion through the weighted average of two HoG patterns and planes;

[0040] Figure 24 Shows the neighboring reconstruction sample points for the DIMD chroma mode;

[0041] Figure 25 Shows the MMVD search points;

[0042] Figure 26 Shows the illustration for the symmetric MVD mode;

[0043] Figure 27 Shows the extended CU region used in BDOF;

[0044] Figure 28 Shows the affine motion model based on control points;

[0045] Figure 29 Shows the affine MVF for each sub-block;

[0046] Figure 30 Shows the position of the inherited affine motion prediction value;

[0047] Figure 31 Shows the control point motion vector inheritance;

[0048] Figure 32 Shows the position of the candidate positions for the constructed affine Merge mode;

[0049] Figure 33 Shows the illustration of the motion vector usage for the proposed combined method;

[0050] Figure 34 Shows the sub-block MV VSB and pixel Δv(i,j);

[0051] Figure 35A and Figure 35B Shows the SbTMVP process in VVC;

[0052] Figure 36 Shows the local illumination compensation;

[0053] Figure 37Shows that subsampling for the short side is not performed;

[0054] Figure 38 Shows decoding - side motion vector refinement;

[0055] Figure 39 Shows the diamond - shaped region in the search area;

[0056] Figure 40 Shows the positions of spatial - domain Merge candidates;

[0057] Figure 41 Shows the candidate pairs considered for redundancy checking of spatial - domain Merge candidates;

[0058] Figure 42 Shows an illustration of motion vector scaling for temporal - domain Merge candidates;

[0059] Figure 43 Shows the candidate positions, C0 and C1, for temporal - domain Merge candidates;

[0060] Figure 44 Shows the VVC spatial - domain neighboring blocks of the current block;

[0061] Figure 45 Shows an illustration of the virtual block in the i - th round of search;

[0062] Figure 46 Shows an example of GPM partitioning grouped by the same angle;

[0063] Figure 47 Shows the unidirectional prediction MV selection for geometric partitioning mode;

[0064] Figure 48 Shows the example generation of the hybrid weight w0 using the geometric partitioning mode;

[0065] Figure 49 Shows the spatial - domain neighboring blocks used to derive spatial - domain Merge candidates;

[0066] Figure 50 Shows the template matching performed on the search area around the initial MV;

[0067] Figure 51 Shows an illustration of the sub - blocks for OBMC application;

[0068] Figure 52 Shows the SBT positions, types, and transform types;

[0069] Figure 53 Shows the neighboring samples used to calculate SAD;

[0070] Figure 54Shows neighboring samples used to calculate SAD for sub-CU level motion information;

[0071] Figure 55 Shows the sorting process;

[0072] Figure 56 Shows the reordering process in the encoder;

[0073] Figure 57 Shows the reordering process in the decoder;

[0074] Figure 58 Shows the spatial part of the convolutional filter;

[0075] Figure 59 Shows the reference region (and its padding) used to derive filter coefficients;

[0076] Figure 60 Shows a square block for gradient calculation;

[0077] Figure 61 Shows a diamond block for gradient calculation;

[0078] Figure 62 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and

[0079] Figure 63 Shows a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.

[0080] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description

[0081] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustration purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than those described below.

[0082] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.

[0083] As used in this disclosure, terms such as "one embodiment", "an embodiment", "example embodiment" etc. indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment must include that specific feature, structure, or characteristic. Additionally, these phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in connection with an example embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.

[0084] It should be understood that although terms such as "first" and "second" etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiment. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.

[0085] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein indicate the presence of the stated features, elements, and / or components etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Example environment

[0086] Figure 1 is a block diagram showing an example video coding and decoding system 100 that can utilize the techniques of this disclosure. As shown, the video coding and decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0087] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.

[0088] Video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded pictures and associated data. The encoded picture is an encoded representation of a picture. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 through the network 130A. The encoded video data may also be stored on the storage medium / server 130B for access by the destination device 120.

[0089] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.

[0090] The video encoder 114 and the video decoder 124 may operate according to video compression standards (such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards).

[0091] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.

[0092] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 the example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.

[0093] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.

[0094] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0095] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for purposes of explanation, these components are shown separately in the Figure 2 example.

[0096] The segmentation unit 201 may segment a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0097] The mode selection unit 203 may select, for example, one coding mode among multiple coding modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select an Intra-Inter Combined Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0098] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the buffer 213 other than the picture associated with the current video block.

[0099] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on a current video block. For example, depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a part of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to parts of a picture composed of macroblocks that are independent of macroblocks in the same picture.

[0100] In some examples, the motion estimation unit 204 can perform uni-directional prediction on a current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0101] Alternatively, in other examples, the motion estimation unit 204 can perform bi-directional prediction on a current video block. The motion estimation unit 204 can search the reference pictures in list 0 to find one reference video block for the current video block, and can also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 can then generate a plurality of reference indexes and a plurality of motion vectors, where the plurality of reference indexes indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 can output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.

[0102] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.

[0103] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.

[0104] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0105] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0106] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0107] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0108] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0109] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0110] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.

[0111] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transformed coefficient video block, respectively, to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.

[0112] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce block effect artifacts in the video block.

[0113] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate the entropy encoded data and output a bitstream including the entropy encoded data.

[0114] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 an example of the video decoder 124 in the system 100 shown.

[0115] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 3 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.

[0116] In Figure 3 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.

[0117] The entropy decoding unit 301 may retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 may decode the entropy-coded video data, and the motion compensation unit 302 may determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 may determine such information, for example, by performing AMVP and the Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, the “Merge mode” may refer to deriving motion information from spatially or temporally adjacent blocks.

[0118] The motion compensation unit 302 may generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision may be included in the syntax element.

[0119] The motion compensation unit 302 may use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 may determine the interpolation filter used by the video encoder 200 according to the received syntax information, and the motion compensation unit 302 may use the interpolation filter to generate a prediction block.

[0120] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a “slice” may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.

[0121] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.

[0122] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the cache 307. The cache 307 provides reference blocks for subsequent motion compensation / intra prediction, and the cache 307 also generates the decoded video for presentation on a display device.

[0123] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to the multi-functional video codec or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview The present disclosure relates to video codec technology. Specifically, the present disclosure relates to a codec tool that utilizes multiple predefined weights and adaptive medians to fuse components to achieve better codec efficiency. The codec tool can be applied to existing video codec standards such as HEVC or multi-functional video codec (VVC). The codec tool is also applicable to future video codec standards or video codecs. 2. Introduction Video coding standards have mainly evolved from the well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure, where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created, working on the VVC standard with the goal of reducing the bit rate by 50% compared to HEVC. 2.1. Color Space and Chroma Subsampling A color space (also known as a color model (or color system)) is an abstract mathematical model that simply describes a color range as a digital tuple, typically 3 or 4 values or color components (e.g., RGB). Basically, a color space is a refinement of a coordinate system and subspace. For video compression, the most commonly used color spaces are YCbCr and RGB. YCbCr, Y'CbCr, or Y Pb / Cb Pr / Cr (also written as YCBCR or Y'CBCR) is a family of color spaces that are used as part of the color image pipeline in video and digital photography systems. Y' is the luminance component, and CB and CR are the blue-difference and red-difference chrominance components. Y' (with an apostrophe) is different from Y, which is luminance, meaning that the light intensity is non-linearly encoded based on gamma-corrected RGB primaries. Chroma subsampling is the practice of encoding an image by achieving a lower resolution for chrominance information than for luminance information, taking advantage of the fact that the human visual system is less sensitive to color differences than to luminance. 2.1.1. 4:4:4 Each of the three Y'CbCr components has the same sampling rate, so there is no chroma subsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2 The two chrominance components are sampled at half the sampling rate of luminance: the horizontal chrominance resolution is halved while the vertical chrominance resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one third with little visual difference. Examples of the nominal vertical and horizontal positions in the 4:2:2 color format are depicted in the Figure 4 VVC working draft. Figure 4 The nominal vertical and horizontal positions of 4:2:2 luma and chroma samples in a picture are shown. 2.1.3. 4:2:0 In 4:2:0, compared to 4:1:1, the horizontal sampling is doubled, but since the Cb and Cr channels are sampled only on every alternate row in this scheme, the vertical resolution is halved. Thus the data rate is the same. Both Cb and Cr are downsampled horizontally and vertically by a factor of 2. There are three variants of the 4:2:0 scheme with different horizontal and vertical positioning. · In MPEG-2, Cb and Cr are horizontally co-located. Cb and Cr are placed between the pixels in the vertical direction (placed with a gap). · In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are placed with a gap, at the mid-position between alternate luma samples. · In 4:2:0 DV, Cb and Cr are horizontally co-located. In the vertical direction, they are co-located on alternate rows. Table 1 SubWidthC and SubHeightC values derived from chroma_format_idc and separate_colour_plane_flag chroma_format_idc separate_colour_plane_flag Chrominance format SubWidthC SubHeightC 0 0 Monochrome 1 1 1 0 4:2:0 2 2 2 0 4:2:2 2 1 3 0 4:4:4 1 1 3 1 4:4:4 1 1 2.2. Encoding process of typical video codecs Figure 5 An example of an encoder block diagram is shown. Figure 5 An example of the encoder block diagram of VVC is shown, which includes three loop filter blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Different from DF which uses predefined filters, SAO and ALF utilize the original samples of the current picture, by adding an offset and by applying a finite impulse response (FIR) filter respectively, and use the encoded side information to signal the offset and filter coefficients to reduce the mean squared error between the original samples and the reconstructed samples. ALF is located at the last processing stage of each picture and can be regarded as a tool to try to capture and fix the artifacts generated by the previous stages. 2.3. Intra-mode codec with 67 intra prediction modes Figure 6shows 67 intra prediction modes. To capture any edge direction presented in natural videos, such as Figure 6 shown, the number of directional intra modes is extended from 33 used in HEVC to 65, and the planar mode and DC mode remain unchanged. These denser directional intra prediction modes apply to all block sizes and apply to both luma intra prediction and chroma intra prediction. In HEVC, each intra-coded block has a square shape and the length of each of its sides is a power of 2. Therefore, no division operation is required to generate the intra prediction value using the DC mode. In VVC, a block can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of a non-square block. 2.3.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. The conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VVC, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode coding and decoding method remains unchanged. Figure 7 shows the reference sample points for wide-angle intra prediction. To support these prediction directions, as Figure 7 shown, a top reference of length 2W+1 and a left reference of length 2H+1 are defined. The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2. Table 2 Intra prediction modes replaced by wide-angle modes Figure 8 shows the discontinuity problem in the case where the direction exceeds 45°. As Figure 8As shown, in the case of wide-angle frame prediction, two vertically adjacent predicted sample points can use two non-adjacent reference sample points. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap Δpα. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that satisfy this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted through these modes, the sample points in the reference cache are directly copied without applying any interpolation. With this modification, the number of sample points that need to be smoothed is reduced. In addition, it aligns the design of non-fractional modes in the conventional prediction mode with the design of non-fractional modes in the wide-angle mode. In VVC, in addition to the 4:2:0 chroma format, the 4:2:2 chroma format and the 4:4:4 chroma format are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC and the number of entries was extended from 35 to 67 to align with the extension of the intra prediction mode. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes with ranges from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the entries in the mapping table to more accurately transform the prediction angles for chroma blocks. 2.4. Intra Prediction Mode Coding and Decoding for Chrominance Components For the chrominance component of an intra PU, the encoder selects the best chroma prediction mode from five modes including planar, DC, horizontal, vertical, and a direct copy of the intra prediction mode for the luma component. The mapping between the intra prediction direction for chroma and the intra prediction mode number is shown in Table 3. When the intra prediction mode number of the chrominance component is 4, the intra prediction direction for the luma component is used for generating intra prediction sample points for the chrominance component. When the intra prediction mode number for the chrominance component is not 4 and is the same as the intra prediction mode number for the luma component, the intra prediction direction 66 is used for generating intra prediction sample points for the chrominance component. 2.5. Inter Prediction For each inter prediction CU, the motion parameters consist of a motion vector, a reference picture index and a reference picture list use index, and additional information required for the new VVC decoding features that will be used for inter prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded / decoded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no coded / decoded motion vector delta or reference picture index. The Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs, including spatial candidates and temporal candidates and additional scheduling introduced in VVC. The Merge mode can be applied to any inter prediction CU, not just for skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and reference picture list use flag for each reference picture list, and other required information are explicitly signaled for each CU. 2.6. Intra Block Copy (IBC) Intra Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding / decoding efficiency of screen content material. Since the IBC mode is implemented as a block-level coding / decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has been reconstructed within the current picture. The luminance block vector of the CU coded / decoded by IBC has integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. The CU coded / decoded by IBC is regarded as a third prediction mode in addition to the intra or inter prediction mode. The IBC mode is applicable to CUs with both width and height less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation for IBC is performed. The encoder performs RD checks on blocks with width or height no greater than 16 luminance samples. For non-Merge modes, the block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a local search based on block matching will be performed. In hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4x4 sub-blocks. For a current block of a larger size, when all the hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference positions, the hash key is determined to match the hash key of the reference block. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the minimum cost is selected. In block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode utilization flag is signaled, and the IBC mode can be signaled as the IBC AMVP mode or the IBC skip / Merge mode as follows: - IBC skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC decoded blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and pairwise candidates. - IBC AMVP mode: The block vector difference is decoded in the same way as the motion vector difference. The block vector prediction method uses two candidates as the prediction values, one from the left neighbor and one from the upper neighbor (if IBC decoded). When either neighbor is not available, the default block vector is used as the prediction value. A flag is signaled to indicate the block vector prediction value index. 2.7. Cross-Component Linear Model Prediction To reduce cross-component redundancy, the cross-component linear model (CCLM) prediction mode is used in VVC, for which the chroma samples are predicted based on the reconstructed luma samples of the same CU by using the following linear model: pred C (i,j) = α·rec L ′(i,j)+β (2-1) where, pred C (i,j) represents the predicted chroma sample in the CU, and rec L (i,j) represents the downsampled reconstructed luma sample of the same CU. The CCLM parameters (α and β) are derived using up to four neighboring chroma samples and their corresponding downsampled luma samples. Assuming the current chroma block dimensions are W×H, W’ and H’ are set as follows: – When the LM mode is applied, W' = W, H' = H; – When the LM_T mode is applied, W’ = W + H; – When applying the LM_L mode, H’ = H + W. The upper neighboring positions are represented as S[0, -1]…S[W' - 1, -1], and the left neighboring positions are represented as S[-1, 0]…S[-1, H' - 1]. Then, four samples are selected as follows: – When the LM mode is applied and both the upper neighboring sample and the left neighboring sample are available, S[W' / 4, -1], S[3*W' / 4, -1], S[-1, H' / 4], S[-1, 3*H' / 4]; – When the LM_T mode is applied or only the upper neighboring sample is available, S[W' / 8, -1], S[3*W' / 8, -1], S[5*W' / 8, -1], S[7*W' / 8, -1]; – When the LM_L mode is applied or only the left neighboring sample is available, S[-1, H' / 8], S[-1, 3* H' / 8], S[-1, 5*H' / 8], S[-1, 7*H' / 8]. The four neighboring luminance samples at the selected positions are downsampled and compared four times to find two larger values: x 0 A and x 1 A , and two smaller values: x 0 B and x 1 B . Their corresponding chrominance sample values are represented as y 0 A, y 1 A, y 0 B and y 1 B. Then x A , x B , y A and y B are derived as follows: X a =(x 0 A +x 1 A +1)>>1; X b =(x 0 B +x 1 B +1)>>1; Y a =(y 0 A +y 1 A +1)>>1; Y b =(y 0 B +y1 B +1) >> 1 (2-2) Finally, the linear model parameters α and β are obtained according to the following equations. β = Y b - α·X b (2-4) Figure 9 Examples of the positioning of the left and upper samples involved in the CCLM mode and the samples of the current block are shown. Figure 9 The positioning of the samples for deriving α and β is shown. The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store the table, the diff value (the difference between the maximum and minimum values) and the parameter α are represented by an exponential notation. For example, diff is approximated with a 4-bit significant part and an exponent. Therefore, the table of 1 / diff is reduced to 16 elements for 16 values of the significant bits as follows: DivTable[] = {0, 7, 6, 5, 5, 4, 4, 3, 3, 2, 2, 1, 1, 0} (2-5) This will have the advantage of reducing both the computational complexity and the memory size required to store the table. In addition to the upper template and the left template being used together to calculate the linear model coefficients, they can also be alternatively used in two other LM modes (referred to as LM_T and LM_L modes). In the LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is extended to (W + H) samples. In the LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is extended to (H + W) samples. In the LM mode, the left template and the upper template are used to calculate the linear model coefficients. To match the chrominance sample positions for 4:2:0 video sequences, two types of downsampling filters are applied to the luma samples to achieve a 2:1 downsampling rate in both the horizontal and vertical directions. The selection of the downsampling filter is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to the "type 0" and "type 2" contents respectively. Note that when the upper reference row is at the CTU boundary, only one luma row (the general row buffer in intra prediction) is used to produce the downsampled luma samples. This parameter calculation is performed as part of the decoding process, rather than just as an encoder search operation. Therefore, no syntax is used to convey the α and β values to the decoder. For chroma intra mode coding and decoding, a total of 8 intra modes are allowed for chroma intra mode coding and decoding. These modes include five regular intra modes and three cross-component linear model modes (LM, LM_T, and LM_L). The chroma modes are signaled and derived as shown in Table 3. The chroma mode coding and decoding directly depends on the intra prediction mode of the corresponding luma block. Since a separate block splitting structure for luma and chroma components is enabled in an I slice, a chroma block can correspond to multiple luma blocks. Therefore, for the chroma DM mode, the intra prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited. Table 3 Derivation of chroma prediction modes from luma modes when CCLM is enabled Regardless of the value of sps_cclm_enabled_flag, a single binarization table is used, as shown in Table 4. Table 4 Unified binarization table for chroma prediction modes Value of intra_chroma_pred_mode Binary string 4 00 0 0100 1 0101 2 0110 3 0111 5 10 6 110 7 111 In Table 4, the first binary bit indicates whether it is a regular mode (0) or an LM mode (1). If it is an LM mode, the next binary bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next 1 binary bit indicates whether it is LM_L (0) or LM_T (1). For this case, when sps_cclm_enabled_flag is 0, the first binary bit of the binarization table corresponding to intra_chroma_pred_mode can be discarded before entropy coding. Or, in other words, the first binary bit is presumed to be 0 and thus not coded. This single binarization table is used for both cases where sps_cclm_enabled_flag is equal to 0 and 1. The first two binary bits in Table 4 are context-coded using their own context model, and the remaining binary bits are bypass-coded. In addition, to reduce the luma-chroma delay in the double tree, when a 64×64 luma coding and decoding tree node is not split (Not Split) (and ISP is not used for 64×64 CUs) or QT is split, the chroma CUs in the 32×32 / 32×16 chroma coding and decoding tree nodes are allowed to use CCLM in the following way: – If the 32×32 chroma node is not split or split by QT split, all chroma CUs in the 32×32 node can use CCLM. – If the 32×32 chrominance node is horizontally BT split and the 32×16 sub - node is not partitioned or uses vertical BT partitioning, then all chrominance CUs in the 32×16 chrominance node can use CCLM. Under all other luma and chroma codec tree partitioning conditions, CCLM is not allowed for chrominance CUs. 2.8. Gradient Linear Model (GLM) Compared with CCLM, GLM uses the luma sample gradient to derive the linear model instead of the down - sampled luma value. In other words, during the CCLM process, the gradient G replaces the low - pass filtered luma samples. Other designs of CCLM (e.g., parameter derivation, prediction sample linear transformation) remain unchanged. C = α·G + β Figure 10A It is shown that G can be calculated by one of four Sobel - based gradient patterns. For signaling, when the CCLM mode is enabled for the current CU, two flags are signaled separately for the Cb / Cr components to indicate whether GLM is enabled for the component; if GLM is enabled for a component, a syntax element is further signaled to select Figure 10A one of the four gradient patterns in 2.9. Multiple - Model Linear Model (MMLM) With MMLM, there can be more than one linear model between the luma samples and chroma samples in a CU. In this method, the neighboring luma samples and neighboring chroma samples of the current block are classified into several groups, and each group is used as a training set to derive a linear model (i.e., specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified based on the same rule as the classification of neighboring luma samples. The neighboring samples can be classified into M groups, where M is 2 or 3. In addition to the original LM mode, the MMLM methods with M = 2 and M = 3 are also designed as two additional chroma prediction modes, called MMLM2 and MMLM3. The encoder selects the best mode during the RDO process and signals that mode. When M equals 2, Figure 10B An example of classifying neighboring samples into two groups is shown. The threshold is calculated as the average of the neighboring reconstructed luma samples. The neighboring samples with Rec'L[x,y] <= threshold are classified into group 1; while the neighboring samples with Rec'L[x,y] > threshold are classified into group 2. Similar to CCLM, there are 3 modes in MMLM, namely MMLM, MMLM_T, and MMLM_L. Two models are derived as follows. Figure 10BAn example of classifying neighboring samples into two groups is shown. The threshold is the average of neighboring samples for luminance reconstruction. The linear model for each class is derived using the Least Mean Square (LMS) method (if enabled) or the minimum / maximum method of VVC. 2.10. Position-Dependent Intra Prediction Combination In VVC, the results of intra prediction for DC, planar, and several angular modes are further modified by the Position-Dependent Intra Prediction Combination (PDPC) method. PDPC is an intra prediction method that calls a combination of boundary reference samples and HEVC-style intra prediction with filtered boundary reference samples. PDPC is applied to the following intra modes without signaling: planar, DC, intra angles less than or equal to horizontal, and intra angles greater than or equal to vertical and less than or equal to 80. If the current block is in BDPCM mode or the MRL index is greater than 0, PDPC is not applied. According to Equation 2-8 below, the predicted sample pred(x’,y’) is predicted using a linear combination of the intra prediction mode (DC, planar, angular) and reference samples: pred(x’,y’) = Clip(0,(1 << BitDepth)-1,(wL × R -1.y’ + wT × R x’ .-1+(64 - wL - wT) × pred(x’,y’) + 32) >> 6) (2-9) where R x,-1 ,R -1,y represent the reference samples located at the top boundary and left boundary of the current sample (x,y), respectively. If PDPC is applied to DC, planar, horizontal, and vertical intra modes, no additional boundary filters are required, as required in the case of the HEVC DC mode boundary filter or horizontal / vertical mode edge filters. The PDPC process for DC and planar modes is the same. For angular modes, if the current angular mode is HOR_IDX or VER_IDX, the left or top reference sample is not used, respectively. The PDPC weights and scaling factors depend on the prediction mode and block size. PDPC is applied to blocks with width and height both greater than or equal to 4. Figure 11A and Figure 8 B shows the definition of the reference samples (R x,-1 and R -1,y ) to which PDPC is applied for various prediction modes. The predicted sample pred(x’,y’) is located at (x’,y’) within the prediction block. For example, for the diagonal mode, the reference sample R x,-1The coordinate x is given by the following formula: x = x’ + y’ + 1, and for the reference sample point R -1,y the coordinate y is similarly given by the following formula: y = x’ + y’ + 1. For other angular patterns, the reference sample points R x,-1 and R -1,y can be located at fractional sample positions. In this case, the sample value at the nearest integer sample position will be used. Figure 11A - 11D illustrates the definition of the samples used by PDPC applied to diagonal and adjacent angular frame-in patterns. Figure 11A illustrates the diagonal upper-right pattern. Figure 11B illustrates the diagonal lower-left pattern. Figure 11C illustrates the adjacent diagonal upper-right pattern. Figure 11D illustrates the adjacent diagonal lower-left pattern. 2.11. Gradient PDPC As Figure 12 shown, the gradient-based method is extended to non-vertical / non-horizontal patterns. Here, the gradient is calculated as r(-1,y) – r(-1 + d, -1), where d is the horizontal displacement depending on the angular direction. There are several points to note here. The gradient term r(-1,y) – r(-1 + d, -1) needs to be calculated once for each row as it does not depend on the x position. The calculation of d is already part of the original intra prediction process that can be reused, so d does not need to be calculated separately. Thus, the precision of d is 1 / 32 pixel. When d is at a fractional position, we have used a two-tap (linear) filter, i.e., if dPos is the displacement with 1 / 32 pixel precision, dInt is the (rounded down) integer part (dPos >> 5), and dFract is the fractional part with 1 / 32 pixel precision (dPos & 31), then r(-1 + d) is calculated as: r(-1 + d) = (32 – dFrac)*r(-1 + dInt) + dFrac*r(-1 + dInt + 1). As described in a, this two-tap filter is executed once for each row (if needed). Finally, the predicted signal is calculated: p(x,y) = Clip(((64 – wL(x))*p(x,y) + wL(x)*(r(-1,y) - r(-1 + d, -1)) + 32) >> 6) where wL(x) = 32 >> ((x << 1) >> nScale2), and nScale2 = (log2(nTbH) + log2(nTbW) – 2) >> 2, which are the same as the vertical / horizontal mode. In short, the same process is applied as the vertical / horizontal mode (in fact, d = 0 indicates the vertical / horizontal mode). Figure 9 Shows the gradient method for non-vertical / non-horizontal modes. Secondly, when (nScale < 0) or PDPC cannot be applied due to the unavailability of the secondary reference sample points, the gradient-based method for non-vertical / non-horizontal modes is activated. We have shown in Figure 13 the values of nScale related to the TB size and angle mode to better visualize the case of using the gradient method. In addition, the flowcharts for the existing PDPC and the proposed PDPC are shown. Figure 13 Shows the nScale values with respect to nTbH and the mode number; for all cases where nScale < 0, the gradient method is used. Figure 14 Shows the flowcharts: the current PDPC (left) and the proposed PDPC (right). 2.12. Secondary MPM The secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, while the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed, and then the first 6 entries in this general MPM table are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is the planar mode. As Figure 15 shown, the remaining entries are composed of the intra modes of the left (L), above (A), bottom-left (BL), top-right (AR), and top-left (AL) neighboring blocks, the band direction modes with an offset added from the first two available band direction modes of the neighboring blocks, and the default mode. If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. Figure 15 Shows the neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list. First, the PMPM flag is parsed. If it is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list is selected; otherwise, the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining modes. 2.13.6 Tapped Intra Interpolation Filter To improve the prediction accuracy, it is proposed to replace the 4-tap cubic interpolation filter with a 6-tap interpolation filter. The filter coefficients are derived based on the same polynomial regression model, but the polynomial order is 6. The filter coefficients are as follows: {0, 0, 256, 0, 0, 0}, / / 0 / 32 position {0, -4, 253, 9, -2, 0}, / / 1 / 32 position {1, -7, 249, 17, -4, 0}, / / 2 / 32 position {1, -10, 245, 25, -6, 1}, / / 3 / 32 position {1, -13, 241, 34, -8, 1}, / / 4 / 32 position {2, -16, 235, 44, -10, 1}, / / 5 / 32 position {2, -18, 229, 53, -12, 2}, / / 6 / 32 position {2, -20, 223, 63, -14, 2}, / / 7 / 32 position {2, -22, 217, 72, -15, 2}, / / 8 / 32 position {3, -23, 209, 82, -17, 2}, / / 9 / 32 position {3, -24, 202, 92, -19, 2}, / / 10 / 32 position {3, -25, 194, 101, -20, 3}, / / 11 / 32 position {3, -25, 185, 111, -21, 3}, / / 12 / 32 position {3, -26, 178, 121, -23, 3}, / / 13 / 32 position {3, -25, 168, 131, -24, 3}, / / 14 / 32 position {3, -25, 159, 141, -25, 3}, / / 15 / 32 position {3, -25, 150, 150, -25, 3}, / / half-pixel position The reference samples for interpolation are taken from the reconstructed samples or the filled samples as in HEVC, so that no conditional check for the availability of the reference samples is required. It is proposed to use a 4-tap cubic interpolation filter instead of the nearest rounding operation to derive the extended intra reference samples. As shown in the example in Figure 16 To derive the value of the reference sample P, a four-tap interpolation filter is used, while in JEM-3.0 or HM, P is directly set to X1. Figure 16 An example of the proposed intra reference mapping is shown. 2.14. Multiple reference line (MRL) intra prediction Multiple reference line (MRL) intra prediction uses more reference lines for intra prediction. In Figure 17 , an example of 4 reference lines is depicted, where the samples of segment A and segment F are not obtained from the reconstructed neighboring samples, but are filled with the closest samples from segment B and segment E respectively. HEVC intra picture prediction uses the nearest reference line (i.e., reference line 0). In MRL, 2 additional lines (reference line 1 and reference line 2) are used. The index (mrl_idx) of the selected reference line is signaled and used to generate the intra prediction value. For reference line indices greater than 0, only the additional reference line mode is included in the MPM list, and only the MPM index is signaled without signaling the remaining modes. The reference line index is signaled before the intra prediction mode, and in the case where a non-zero reference line index is signaled, the planar mode is excluded from the intra prediction modes. Figure 17 An example of four reference lines adjacent to the prediction block is shown. MRL is disabled for the first row of blocks inside the CTU to prevent the use of extended reference samples outside the current CTU row. Additionally, PDPC is disabled when the additional lines are used. For the MRL mode, the derivation of the DC value in the DC intra prediction mode for non-zero reference line indices is aligned with the derivation for reference line index 0. MRL requires storing 3 neighboring luma reference lines of the CTU to generate the prediction. The cross-component linear model (CCLM) tool also requires 3 neighboring luma reference lines for its downsampling filter. The definition of MRL using the same 3 lines is aligned with CCLM to reduce the storage requirements for the decoder. 2.15. Intra sub-partitioning (ISP) Intra sub-partitioning (ISP) divides the luma intra prediction block vertically or horizontally into 2 or 4 sub-partitions according to the block size. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is greater than 4×8 (or 8×4), the corresponding block will be divided into 4 sub-partitions. It has been noted that M×128 (M≤64) and 128×N (N≤64) ISP blocks may cause potential problems for the 64×64 VDPU. For example, an M×128 CU in the single-tree case has an M×128 luma TB and two corresponding Chroma TB. If the CU uses ISP, the luminance TB will be divided into 4 M×32 TBs (only horizontal division is possible), and each of these TBs is smaller than a 64×64 block. However, in the current ISP design, chroma blocks are not divided. Therefore, the sizes of both chroma components will be larger than 32×32 blocks. Similarly, using ISP with 128×NCU can cause a similar situation. Therefore, these two cases are problems for the 64×64 decoder pipeline. Thus, the CU size that can use ISP is limited to a maximum of 64×64. Figure 18A and Figure 18B Examples showing two possibilities are presented. All sub - partitions satisfy the condition of having at least 16 samples. In ISP, dependence of the 1×N / 2×N sub - block prediction on the reconstructed value of the previously decoded 1×N / 2×N sub - block of the coded - decoded block is not allowed, such that the minimum prediction width of the sub - block becomes four samples. For example, an 8×N (N>4) coded - decoded block coded using ISP with vertical division is divided into two prediction regions each of size 4×N and four transforms of size 2×N. Additionally, a 4×N coded - decoded block coded using ISP with vertical division is predicted using the full 4×N block; four 1×N transforms are used. Although 1×N and 2×N transform sizes are allowed, it is asserted that the transforms of these blocks in the 4×N region can be executed in parallel. For example, when the 4×N prediction region contains four 1×N transforms, there is no transform in the horizontal direction; the transforms in the vertical direction can be executed as a single 4×N transform in the vertical direction. Similarly, when the 4×N prediction region contains two 2×N transform blocks, the transform operations of the two 2×N blocks in each direction (horizontal and vertical) can be performed in parallel. Therefore, no additional latency is added when processing these smaller blocks compared to processing an intra - frame block of 4×4 conventional coding - decoding. Figure 18A and 15 B shows the sub - partition depending on the block size. Figure 18A Examples of sub - partitions of 4×8 and 8×4 CUs are presented. Figure 18B Examples of sub - partitions of CUs other than 4×8, 8×4, and 4×4 are presented. Table 5 Entropy Coding Coefficient Group Sizes Block size Coefficient set size 1×N, N≥16 1×16 N×1, N≥16 16×1 2×n, n≥8 2×8 n×2, n≥8 8×2 All other possible M×N cases For each sub - partition, the reconstructed samples are obtained by adding the residual signal to the predicted signal. Here, the residual signal is generated by processes such as entropy decoding, inverse quantization, and inverse transformation. Thus, the reconstructed sample values of each sub - partition can be used to generate the prediction for the next sub - partition, and each sub - partition is processed repeatedly. Also, the first sub - partition to be processed is the one containing the top - left samples of the CU, and then it continues down (for horizontal partitioning) or to the right (for vertical partitioning). As a result, the reference samples used to generate the sub - partition prediction signals are only on the left and above the row. All sub - partitions share the same intra - frame mode. The following is a summary of the interaction between ISP and other codec tools. – Multiple Reference Lines (MRL): If the block has an MRL index other than 0, the ISP codec mode is presumed to be 0, so the ISP mode information will not be sent to the decoder. – Entropy - coded coefficient group size: The size of the entropy - coded sub - blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 5. Note that the new size only affects the blocks generated by ISP where one of the dimensions is less than 4 samples. In all other cases, the coefficient group remains 4×4 dimension. – CBF coding: Assume that at least one sub - partition has a non - zero CBF. Thus, if n is the number of sub - partitions and the first n - 1 sub - partitions have produced zero CBF, the CBF of the nth sub - partition is presumed to be 1. – Transform size constraint: All ISP transforms with a length greater than 16 points use DCT - II. – MTS flag: If the CU uses the ISP codec mode, the MTS CU flag will be set to 0 and it will not be sent to the decoder. Thus, the encoder will not perform RD tests for the different available transforms for each resulting sub - partition. Instead, the transform selection for the ISP mode will be fixed and selected based on the intra - frame mode utilized, the processing order, and the block size. Thus, no signaling is required. For example, let t H and t V be the horizontal and vertical transforms selected for a w×h sub - partition respectively, where w is the width and h is the height. Then the transform is selected according to the following rules: – If w = 1 or h = 1, there is no horizontal transform or vertical transform respectively. – If w≥4 and w≤16, t H = DST - VII, otherwise t H = DCT - II. – If h≥4 and h≤16, t V = DST - VII, otherwise t V = DCT - II. In the ISP mode, all 67 intra prediction modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 samples long. Additionally, the conditions for the reference sample filtering process (reference smoothing) and the intra interpolation filter selection no longer exist, and the cubic (DCT-IF) filter is always applied for fractional position interpolation in the ISP mode. 2.16. Matrix-Weighted Intra Prediction (MIP) The matrix-weighted intra prediction (MIP) method is a newly added intra prediction technique in VVC. To predict the samples of a rectangular block with width W and height H, the matrix-weighted intra prediction (MIP) takes as input a row of H reconstructed neighboring boundary samples to the left of the block and a row of W reconstructed neighboring boundary samples above the block. If the reconstructed samples are not available, they are generated in the same way as in conventional intra prediction. The generation of the prediction signal is based on the following three steps, namely averaging, matrix-vector multiplication, and linear interpolation, as Figure 19 shown. Figure 19 Shows the matrix-weighted intra prediction process. 2.16.1. Averaging Neighboring Samples Among the boundary samples, four or eight samples are selected by averaging based on the block size and shape. Specifically, according to a predefined rule depending on the block size, by averaging the neighboring boundary samples, the input boundaries bdry top and bdry left are reduced to smaller boundaries and Then, the two reduced boundaries and are concatenated to the reduced boundary vector bdry red , which has a size of 4 for a 4×4-shaped block and a size of 8 for all other shaped blocks. If mode refers to the MIP mode, this concatenation is defined as follows: 2.16.2. Matrix Multiplication Taking the averaged samples as input, matrix-vector multiplication is performed and then an offset is added. The result is a reduced prediction signal on the downsampled set of samples in the original block. From the reduced input vector bdry red , a reduced prediction signal pred red is generated, which is a signal on the downsampled block with width W red and height H red . Here, W red and H red are defined as: Compute the reduced prediction signal pred by computing the matrix-vector product and adding an offset red : pred red = A · bdry red + b (2-13) Here, if W = H = 4, then A is a matrix with W red · H red rows and 4 columns, and in all other cases has 8 columns. b is a vector with dimension W red · H red . The matrix A and the offset vector b are taken from one of the sets S0, S1, S2. An index idx = idx(W,H) is defined as follows: Here, each coefficient of the matrix A is represented with 8-bit precision. The set S0 consists of 16 matrices i ∈ {0,…,15}, each matrix having 16 rows and 4 columns, and 16 offset vectors i ∈ {0,…,16}, each offset vector having dimension 16. The matrices and offset vectors of this set are used for blocks of size 4×4. The set S1 consists of 8 matrices i ∈ {0,…,7}, each matrix having 16 rows and 8 columns, and 8 offset vectors i ∈ {0,…,7}, each offset vector having dimension 16. The set S2 consists of 6 matrices i ∈ {0,…,5}, each matrix having 64 rows and 8 columns, and 6 offset vectors i ∈ {0,…,5}, each offset vector having dimension 64. 2.16.3. Interpolation The prediction signal at the remaining positions is generated by linear interpolation from the prediction signals on the downsampled set, which is a single-step linear interpolation in each direction. The interpolation is first performed in the horizontal direction and then in the vertical direction, independent of the shape or size of the block. 2.16.4. Signaling of the MIP mode and coordination with other codec tools For each coding unit (CU) in the intra mode, a flag indicating whether to apply the MIP mode is sent. If the MIP mode is to be applied, the MIP mode (predModeIntra) is signaled. For the MIP mode, the transpose flag (isTransposed) (which determines whether the mode is transposed) and the MIP mode Id (modeId) (which determines which matrix is to be used for a given MIP mode) are derived as follows. isTransposed = predModeIntra & 1 modeId = predModeIntra >> 1 (2 - 15) The MIP encoding / decoding mode is coordinated with other encoding / decoding tools by considering the following aspects: – Enable LFNST for MIP on large blocks. The LFNST transform in planar mode is used here. – The derivation of reference samples for MIP is performed in exactly the same way as that for conventional intra prediction modes. – For the upsampling step used in MIP prediction, the original reference samples are used instead of the downsampled reference samples. – Clipping is performed before upsampling instead of after upsampling. – MIP allows up to 64×64 regardless of the maximum transform size. The number of MIP modes is 32 for sizeId = 0, 16 for sizeId = 1, and 12 for sizeId = 2. 2.17. Intra mode derivation on the decoder side In JEM - 2.0, the intra modes are extended from 35 modes in HEVC to 67 modes, and they are derived at the encoder and signaled explicitly to the decoder. In JEM - 2.0, a large amount of overhead is spent on intra mode encoding / decoding. For example, in all intra encoding / decoding configurations, the intra mode signaling overhead can be as high as 5% - 10% of the total bitrate. This paper proposes an intra mode derivation method on the decoder side to reduce the intra mode encoding / decoding overhead while maintaining the prediction accuracy. To reduce the overhead of intra mode signaling, this paper proposes a decoder - side intra mode derivation (DIMD) method. In the proposed method, instead of signaling the intra mode explicitly, information is derived from the neighboring reconstructed samples of the current block at both the encoder and the decoder. There are two ways to use the intra mode derived by DIMD: 1) For a 2N×2N CU, when the corresponding CU - level DIMD flag is enabled, the DIMD mode is used as the intra mode for intra prediction; 2) For an N×N CU, the DIMD mode is used to replace a candidate in the existing MPM list to improve the efficiency of intra mode encoding / decoding. 2.17.1. Template - based intra mode derivation As Figure 20 shown, the target represents the current block (block size is N) for which the intra prediction mode is to be estimated. The template (denoted by Figure 20The pattern area (indicated in) specifies a set of reconstructed samples that are used to derive the intra mode. The template size is expressed as the number of samples extending above and to the left of the target block within the template, i.e., L. In the current implementation, the template size used for 4×4 and 8×8 blocks is 2 (i.e., L = 2), and the template size used for 16×16 and larger blocks is 4 (i.e., L = 4). As defined in JEM-2.0, the reference of the template (indicated by the dashed area in Figure 20 is the set of neighboring samples above and to the left of the template. Different from the template samples that always come from the reconstructed area, the reference samples of the template may not have been reconstructed when encoding / decoding the target block. In this case, the existing reference sample replacement algorithm of JEM-2.0 is used to replace the unavailable reference samples with available reference samples. Figure 20 shows the target samples, template samples, and reference samples of the template used in DIMD. For each intra prediction mode, DIMD calculates the sum of absolute differences (SAD) between the reconstructed template samples and their predicted samples obtained from the reference samples of the template. The intra prediction mode that produces the minimum SAD is selected as the final intra prediction mode of the target block. 2.17.2. DIMD for Intra 2N×2N CU For intra 2N×2N CU, DIMD is used as an additional intra mode, which is adaptively selected by comparing the DIMD intra mode with the best normal intra mode (i.e., signaled explicitly). A flag is signaled for each intra 2N×2N CU to indicate the use of DIMD. If the flag is 1, the intra mode derived by DIMD is used to predict the CU; otherwise, DIMD is not applied, and the intra mode signaled explicitly in the bitstream is used to predict the CU. When DIMD is enabled, the chrominance components always reuse the same intra mode derived for the luminance component, i.e., the DM mode. In addition, for each CU encoded / decoded by DIMD, the blocks within the CU can adaptively select to derive their intra mode at the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is signaled to indicate at which level DIMD is executed. If the flag is 0, it means that DIMD is executed at the PU level, and all TUs within the PU use the same derived intra mode for their intra prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is executed at the TU level, and each TU within the PU derives its own intra mode. In addition, when DIMD is enabled, the number of angular directions is increased to 129, and DC and planar modes remain unchanged. To accommodate the increased granularity of the intra-frame mode, the accuracy of intra-frame interpolation filtering for DIMD-encoded / decoded CUs is increased from 1 / 32 pixel to 1 / 64 pixel. Additionally, to use the derived intra-frame mode of DIMD-encoded / decoded CUs as MPM candidates for neighboring intra-blocks, these 129 directions of DIMD-encoded / decoded CUs are converted to "normal" intra-frame modes (i.e., 65 angular intra-frame directions) before being used as MPMs. 2.17.3. DIMD for Intra N×N CUs In the proposed method, the intra-frame mode of intra N×N CUs is always signaled. However, to improve the efficiency of intra-frame mode coding / decoding, the intra-frame modes derived from DIMD are used as MPM candidates for predicting the intra-frame modes of four PUs in a prediction CU. To avoid increasing the overhead of MPM index signaling, DIMD candidates are always placed in the first position in the MPM list, and the last existing MPM candidate is removed. Additionally, a deduplication operation is performed such that if a DIMD candidate is redundant, it will not be added to the MPM list. 2.17.4. Intra-frame Mode Search Algorithm for DIMD To reduce the encoding / decoding complexity, a straightforward fast intra-frame mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra-frame mode search. Specifically, an initial candidate list is created by selecting N fixed modes from the allowed intra-frame modes. Then, the SAD is calculated for all candidate intra-frame modes, and an intra-frame mode that minimizes the SAD is selected as the starting intra-frame mode. To achieve a good complexity / performance trade-off, the initial candidate list consists of 11 intra-frame modes, including DC, planar, and every 4th mode among the 33 angular intra-frame directions defined in HEVC, i.e., intra-frame modes 0, 1, 2, 6, 10…30, 34. If the starting intra-frame mode is DC or planar, it is used as the DIMD mode. Otherwise, based on the starting intra-frame mode, a refinement process is then applied, where the best intra-frame mode is identified through an iterative search. It works by comparing the SAD values of three intra-frame modes separated by a given search interval in each iteration and maintaining the intra-frame mode that minimizes the SAD. Then the search interval is reduced to half, and the selected intra-frame mode from the last iteration will be used as the center intra-frame mode for the current iteration. For the current DIMD implementation with 129 angular intra-frame directions, up to 4 iterations are used in the refinement process to find the best DIMD intra-frame mode. 2.18. Decoder-side Intra-frame Mode Derivation by Calculating Gradients of Neighboring Samples Three angular modes are selected from the Histogram of Oriented Gradients (HoG) calculated from the neighboring pixels of the current block. Once the three modes are selected, their prediction values are calculated normally, and then their weighted average is used as the final prediction value of the block. To determine the weights, the corresponding magnitudes in the HoG are used for each of the three modes. The DIMD mode is used as an alternative prediction mode and is always checked in the FullRD mode. Some aspects of the current version of DIMD are modified in signaling, HoG calculation, and prediction fusion. The purpose of this modification is to improve the codec performance and to address the complexity issue (i.e., throughput of 4x4 blocks) raised during the last meeting. The following sections describe the modifications in each aspect. 2.18.1. Signaling Figure 21 The order of the parsing flags / indexes integrated with the proposed DIMD in VTM5 is shown. It can be seen that a single CABAC context is first used to parse the DIMD flag of the block, and this single CABAC context is initialized to the default value 154. If flag == 0, the parsing continues normally. Otherwise (if flag == 1), only the ISP index is parsed, and the following flags / indexes are presumed to be zero: BDPCM flag, MIP flag, MRL index. In this case, the entire IPM parsing is also skipped. During the parsing phase, when a regular non-DIMD block queries the IPM of its DIMD neighbor, the mode PLANAR_IDX is used as the virtual IPM of the DIMD block. Figure 21 The proposed intra block decoding process is shown. 2.18.2. Texture analysis The texture analysis of DIMD includes the Histogram of Oriented Gradients (HoG) calculation ( Figure 22 ). The Histogram of Oriented Gradients calculation is performed by applying horizontal and vertical Sobel filters to the pixels in a template of width 3 around the block. However, if the upper template pixels fall into different CTUs, they will not be used for texture analysis. Once calculated, the IPMs corresponding to the two highest histogram bins are selected for the block. In the previous version, all pixels in the middle row of the template participated in the HoG calculation. However, the current version improves the throughput of this process by applying the Sobel filter more sparsely on 4x4 blocks. For this purpose, only one pixel from the left and one pixel from the top are used. As Figure 22 shown. In addition to reducing the number of operations for gradient calculation, this feature also simplifies the selection of the best two patterns from the HoG, because the resulting HoG cannot have more than two non-zero magnitudes. Figure 22 Shows the HoG calculation from a template with a width of 3 pixels. 2.18.3. Prediction Fusion The current version of this method also uses the fusion of three prediction values for each block. However, the selection of the prediction mode is different, and the combined hypothesis intra prediction method proposed in [2] is utilized, where the planar mode is considered to be used in combination with other modes when calculating the intra prediction candidates. In the current version, the two IPMs corresponding to the two highest HoG bars are combined with the planar mode. Prediction fusion is applied as a weighted average of the above three prediction values. For this purpose, the weight of the plane is fixed at 21 / 64 (~1 / 3). Then the remaining 43 / 64 (~2 / 3) of the weight is shared between the two HoG IPMs, proportional to the magnitudes of their HoG bars. Figure 23 Illustrates this process. Figure 23 Shows the prediction fusion by weighted averaging of two HoG modes and the plane. 2.19. DIMD Chrominance Mode As Figure 24 shown, the DIMD chrominance mode uses the DIMD derivation method to derive the intra prediction mode of the current block's chrominance based on the neighboring reconstructed Y, Cb, and Cr samples in the second neighboring rows and columns. Specifically, for each co-located reconstructed luminance sample and the reconstructed Cb and Cr samples of the current chrominance block, the horizontal gradient and the vertical gradient are calculated to construct the HoG. Then, the intra prediction mode with the maximum histogram magnitude value is used to perform the intra prediction of the current chrominance block's chrominance. Figure 24 Shows the neighboring reconstructed samples for the DIMD chrominance mode. When the intra prediction mode derived from the DIMD chrominance mode is the same as the intra prediction mode derived from the DM mode, the intra prediction mode with the second largest histogram magnitude value is used as the DIMD chrominance mode. The CU-level flag is signaled to indicate whether the proposed DIMD chrominance mode is applied. 2.20. Template-Based Intra Mode Derivation (TIMD) This paper proposes a template-based intra mode derivation (TIMD) method using the MPM, where the TIMD mode is derived from the MPM using neighboring templates. The TIMD mode is used as an additional intra prediction method for the CU. 2.20.1. TIMD Mode Derivation For each intra prediction mode in MPM, the SATD between the prediction of the template and the reconstructed samples is calculated. The intra prediction mode with the minimum SATD is selected as the TIMD mode and is used for the intra prediction of the current CU. The position-dependent intra prediction combination (PDPC) is included in the derivation of the TIMD mode. 2.20.2. TIMD Signaling A flag is signaled in the sequence parameter set (SPS) to enable / disable the proposed method. When the flag is true, a CU-level flag is signaled to indicate whether the proposed TIMD method is used. The TIMD flag is signaled after the MIP flag. If the TIMD flag is equal to true, the remaining syntax elements related to the luma intra prediction mode, including MRL, ISP, and the normal parsing stage for the luma intra prediction mode, are skipped. 2.20.3. Interaction with New Decoding Tools The DIMD method using planar prediction fusion is integrated into EE2. When the EE2 DIMD flag is equal to true, the proposed TIMD flag is not signaled and is set to be equal to false. Similar to PDPC, gradient PDPC is also included in the derivation of the TIMD mode. When the secondary MPM is enabled, both the primary MPM and the secondary MPM are used to derive the TIMD mode. The 6-tap interpolation filter is not used for the derivation of the TIMD mode. 2.20.4. Modification of MPM List Construction in TIMD Mode Derivation During the construction of the MPM list, the intra prediction mode of the neighboring block is derived as planar when it is inter-coded. To improve the accuracy of the MPM list, when the neighboring block is inter-coded, the propagated intra prediction mode is derived using the motion vector and the reference picture and is used in the construction of the MPM list. This modification is only used for the derivation of the TIMD mode. 2.20.5. TIMD with Fusion This paper proposes that for the intra mode derived using the TIMD method, instead of selecting only one mode with the minimum SATD cost, the top two modes with the minimum SATD cost are selected, then they are fused by weights, and this weighted intra prediction is used to encode and decode the current CU. The costs of the two selected modes are compared with a threshold, and a cost factor of 2 is applied in the test as follows: costMode2 < 2′costMode1. If this condition is true, fusion is applied; otherwise, only mode1 is used. The weights of the modes are calculated according to their SATD costs as follows: weight1 = costMode2 / (costMode1 + costMode2c); weight2 = 1 - weight1. 2.21. Merge Mode with MVD (MMVD) In addition to the Merge mode, in the case where the implicitly derived motion information is directly used for the prediction sample generation of the current CU, a Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag is sent to specify whether the MMVD mode is used for the CU. In MMVD, after a Merge candidate is selected, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, the index for specifying the motion amplitude, and the index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The MMVD candidate flag is signaled to specify which one is used between the first Merge candidate and the second Merge candidate. Figure 25 The MVD search points are shown. The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point. As Figure 25 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 6. Table 6 - Relationship between Distance Index and Predefined Offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown in Table 7. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where both lists point to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 7 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV where the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture), and the POC difference in list 0 is greater than the POC difference in list 1, the symbol in Table 7 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Otherwise, if the POC difference in list 1 is greater than the POC difference in list 0, the symbol in Table 7 specifies the sign of the MV offset added to the list 1 MV component of the starting MV, and the sign of the list 0 MV has the opposite value. The MVD is scaled according to the POC differences in each direction. If the POC differences in both lists are the same, no scaling is required. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, then as Figure 26 described in, the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb. If the POC difference of L1 is greater than the POC difference of L0, the MVD of list 0 is scaled in the same way. If the starting MV is unidirectionally predicted, the MVD is added to the available MV. Table 7 - Signs of MV Offsets Specified by the Direction Index Direction index 00 01 10 11 x - axis + - N / A N / A y - axis N / A N / A + - 2.22. Symmetric MVD Coding and Decoding In VVC, in addition to the regular unidirectional prediction mode MVD signaling and bidirectional prediction mode MVD signaling, the symmetric MVD mode for bidirectional prediction MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indices of both list 0 and list 1 and the MVD of list 1 is not signaled but derived. The decoding process of the symmetric MVD mode is as follows: 1. At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0. – Otherwise, if the nearest reference picture in List 0 and the nearest reference picture in List 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the List 0 reference picture and the List 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2. At the CU level, if the CU is coded / decoded using bidirectional prediction and BiDirPredFlag is equal to 1, then the symmetry mode flag indicating whether the symmetry mode is used is signaled explicitly. When the symmetry mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indices for List 0 and List 1 are set to be equal to the reference picture pair, respectively. MVD1 is set to be equal to (-MVD0). The final motion vector is shown in the following formula. Figure 26 The symmetric MVD mode is shown. In the encoder, symmetric MVD motion estimation starts from the initial MV evaluation. A set of initial MV candidates includes the MVs obtained from unidirectional prediction search, the MVs obtained from bidirectional prediction search, and the MVs from the AMVP list. The one with the lowest distortion rate cost is selected as the initial MV for the symmetric MVD motion search. 2.23. Bidirectional Optical Flow (BDOF) The bidirectional optical flow (BDOF) tool is included in VVC. BDOF (previously called BIO) is included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version, which requires much less computation, especially in terms of the number of multiplications and the multiplier size. BDOF is used to refine the bidirectional prediction signal of the CU at the 4x4 sub-block level. BDOF is applied to a CU if all of the following conditions are met: – The CU is coded / decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other of the two reference pictures is after the current picture in display order. – The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same. – Both of the two reference pictures are short-term reference pictures. – The CU is not coded / decoded using the affine mode or the SbTMVP Merge mode. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. –WP is not enabled for the current CU. –CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As its name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bi-predicted sample values in the 4x4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals, and k = 0, 1, are calculated by directly computing the difference between two neighboring samples, i.e., where I (k) (i,j) is the sample value at the prediction signal coordinates (i,j) in the list k, k = 0, 1, and shift1 is calculated based on the luma bit depth bitDepth as shift1 = max(6, bitDepth - 6). Then, the auto-correlations and cross-correlations of the gradients S1, S2, S3, S5 and S6 are calculated as follows where where Ω is a 6×6 window around the 4×4 sub-block, and the values of n a and n b are set to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then, using the cross-correlation terms and auto-correlation terms, the motion refinement (v x ,v y ) is derived using the following method: where th′ BIO = 2 max(5,BD-7) . is the floor function, and Based on the motion refinement and gradients, the following adjustment is calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bi-predicted samples in the following manner: pred BDOF (x,y) = (I (0) (x, y) + I (1) (x,y) + b(x,y) + o offset ) >> shift (2 - 22) These values are selected so that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient values, some prediction samples I (k) (i,j) in list k (k = 0,1) outside the current CU boundary need to be generated. As Figure 25 depicted, the BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, the prediction samples (white positions) in the extended region are generated by directly taking the reference samples at the nearby integer positions (using the floor() operation on the coordinates) without using interpolation, and the regular 8 - tap motion - compensated interpolation filter is used to generate the prediction samples (gray positions) inside the CU. These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e., repeated) from their nearest neighbors. Figure 27 Shows the extended CU region used in BDOF. When the width and / or height of the CU is greater than 16 luma samples, it is divided into sub - blocks with width and / or height equal to 16 luma samples, and the sub - block boundaries are considered as the CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub - block, the BDOF process can be skipped. When the SAD between the initial L0 prediction sample and the L1 prediction sample is less than the threshold, the BDOF process is not applied to the sub - block. The threshold is set to be equal to (8 * W * (H >> 1), where W indicates the sub - block width and H indicates the sub - block height. To avoid the additional complexity of SAD calculation, the SAD calculated in the DVMR process between the initial L0 prediction sample and the L1 prediction sample is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., the luma_weight_lx_flag of any one of the two reference pictures is 1, the BDOF is also disabled. When the CU is encoded / decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.24. Intra - Inter Joint Prediction (CIIP) 2.25. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, in the real world, there are many kinds of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As Figure 28 shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). Figure 28 Shows the affine motion model based on control points. In Figure 28 , both the 4-parameter affine model and the 6-parameter affine model are shown. For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: where (mv 0x , mv 0y ) is the motion vector of the upper left control point, (mv 1x , mv 1y ) is the motion vector of the upper right control point, and (mv 2x , mv 2y ) is the motion vector of the lower left control point. To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector of each 4×4 luminance sub-block, the motion vector of the central sample of each sub-block is calculated according to the above equations (as Figure 29 shown) and rounded to 1 / 16 fractional precision. Then a motion compensation interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size of the chrominance component is also set to 4×4. The MV of a 4×4 chrominance sub-block is calculated as the average of the MVs of four corresponding 4×4 luminance sub-blocks. Figure 29 Shows the affine MVF of each sub-block. Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.25.1. Affine Merge Prediction The AF_MERGE mode can be applied to CUs with both width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be no more than five CPMV candidates, and one to be used for the current CU is indicated by a signaling index. The following three types of CPMV candidates are used to form the affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs. – Constructed affine Merge candidate CPMVPs derived using the translational MVs of neighboring CUs. – Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as Figure 30 shown. For the prediction values on the left, the scan order is A0->A1, and for the prediction values on the upper side, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When the neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidates in the affine Merge list of the current CU. As shown in the figure, if the neighboring lower-left block A is coded in the affine mode, the motion vectors v2, v3, and v4 of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated based on v2 and v3. When block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated based on v2, v3, and v4. Figure 30 Shows the positions of the inherited affine motion prediction values. Figure 31 Shows the inheritance of the control point motion vectors. The constructed affine candidates refer to constructing candidates by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the specified spatial neighbors and temporal neighbors shown in Figure 31 , and CPMV k (k = 1, 2, 3, 4) represents the kth control point. For CPMV1, the B2->B3->A2 block is checked, and the MV of the first available block is used. For CPMV2, the B1->B0 block is checked, and for CPMV3, the A1->A0 block is checked. The TMVP is used as CPMV4 (if available). After obtaining the MVs of the four control points, affine Merge candidates are constructed based on this motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}. Combinations of 3 CPMVs construct 6-parameter affine Merge candidates, and combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of the control point MVs are discarded. Figure 32 Shows the positioning of candidate positions for the affine Merge mode used for construction. After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list. 2.25.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with both width and height greater than or equal to 16. In the bitstream, an affine flag at the CU level is signaled to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and it is generated by sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMVs of neighboring CUs. – Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs. – Translational MVs from neighboring CUs. – Zero MVs. The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting the inherited affine motion prediction values into the candidate list, the deduplication process is not applied. The constructed AMVP candidates are from Figure 15It is derived from the specified spatial neighborhood shown. The same checking order as in affine Merge candidate construction is used. Additionally, the reference picture indices of neighboring blocks are also checked. Using the first block in the checking order, this block is inter-frame coded and has the same reference picture as in the current CU. There is only one. When the current CU is coded using the 4-parameter affine mode and both mv0 and mv1 are available, they are added as a candidate in the affine AMVP list. When the current CU is coded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable. If the affine AMVP list candidate is still less than 2 after checking the inherited affine AMVP candidate and the constructed AMVP candidate, mv0, mv1, and mv2 will be added in order as translational MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list. 2.25.3. Affine Motion Information Storage In VVC, the CPMVs of affine CUs are stored in a separate cache. The stored CPMVs are only used to generate the inherited CPMVs in the affine Merge mode and the inherited CPMVs in the affine AMVP mode for the most recently coded CU. The sub-block MVs derived from the CPMVs are used for motion compensation, MV derivation for the Merge / AMVP list of translational MVs, and deblocking. To avoid picture line caching for additional CPMVs, the affine motion data inheritance from CUs in the CTU above is processed differently from the inheritance from ordinary neighboring CUs. If the candidate CU for affine motion data inheritance is in the row above the CTU, the left-bottom and right-bottom sub-block MVs in the line cache are used instead of the CPMVs for affine MVP derivation. In this way, the CPMVs are only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to a 4-parameter model. As Figure 16 shown, along the top boundary of the CTU, the left-bottom and right-bottom sub-block motion vectors of the CU are used for the affine inheritance of the CU in the bottom of the CTU. Figure 33 A diagram showing the use of motion vectors of the proposed combined method is presented. 2.25.4. Prediction Refinement Using Optical Flow for Affine Modes Compared with pixel-based motion compensation, sub-block based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a finer motion compensation granularity, Prediction Refinement by Optical Flow (PROF) is used to refine the sub-block based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after the sub-block based affine motion compensation is performed, the luma prediction samples are refined by adding the differences derived from the optical flow equations. PROF is described in the following four steps: Step 1) Sub-block based affine motion compensation is performed to generate the sub-block prediction I(i,j). Step 2) Using a 3-tap filter [-1,0,1], the spatial gradients g x (i,j) and g y (i,j) are calculated at each sample position. The gradient calculation is exactly the same as that in BDOF. g x (i, j) = (I(i+1,j) >> shift1) - (I(i-1,j) >> shift1) (2-25) g y (i, j) = (I(i,j+1) >> shift1) - (I(i,j-1) >> shift1) (2-26) shift1 is used to control the accuracy of the gradient. The sub-block (i.e., 4x4) prediction is extended by one sample on each side for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extended boundaries are copied from the nearest integer pixel positions in the reference picture. Step 3) Luma prediction refinement is calculated by the following optical flow equation. ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) (2-27) where as Figure 34 shown, Δv(i,j) is the sample MV calculated for the sample position (i,j), denoted as v(i,j), which is the difference from the sub-block MV of the sub-block to which the sample (i,j) belongs. Δv(i,j) is quantized in units of 1 / 32 luma sample accuracy. Figure 34 shows the sub-block MV V SB and the pixel Δv(i, j) (shown by the gray arrow). Since the affine model parameters and the sample positions relative to the sub - block center do not change from sub - block to sub - block, Δv(i,j) can be calculated for the first sub - block and reused for other sub - blocks in the same CU. Let dx(i,j) and dy(i,j) be the horizontal and vertical offsets of the sample position (i,j) to the sub - block center (x SB ,y SB ). The Δv(x,y) can be derived by the following equation. To maintain accuracy, the center (x SB ,y SB ) of the sub - block is calculated as ((W SB - 1) / 2,(H SB - 1) / 2), where W SB and H SB are the width and height of the sub - block respectively. For the 4 - parameter affine model, For the 6 - parameter affine model, where (v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y ) are the motion vectors of the control points at the upper - left, upper - right and lower - left corners, and w and h are the width and height of the CU. Step 4) Finally, the luminance prediction refinement ΔI(i,j) is added to the sub - block prediction I(i,j). The final prediction I’ is generated by the following equation. I′(i,j) = I(i,j)+ΔI(i,j) (2 - 32) PROF is not applicable to affine - coded CUs in two cases: 1) all control - point MVs are the same, which means the CU has only translational motion; 2) the affine motion parameters are greater than the specified limit, because the sub - block - based affine MC is degraded to CU - based MC to avoid large memory access bandwidth requirements. A fast encoding method is applied to reduce the encoding complexity of affine motion estimation using PROF. In the following two cases, PROF is not applied during the affine motion estimation stage: a) If the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) If the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low-latency picture, then PROF is not applied because the improvement introduced by PROF is small in this case. In this way, the affine motion estimation using PROF can be accelerated. 2.26. Sub-block based Temporal Motion Vector Prediction (SbTMVP) VVC supports the sub-block based Temporal Motion Vector Prediction (SbTMVP) method. Similar to the Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and the Merge mode of the CUs in the current picture. The same co-located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects: – TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; – While TMVP obtains the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right block or the center block relative to the current CU), SbTMVP applies a motion displacement before obtaining the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process is shown in Figure 18A SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 18A is checked. If A1 has a motion vector using the co-located picture as its reference picture, then that motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current picture) to obtain the sub-CU level motion information (motion vector and reference index) from the co-located picture as shown in Figure 18B Figure 18B ​The example in [ ] assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample point) in the co-located picture is used to derive the motion information of the sub-CU. After identifying the motion information of the co-located sub-CU, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. Figure 35A and Figure 35B shows the SbTMVP process in VVC. Figure 35A shows the spatial neighboring blocks used by ATVMP. Figure 35B shows the derivation of the sub-CU motion field by applying the motion displacement from the spatial neighboring CU and scaling the motion information from the corresponding co-located sub-CU. In VVC, a combined sub-block-based Merge list containing both SbTMVP candidates and affine Merge candidates is used for signaling of the sub-block-based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block-based Merge candidates, followed by the affine Merge candidates. The size of the sub-block-based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8x8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs with width and height both greater than or equal to 8. The encoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P-slice or B-slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.27. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in units of quarter luminance samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be encoded and decoded with different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AVMP mode), the MVD of the current CU can be adaptively selected as follows: – Normal AMVP mode: quarter-luma samples, half-luma samples, integer-luma samples, or quadruple-luma samples. – Affine AMVP mode: quarter-luma samples, integer-luma samples, or 1 / 16-luma samples. If the current CU has at least one non-zero MVD component, the MVD resolution indication at the CU level is signaled conditionally. If all MVD components (i.e., both the horizontal MVD and the vertical MVD for reference list L0 and reference list L1) are zero, quarter-luma sample MVD resolution is assumed. For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate that half-luma samples or other MVD precision (integer or quadruple-luma samples) is used for normal AMVP CUs. In the case of half-luma samples, the half-luma sample positions use a 6-tap interpolation filter instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer-luma samples or quadruple-luma sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer-luma sample MVD precision or 1 / 16-luma sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luma samples, half-luma samples, integer-luma samples, or quadruple-luma samples), the motion vector prediction value for the CU is rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded to zero (i.e., negative motion vector prediction values are rounded to positive infinity and positive motion vector prediction values are rounded to negative infinity). The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM11, the RD check for MVD accuracy other than quarter-luminance samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for quarter-luminance sample MVD accuracy and the RD cost for integer-luminance sample MV accuracy are calculated. Then, the RD cost for integer-luminance sample MVD accuracy is compared with the RD cost for quarter-luminance sample MVD accuracy to decide whether it is necessary to further check the RD cost for four-luminance sample MVD accuracy. When the RD cost for quarter-luminance sample MVD accuracy is much smaller than the RD cost for integer-luminance sample MVD accuracy, the RD check for four-luminance sample MVD accuracy is skipped. Then, if the RD cost for integer-luminance sample MVD accuracy is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for half-luminance sample MVD accuracy is skipped. For the affine AMVP mode, if the rate-distortion cost of the affine inter prediction mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, Merge / skip mode, quarter-luminance sample MVD accuracy normal AMVP mode, and quarter-luminance sample MVD accuracy affine AMVP mode, the 1 / 16-luminance sample MV accuracy and 1-pixel MV accuracy affine inter prediction modes are not checked. In addition, in the 1 / 16-luminance sample and quarter-luminance sample MV accuracy affine inter prediction modes, the affine parameters obtained in the quarter-luminance sample MV accuracy affine inter prediction mode are used as the starting search points. 2.28. Bi-directional prediction with CU-level weights (BCW) In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred ((8 - w)*P0 + w*P1 + 4) >> 3 (2 - 33) Five weights are allowed in weighted-average bi-directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi-directional prediction CU, the weight w is determined in one of two ways: 1) For non-Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luminance samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-delay picture, the unequal weights for 1-pixel and 4-pixel motion vector precisions are only conditionally checked. – When combined with affine, the unequal-weight affine ME is performed if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bi-prediction are the same, the unequal weights are only conditionally checked. – When specific conditions are met, the unequal weights are not searched, depending on the POC distance between the current picture and its reference picture, the coding / decoding QP, and the temporal level. The BCW weight index is decoded using a context-coded bit followed by a bypass-coded bit. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, the bypass coding signals additional bits to indicate which unequal weight is used. Weighted Prediction (WP) is a coding / decoding tool supported by the H.264 / AVC and HEVC standards for efficiently coding / decoding video content with fades. Support for WP is also added in the VVC standard. WP allows signaling of weighting parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the (multiple) weights and (multiple) offsets corresponding to the (multiple) reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For a Merge CU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.29. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is a codec tool used to address the problem of local illumination changes between the current picture and its temporal reference picture. LIC is based on a linear model, where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples of the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x,y) = α·P r (x + v x , y + v y ) + β where P(x,y) is the predicted signal of the current block at coordinates (x,y); P r (x + v x , y + v y ) is the reference block pointed to by the motion vector (v x , v y ); α and β are the corresponding scaling factor and offset applied to the reference block. Figure 19 The LIC process is shown. In Figure 19 , when LIC is applied to a block, the Least Mean Square Error (LMSE) method is adopted. By minimizing the difference between the neighboring samples of the current block (i.e., the template T in Figure 19 ) and its corresponding reference samples in the temporal reference picture (i.e., T0 or T1 in Figure 19 ), the values of the LIC parameters (i.e., α and β) are derived. Additionally, to reduce the computational complexity, both the template samples and the reference template samples are downsampled (adaptive downsampling) to derive the LIC parameters, i.e., only the shaded samples in Figure 19 are used to derive α and β. To improve the codec performance, as shown in Figure 20 , the short side is not downsampled. Figure 36 Local Illumination Compensation is shown. Figure 37 It shows that downsampling for the short side is not performed. 2.30. Decoder-Side Motion Vector Refinement (DMVR) To improve the accuracy of the Merge mode MV, decoder-side motion vector refinement based on Bilateral Matching (BM) is applied in VVC. In the bi-predictive operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and list L1. As shown in Figure 21 , the Sum of Absolute Differences (SAD) between two blocks based on each MV candidate (e.g., MV0’ and MV1’) around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predictive signal. Figure 38Decoder-side motion vector refinement is shown. In VVC, the application of DMVR is restricted and is only applied to CUs encoded and decoded using the following modes and functions: – CU-level Merge mode with bi-predictive MVs. – For the current picture, one reference picture is past and the other reference picture is future. – The distances (i.e., POC differences) from the two reference pictures to the current picture are the same. – Both reference pictures are short-term reference pictures. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current block. – The CIIP mode is not used for the current block. The refined MVs derived through the DMVR process are used to generate inter-predicted samples and are also used for temporal motion vector prediction for future picture coding. The original MVs are used for the deblocking process and are also used for spatial motion vector prediction for future CU coding. Additional features of DMVR are mentioned in the following sub-articles. 2.30.1. Search scheme In DVMR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point examined by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0 + MV_offset (2-34) MV1′ = MV1 - MV_offset (2-35) where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample phase of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search phase. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between the reference blocks pointed to by the initial MV candidates reduces the SAD value by 1 / 4. After integer sample point search, fractional sample point refinement is performed. To save computational complexity, fractional sample point refinement is derived using the parametric error surface equation instead of performing additional search using SAD comparison. Fractional sample point refinement is conditionally invoked based on the output of the integer sample point search phase. When the integer sample point search phase ends at the center with the minimum SAD in the first iteration or the second iteration search, fractional sample point refinement is further applied. In sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C(2 - 36) where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) (2 - 37) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0))) (2 - 38) The values of x min and y min are automatically constrained between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined incremental MV. 2.30.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are centered around the initial fractional pixel MV with integer sample offsets, so samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, within a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples than the normal MC process, the samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV will be filled with samples from those available samples. 2.30.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it is further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16. 2.31. Multi-pass Decoder-side Motion Vector Refinement In this paper, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the coded / decoded block. In the second pass, BM is applied to each 16x16 sub-block within the coded / decoded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction. 2.31.1. First Pass – Block-based Bilateral Matching MV Refinement In the first pass, refined MVs are derived by applying BM to the coded / decoded block. Similar to decoder-side motion vector refinement (DMVR), refined MVs are searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between two reference blocks in L0 and L1. BM performs a local search to derive the integer sample accuracy intDeltaMV and the half-pixel sample accuracy halfDeltaMV. The local search applies a 3×3 square search pattern to loop within the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the distorted DC effect between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV or halfDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. The current fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as: · MV0_pass1 = MV0 + deltaMV; · MV1_pass1 = MV1 – deltaMV. 2.31.2. Second Pass – Sub-block Based Bilateral Matching MV Refinement In the second pass, the refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) for the reference picture lists L0 and L1 obtained through the first pass. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM performs a full search to derive the integer sample accuracy intDeltaMV. The full search has a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference sub-blocks, as: bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into no more than 5 diamond search areas, as Figure 22As shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond-shaped area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order from the upper left corner of the area to the lower right corner. When the minimum bilCost within the current search area is less than a threshold equal to sbW * sbH, the full-pixel full search is terminated; otherwise, the full-pixel full search continues to the next search area until all search points have been examined. Figure 39 The diamond-shaped areas in the search area are shown. BM performs a local search to derive the half-sample accuracy halfDeltaMv. The search pattern and cost function are the same as those defined in Section 2.9.1. Existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). Then, the refined MV for the second pass is derived as: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2); · MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2). 2.31.3. Third Pass – Sub-Block Based Bidirectional Optical Flow MVV Refinement In the third pass, refined MVs are derived by applying BDOF to 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied starting from the refined MVs of the parent and child blocks in the second pass to derive scaled Vx and Vy without clipping. The derived bioMv (Vx, Vy) is rounded to 1 / 16 sample accuracy and clipped between -32 and 32. The refined MVs for the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as: · MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv; · MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv. 2.32. Sample-Based BDOF In sample-based BDOF, the motion refinement (Vx, Vy) is not derived based on blocks but is performed for each sample. The coding / decoding block is partitioned into 8×8 sub - blocks. For each sub - block, whether to apply BDOF is determined by checking the SAD between two reference sub - blocks against a threshold. If it is decided to apply BDOF to the sub - block, for each sample point in the sub - block, a sliding 5x5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the central sample point of the window. 2.33. Extended Merge Prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: 1) Spatial MVPs from spatially neighboring CUs. 2) Temporal MVPs from co - located CUs. 3) History - based MVPs from the FIFO table. 4) Paired - average MVPs. 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU code in the Merge mode, the index of the best Merge candidate is encoded using truncated unary binary (TU). The first binary bit (bin) of the Merge index is coded / decoded using context, while bypass coding is used for the other binary bits. The derivation process for each category of Merge candidates is provided in this section. Similar to what is done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a region of a specific size. 2.33.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates at the shown positions, up to four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Only when one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or are intra - coded, is position B2 considered. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thus improving the coding / decoding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs Figure 24 linked by arrows in are considered, and a candidate is added to the list only if the corresponding candidates used for the redundancy check do not have the same motion information. Figure 40The positions of spatial merge candidates are shown. Figure 41 Candidate pairs considered for redundancy check of spatial merge candidates are shown. 2.33.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. In particular, in the derivation of the temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list to be used for the derivation of the co-located CU is explicitly signaled in the slice header. Figure 25 As shown by the dotted line in , the scaled motion vector for the temporal Merge candidate is obtained by scaling the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set equal to zero. Figure 42 An illustration of motion vector scaling for temporal Merge candidates is shown. like Figure 26 As shown, the position of the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is not available, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate. Figure 43 The candidate positions for the time domain Merge candidates C0 and C1 are shown. 2.33.3. Merge candidate derivation based on history After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of the previous codec block is stored in a table and used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, and then all HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest several HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. For spatial or temporal Merge candidates, redundancy checking is applied to the HMVP candidates. To reduce the number of redundancy checking operations, the following simplifications are introduced: The number of HMPV candidates for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table. Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the Merge candidate list construction process from HMVP is terminated. 2.33.4. Pairwise average Merge candidate derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the Merge indices of the Merge candidate list. The average motion vector is calculated separately for each reference list. If both motion vectors are available in a list, they are averaged even if the two motion vectors point to different reference pictures; if only one motion vector is available, that motion vector is directly used; if no motion vector is available, this list is kept invalid. When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum Merge candidate number is reached. 2.33.5. Merge estimation region The Merge Estimation Region (MER) allows for the independent derivation of the Merge candidate list for a CU within the same MER. Candidate blocks that are within the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. Additionally, the update process for the history-based motion vector prediction candidate list is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2parMrglevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set in the form of log2_parallel_merge_level_minus2. 2.34. New Merge Candidates 2.34.1. Non-Adjacent Merge Candidate Derivation In VVC, Figure 27 the five spatially neighboring blocks and one temporally neighboring block shown are used to derive Merge candidates. It is proposed to derive additional Merge candidates from positions non-adjacent to the current block using the same style as in VVC. To achieve this, for each search round i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated by the following formula: Offsetx = -i × gridX, Offsety = -i × gridY where Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid. Second, the width and height of the virtual block are calculated by the following formula: newWidth = i × 2 × gridX + currWidth newHeight = i × 2 × gridY + currHeight. where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block. gridX and gridY are currently set to currWidth and currHeight respectively. Figure 44Shows the VVC spatial neighboring blocks of the current block. Figure 28 Shows the relationship between the virtual block and the current block. Figure 45 Shows the illustration of the virtual block in the i-th round of search. After generating the virtual block, block A i 、B i 、C i 、D i and E i can be regarded as the VVC spatial neighboring blocks of the virtual block, and their positions are obtained using the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i 、B i 、C i 、D i and E i are the spatial neighboring blocks used in the VVC Merge mode. When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighboring blocks are used. The non-adjacent spatial Merge candidates are inserted into the Merge list after the temporal Merge candidates in the order of B1->A1->C1->D1->E1. 2.34.2.STMVP It is proposed to use three spatial Merge candidates and one temporal Merge candidate to derive the average candidate as the STMVP candidate. STMVP is inserted before the top-left spatial Merge candidate. The STMVP candidate is deduplicated together with all previous Merge candidates in the Merge list. For the spatial candidates, the first three candidates in the current Merge candidate list are used. For the temporal candidate, the same position as in VTM / HEVC at the same position is used. For the spatial candidates, the first, second, and third candidates inserted before STMVP in the current Merge candidate list are denoted as F, S, and T. The temporal candidate with the same position as in VTM / HEVC used in TMVP is denoted as Col. The motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows: 1) If the reference indices of the four Merge candidates are all valid and equal to zero in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_S + mvLX_T + mvLX_Col) >> 2 2) If the reference indices of three out of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F × 3 + mvLX_S × 3 + mvLX_Col × 2) >> 3 or mvLX = (mvLX_F × 3 + mvLX_T × 3 + mvLX_Col × 2) >> 3 or mvLX = (mvLX_S × 3 + mvLX_T × 3 + mvLX_Col × 2) >> 3. 3) If the reference indices of two out of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_Col) >> 1 or mvLX = (mvLX_S + mvLX_Col) >> 1 or mvLX = (mvLX_T + mvLX_Col) >> 1. Note: If the temporal candidate is not available, the STMVP mode is turned off. 2.34.3. Merge List Size If both non - adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8. 2.35. Geometric Partitioning Mode (GPM) In VVC, the geometric partitioning mode is supported for inter - prediction. The CU - level flag is used as a kind of Merge mode to signal the geometric partitioning mode. Other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub - block Merge mode. For each possible CU size w×h = 2 m ×2 n , where m,n ∈ {3…6} excluding 8x64 and 64x8, the geometric partitioning mode supports a total of 64 partitions. When using this mode, the CU is divided into two parts by a geometrically - located line( Figure 29)。The position of the dividing line is derived mathematically from the perspective of a specific division and an offset parameter. Each part in the geometric division of the CU is inter - frame predicted using its own motion; only unidirectional prediction is allowed for each division, that is, each part has a motion vector and a reference index. The unidirectional prediction motion constraint is applied to ensure that, like traditional bidirectional prediction, each CU only requires two motion - compensated predictions. The unidirectional prediction motion for each division is derived using the process described in 2.34.1. Figure 46 An example of GPM division grouped by the same angle is shown. If the geometric division mode is used for the current CU, the geometric division index (angle and offset) indicating the division mode of the geometric division and two Merge indices (one for each division) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization for the GPM Merge index is specified. After predicting each part of the geometric division, as in 2.34.2, a hybrid process with adaptive weights is used to adjust the sample values along the geometric division edge. This is the prediction signal for the entire CU, and the transformation process and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric division mode is stored, as described in 2.34.3. 2.35.1. Unidirectional Prediction Candidate List Construction In 2.32, the unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (X equals the parity of n) of the n - th extended Merge candidate is used as the n - th unidirectional prediction motion vector for the geometric division mode. These motion vectors are marked with "x" in Figure 30 . If the corresponding LX motion vector of the n - th extended Merge candidate does not exist, the L(1 - X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric division mode. Figure 47 An example of unidirectional prediction MV selection for the geometric division mode is shown. 2.35.2. Hybrid Along the Geometric Division Edge After predicting each part of the geometric division using the own motion of each part of the geometric division, a hybrid is applied to the two prediction signals to derive the samples around the geometric division edge. The hybrid weight for each position of the CU is derived based on the distance between the independent position and the division edge. The distance from the position (x, y) to the division edge is derived as: where i and j are the indices of the angle and offset of the geometric partitioning, which depend on the geometric partitioning index transmitted through the signal. ρ x,j and ρ y,j The sign of depends on the angle index i. The weight of each part of the geometric partitioning is derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) (2-43) partIdx depends on the angle index i. An example of the weight w0 is shown in Figure 31 is shown. Figure 48 An example generation of the hybrid weight w0 using the geometric partitioning pattern is shown. 2.35.3. Motion Field Storage for Geometric Partitioning Mode Mv1 from the first part of the geometric partitioning, Mv2 from the second part of the geometric partitioning, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded / decoded in the geometric partitioning mode. The type of motion vector stored for each independent position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx ≤ 0? (1 - partIdx) : partIdx (2-46) where motionIdx is equal to d(4x + 2, 4y + 2), which is recomputed from Equation (2-18). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, then the combined Mv from Mv1 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-predictive motion vector. Otherwise, if Mv1 and Mv2 are from the same list, then only the uni-predictive motion Mv2 is stored. 2.36. Multiple Hypothesis Prediction In multiple hypothesis prediction (MHP), on top of the inter-frame AMVP mode, the regular Merge mode, the affine Merge, and the MMVD mode, at most two additional prediction values are signaled. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 = (1 - α n+1 )pn +α n+1 h n+1 The weighting factor α is specified according to Table 8 below. Table 8 Weighting factors for MHP add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For the inter-frame AMVP mode, MHP is applied only when unequal weights in BCW are selected in the bi-prediction mode. Additional assumptions can be the Merge mode or the AMVP mode. In the Merge mode, the motion information is indicated by the Merge index, and the Merge candidate list is the same as that in the geometric partitioning mode. In the AMVP mode, the reference index, the MVP index, and the MVD are signaled. 2.37. Non-adjacent spatial candidates Non-adjacent spatial Merge candidates are inserted after the TMVP in the regular Merge candidate list. The pattern of the spatial Merge candidates is shown in Figure 32 above. The distance between the non-adjacent spatial candidates and the current coded block is based on the width and height of the current coded block. Figure 49 The spatial neighboring blocks used to derive the spatial Merge candidates are shown. 2.38. Template matching (TM) Template matching (TM) is a decoder-side MV derivation method for refining the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference picture (i.e., of the same size as the template). As Figure 33 shown, within the [-8, +8] pixel search range, a better MV is searched around the initial motion of the current CU. Template matching is adopted in this paper with two modifications: determining the search step based on the AMVR mode, and TM can be cascaded using the bilateral matching process in the Merge mode. Figure 50Performing template matching on a search area around the initial MV is shown. In the AMVP mode, one that achieves the minimum difference between the current block template and the reference block template is selected based on the template matching error to determine the MVP candidate, and then the TM only performs MV refinement on this specific MVP candidate. The TM refines this MVP candidate by using an iterative diamond search, starting from the full pixel MVD accuracy within the [-8, +8] pixel search range (or 4 pixels for the 4-pixel AMVR mode). The AMVP candidate can be further refined by using a cross search with the full pixel MVD accuracy (or 4 pixels for the 4-pixel AMVR mode), and then successively using half pixels and quarter pixels according to the AMVR mode specified in Table 9. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. Table 9 Search Patterns for AMVR and Merge Mode Utilizing AMVR In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 9, the TM can be performed all the way to the 1 / 8 pixel MVD accuracy, or skip those accuracies that exceed the half pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in the half pixel mode). In addition, when the TM mode is enabled, the template matching can operate as an independent process, or as an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether the BM can be enabled according to its enabling conditions check. 2.39. Overlapped Block Motion Compensation (OBMC) Overlapped Block Motion Compensation (OBMC) has been used in H.263 before. In JEM, different from H.263, OBMC can be turned on and off using the syntax at the CU level. When OBMC is used in JEM, OBMC is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. In addition, it also applies to the luminance component and the chrominance component. In JEM, the MC block corresponds to the coding / decoding block. When coding / decoding the CU in the sub-CU mode (including sub-CU Merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To handle the CU boundaries in a unified manner, OBMC is performed on all MC block boundaries at the sub-block level, where the sub-block size is set to be equal to 4×4, as Figure 34 shown. When OBMC is applied to the current sub - block, in addition to the current motion vector, the motion vectors of four connected neighboring sub - blocks (if available and different from the current motion vector) are also used to derive the prediction block for the current sub - block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub - block. Denote the prediction block based on the motion vector of the neighboring sub - block as P N , where N indicates the indices of the neighboring upper, lower, left, and right sub - blocks, and the prediction block based on the motion vector of the current sub - block is denoted as P C . When P N is based on the motion information of neighboring sub - blocks containing the same motion information as the current sub - block, OBMC is not performed starting from P N . Otherwise, each sample of P N is added to the same sample in P C , that is, the four rows / columns of P N are added to P C . The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . Except for small MC blocks (i.e., when the height or width of the coded / decoded block is equal to 4, or when the CU is coded / decoded using the sub - CU mode), for small MC blocks, only two rows / columns of P N are added to P C . In this case, the weighting factors {1 / 4, 1 / 8} are used for P N , while the weighting factors {3 / 4, 7 / 8} are used for P C . For P N generated based on the motion vectors of vertical (horizontal) neighboring sub - blocks, the samples in the same row (column) of P N are added to P C with the same weighting factor. Figure 51 FIG. shows a diagram of the sub - block to which OBMC is applied. In JEM, for a CU with a size less than or equal to 256 luma samples, a signal is transmitted at the CU level flag to indicate whether OBMC is applied to the current CU. For a CU with a size greater than 256 luma samples or not coded / decoded using the AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its influence is considered during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top neighboring block and the left neighboring block is used to compensate the top boundary and the left boundary of the original signal of the current CU, and then the normal motion estimation process is applied. 2.40. Multiple - transform Selection (MTS) for Kernel Transform In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of inter- and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 10 shows the basis functions of the selected DST / DCT. Table 10 Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that of the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter respectively. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applicable to luma. The MTS signaling is skipped when one of the following conditions is met: - The position of the last significant coefficient of the luma TB is less than 1 (i.e., only DC). - The last significant coefficient of the luma TB is within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 11. A unified transform selection for ISP and implicit MTS is used by eliminating the intra-mode and block-shape dependencies. If the current block is in ISP mode or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are on, only DST7 is used for the horizontal transform kernel and the vertical transform kernel. In terms of transform matrix precision, an 8-bit primary transform kernel is used. Thus, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8) all use an 8-bit primary transform kernel. Table 11 Transform and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are zeroed. Only the coefficients within the 16×16 low-frequency region are retained. Similar to HEVC, the residual of a block can be coded and decoded using the transform skip mode. To avoid redundancy in syntax coding and decoding, the transform skip flag is not signaled when the MTS_CU_flag at the CU level is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, the implicit MTS can still be enabled. 2.41. Sub-Block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, for a CU, only a sub-part of the residual block is coded. When the cu_cbf of an inter-predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is coded. For the former case, the inter MTS information is further parsed to determine the transform type of the CU. For the latter case, a part of the residual block is coded by a presumed adaptive transform while the other part of the residual block is zeroed. When SBT is used for an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, as Figure 35A and Figure 35B shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to a binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to an asymmetric binary tree (ABT) partition. In the ABT partition, only small regions contain non-zero residuals. If one dimension of a CU is 8 (in terms of luminance samples), the 1:3 / 3:1 partition along that dimension is not allowed. A CU can have up to 8 SBT modes. Position-dependent transform kernel selection is applied to the luminance transform blocks in SBT-V and SBT-H (chrominance TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are specified in Figure 35A and Figure 35B . For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Thus, the sub-block transform jointly specifies the TU slicing, cbf, and the horizontal and vertical kernel transform types of the residual block. SBT is not applied to CUs coded in the inter-intra joint mode. Figure 52Shows the SBT position, type, and transform type. 2.42. Adaptive Merge Candidate Reordering Based on Template Matching To improve the encoding and decoding efficiency, after constructing the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list according to the ascending template matching cost. It is operated in subgroups. The template matching cost is measured by the SAD (Sum of Absolute Differences) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes bidirectional prediction motion information, then as Figure 36 shown, the corresponding reference sample is the average of the corresponding reference samples in reference list 0 and the corresponding reference samples in reference list 1. If the Merge candidate contains motion information at the sub-CU level, then as Figure 37 shown, the corresponding reference sample consists of the neighboring samples of the corresponding reference sub-block. As Figure 38 shown, the sorting process is operated in subgroups. The first three Merge candidates are sorted together. The next three Merge candidates are sorted together. The template size (width of the left template or height of the upper template) is 1. The subgroup size is 3. Figure 53 Shows the neighboring samples for calculating the SAD. Figure 54 Shows the neighboring samples for calculating the SAD for motion information at the sub-CU level. Figure 55 Shows the sorting process. 2.43. Adaptive Merge Candidate List Assume the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first subgroup, and the subsequent 3 Merge candidates are taken as the second subgroup (i.e., the last subgroup). For the encoder, as Figure 39 shown, after the Merge candidate list is constructed, some Merge candidates are adaptively reordered in ascending order of the Merge candidate cost. More specifically, the template matching costs of the Merge candidates in all subgroups except the last subgroup are calculated; then, the Merge candidates in their own subgroups except the last subgroup are reordered; finally, the final Merge candidate list will be obtained. Figure 56 Shows the reordering process in the encoder. For the decoder, after the Merge candidate list is constructed, as Figure 40 shown, some / none of the Merge candidates are adaptively reordered in ascending order of the Merge candidate cost. At Figure 40Among them, the subgroup where the selected (transmitted via signal) Merge candidate is located is called the selected subgroup. Figure 57 Illustrates the reordering process in the decoder. More specifically, if the selected Merge candidate is in the last subgroup, after the selected Merge candidate is derived, the Merge candidate list construction process is terminated, no reordering is performed, and the Merge candidate list is not changed; otherwise, the process is as follows: After all the Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; calculate the template matching cost of the Merge candidates in the selected subgroup; reorder the Merge candidates in the selected subgroup; finally, a new Merge candidate list will be obtained. For both the encoder and the decoder: The template matching cost is derived as a function of T and RT, where T is the set of sample points in the template and RT is the set of reference sample points for the template. When deriving the reference sample points of the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel accuracy. The reference sample points (RT) of the template for bidirectional prediction are derived by weighted averaging the reference sample points (RT0) of the template in reference list 0 and the reference sample points (RT1) of the template in reference list 1 as follows. RT = ((8 - w) * RT0 + w * RT1 + 4) >> 3 (2 - 47) Among them, the weight (8 - w) of the reference template in reference list 0 and the weight (w) of the reference template in reference list 1 are determined by the BCW index of the Merge candidate. The BCW index equal to {0, 1, 2, 3, 4} corresponds to w equal to {-2, 3, 4, 5, 10} respectively. If the local illumination compensation (LIC) flag of the Merge candidate is true, the LIC method is used to derive the reference sample points of the template. The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT. The template size is 1. This means that the width of the left template and / or the height of the upper template is 1. If the codec mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered. If the codec mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered. 2.44. Geometric prediction mode with motion vector difference In the Geometric Prediction Mode with Motion Vector Difference (GMVD), each geometric partition in the GPM can decide whether to use GMVD. If GMVD is selected for a geometric region, the MV of that region is calculated as the sum of the MV of the Merge candidate and the MVD. All other processing remains the same as in GPM. With GMVD, the MVD is signaled in the form of a direction and distance pair. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels), and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag is equal to 1, the MVD in GMVD is also left-shifted by 2 as in MMVD. 2.45. Affine MMVD In Affine MMVD, an affine Merge candidate (referred to as the base affine Merge candidate) is selected, and the MV of the control points is further refined by the signaled MVD information. The MVD information of the MV of all control points is the same in one prediction direction. When the starting MV is a bi-predictive MV where the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture while the POC of the other reference is less than the POC of the current picture), the MV offset added to the list 0 MV component of the starting MV and the MV offset of the list 1 MV have opposite values; otherwise, when the starting MV is a bi-predictive MV where both lists point to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the MV offset added to the list 0 MV component of the starting MV and the MV offset of the list 1 MV are the same. 2.46. Adaptive Decoder-Side Motion Vector Refinement (ADMVR) In ECM-2.0, if the selected Merge candidate meets the DMVR condition, the multi-pass decoder-side motion vector refinement (DMVR) method is applied in the regular Merge mode. In the first pass, bilateral matching (BM) is applied to the coded / decoded block. In the second pass, BM is applied to each 16x16 sub-block within the coded / decoded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The adaptive decoder-side motion vector refinement method consists of two new Merge modes, which are introduced to refine the MV in only one direction (L0 or L1) of the bi-prediction of the Merge candidates that meet the DMVR condition. The multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vectors. However, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is set to zero. Similar to the regular Merge mode, the Merge candidates for the proposed Merge mode are derived from spatial neighboring coded / decoded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that meet the DMVR condition are added to the candidate list. The same Merge candidate list (i.e., ADMVR Merge list) is used by the two proposed Merge modes, and the Merge index is coded / decoded in the same way as the regular Merge mode. 2.47. Convolutional Cross-Component Model (CCCM) for Intra Prediction It is proposed to apply the convolutional cross-component model (CCCM) in a similar spirit to that done by the current CCLM mode to predict chroma samples from the reconstructed luma samples. Similar to CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower-resolution chroma grid. Similarly, similar to CCLM, there is an option to use single-model or multi-model variants of CCCM. The multi-model variant uses two models, one model is derived for samples above the average luma reference value, and the other model is derived for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples. 2.47.1. Convolutional Filter The proposed convolutional 7-tap filter consists of a 5-tap plus-shaped spatial component, a non-linear term, and a bias term. The input of the spatial 5-tap component of the filter consists of the central (C) luma sample co-located with the chroma sample to be predicted and its upper / northern (N), lower / southern (S), left / western (W), and right / eastern (E) neighbors as shown below. Figure 58 The spatial part of the convolutional filter is shown. The non-linear term P is represented as a power of 2 in the central luma sample C and is scaled to the sample value range of the content: P = (C * C + midVal) >> bitDepth. That is, it is calculated for 10-bit content as: P = (C * C + 512) >> 10. The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the middle chroma value (512 for 10-bit content). The output of the filter is computed as the convolution between the filter coefficients c i and the input values, and is clipped to the range of valid chroma samples: predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B. 2.47.2. Computation of filter coefficients The filter coefficients c i are computed by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 59 Fig. shows a reference region consisting of 6 rows of chroma samples above and to the left of the PU. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to include only the available samples. The extension of the region shown in blue is needed to support the "side samples" of the plus-shaped spatial filter and is filled when in an unavailable region. Figure 59 Fig. shows the reference region (and its filling) for deriving the filter coefficients. The MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL-factorized, and the final filter coefficients are computed using back substitution. This process generally follows the computation of the ALF filter coefficients in ECM. However, LDL factorization is chosen instead of Cholesky factorization to avoid using square root operations. The proposed method uses only integer operations. 2.47.3. Bitstream signaling The use of this mode is signaled using a CABAC-encoded / decoded PU-level flag. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is signaled only if the intra prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-model CCCM). 2.47.4. Encoder operation The encoder performs two new RD checks in the chroma prediction mode loop, one for checking the single-model CCCM mode and one for checking the multi-model CCCM mode. 2.48. Chroma fusion 2.48.1. Chroma fusion using DM, default, and MMLM modes In ECM 6.0, there is a chroma fusion method, that is, the DM mode and four default modes can be fused with the MMLM_LT mode, as follows: pred = (w0 * pred0 + w1 * pred1 + (1 << (shift - 1))) >> shift where pred0 is the predicted value obtained by applying the non-LM mode, pred1 is the predicted value obtained by applying the MMLM_LT mode, and pred is the final predicted value of the current chroma block. The preset weight combination {w0, w1} is equal to {1, 3}, {3, 1}, or {2, 2}. 2.48.2. Chroma Fusion Using LDL Decomposition A new chroma fusion method is proposed as follows: pred C (i, j) = α0 · rec′ L (i, j) + α1 · pred′ C (i, j) + α2 · midValue where pred C (i, j) represents the final predicted chroma sample, rec′ L (i, j) represents the reconstructed luminance sample. pred′ C (i, j) represents the predicted chroma sample obtained by applying the non-LM mode. For 10-bit content, midValue is set to be equal to 512. The model parameters α0, α1, and α2 are derived from the same adjacent samples (two rows) based on the LDL decomposition method used in CCCM. As shown in Table 2-12, there are three chroma fusion methods. Table 2-12 - Chroma Fusion Modes in the Proposed Method Value Name 0 Disable 1 Preset weight 2 Single - linear model 3 Multi - linear model 3. Problems 1. In the current design of the codec, the chroma fusion method can only fuse using two weights. However, this way may limit the compression efficiency because these two weights are not always beneficial for compression, especially when the blocks or frames have different signal and texture distributions. 2. In the current design of the codec, the chroma fusion method can use a fixed median for fusion. However, for different blocks or frames with different signal and texture distributions, the fixed median may not be the optimal bias value. 4. Detailed Solutions The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow way. In addition, these solutions can be combined in any way. In the present disclosure, the term "block" may represent a coding block (CB), or a coding unit (CU), or a prediction block (PB), or a prediction unit (PU), or a transform block (TB), or a transform unit (TU), or a coding tree block (CTB), or a coding tree unit (CTU), or a rectangular region of samples / pixels. In the following discussion, SatShift(x, n) is defined as: Shift(x,n) is defined as Shift(x,n) = (x + offset0) >> n. In one example, offset0 and / or offset1 are set to (1 << n) >> 1 or (1 << (n - 1)). In another example, offset0 and / or offset1 are set to 0. In another example, offset0 = offset1 = ((1 << n) >> 1) - 1 or ((1 << (n - 1))) - 1. Clip3(min, max, x) is defined as: In the following discussion, "picture", "image", and "frame" may have the same meaning. Chrominance fusion 1. Instead of using a set of predefined weights for chroma fusion, multiple sets of weights can be used for chroma fusion, where more than one set of predefined weights is used, or at least one set of derived weights of the single / multi-linear model defined in the background is not used. Represent chroma fusion as F(), such as B = F(p0, p1,..., p n ), where p n is the nth prediction method. a. In one example, more than one set of predefined weights can be used. i. In one example, whether to use the predefined weights and / or which set of predefined weights to use can be signaled in the bitstream, or derived using coding information. ii. In one example, the number of sets can depend on the following coding information: 1) Block dimension, and / or block size, and / or block depth; 2) Slice / picture type and / or partition tree type (single tree or dual tree, or local dual tree); 3) Temporal layer identifier; 4) QP. b. In one example, a chroma fusion may refer to mixing a set of prediction values (p0, p1,..., pn )。 i. In one example, B = F(p i ), where p i is the predicted value of the i-th prediction method, and w i is the fusion weight corresponding to the predicted value. 1) In one example, the weight w i is a fixed predefined value. 2) In one example, the weight w i is a value that varies depending on the context of the codec. ii. In one example, B = F(p i ), where p i is the predicted value of the i-th prediction method, w i is the fusion weight corresponding to the predicted value, b is the offset, and n is the number of right shifts. c. In one example, a prediction fusion can refer to a non-linear method of mixing a set of predicted values (p0, p1, …, p n ). i. In one example, the non-linear fusion method can employ a convolutional neural network. 1) In one example, C(p i ) = conv2(w i , p i ) + b i (3) where p i is the predicted value of the i-th prediction method, w i is the weight of the convolutional layer, and b i is the bias of the convolutional layer. B(p i ) = M(C(p1), C(p2), …, C(p i )) (4) where M is the fusion method for i predicted values. R(p i ) is the rectified linear unit (ReLU), which is an activation function. d. The p i disclosed above can be the sample points of the current component. i. The p i disclosed above can be the sample points of different components. ii. The pi It can be a sample after filtering or rescaling. e. In one example, the total number n of predicted values in (1) is equal to 2. i. In one example, the two predicted values (p1, p2) come from the linear model (LM) mode and the non-LM mode. F(p i ) = w0p0 + w1p1 1) In one example, (w0, w1) are fixed predefined values. 2) In one example, (w0, w1) are values that vary depending on the context of the codec. f. In one example, one or more sets of weights can be signaled in the bitstream. i. In one example, one or more sets of weights can be signaled at the block level / sequence level / group of pictures level / picture level / strip level / slice group level, such as in the coding / decoding structures of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. ii. In one example, a first syntax element can be signaled to indicate whether chroma fusion with multiple weights is applied. iii. In one example, a second syntax element can be signaled to indicate which set of weights among the weights is applied or not applied. 1) In one example, the second syntax element can be signaled only when the first syntax indicates that chroma fusion with multiple weights is applied. iv. In one example, a single SE can be signaled, and the single SE indicates whether chroma fusion is used and which set of weights is used. v. The (multiple) syntax elements (SE) can be coded using fixed-length coding, exponential-Golomb coding, truncation (unary) coding, etc. vi. The (multiple) SE can be coded using at least one context in arithmetic coding. vii. The (multiple) SE can be bypass-coded. viii. The (multiple) SE can be signaled only when chroma fusion with multiple weights is allowed to be used. ix. The (multiple) SE can be signaled in the SEI or VUI message. x. In one example, one or more sets of weights can be signaled in a predictive manner. 1) In one example, the weights within a set can be signaled in a predictive manner. xi. In one example, multiple sets of weights for the first video unit may not be transmitted via a signal. Instead, one or more sets of weights transmitted via a signal for the second video unit may be used, where the second video unit is encoded / decoded before the first video unit. g. In one example, one or more sets of derived weights may be used. i. In one example, the derived weights may be obtained using the reconstructed luma samples, and / or neighboring reconstructed luma / chroma samples, and / or predicted chroma samples. ii. In one example, the derived weights may be obtained using the least mean square (LMS) method (e.g., used in CCLM / MMLM), or the LDL decomposition method (e.g., used in CCCM), or the Cholesky decomposition method (e.g., used in ALF). h. In one example, multiple sets of weights may depend on the color format and / or color component. i. In one example, multiple sets of weights may be the same for two chroma components. 1) Alternatively, multiple sets of weights may be different for two chroma components. ii. In one example, how to obtain the derived weights may be different for different color formats. 1) In one example, the downsampled reconstructed luma samples may be used for 4:2:0 / 4:2:2 color formats, and the non-downsampled reconstructed luma samples may be used for 4:4:4 color formats. i. Multiple weights in chroma fusion may be different in terms of encoding / decoding information, context, and background. i. In one example, the weights may depend on the sub-picture, slice, strip, or block size of CTU / CU / PU / TU / CTB / CB / PB / TB. 1) In one example, for non-LM intra prediction mode, when the block size is 256×256, w i is p, and when the block size is 128×128, w i is q. ii. In one example, the weights may depend on the video content of the video unit. 1) In one example, for LM intra prediction mode, when the video content is a natural sequence, w i is p, and when the video content is a screen content (SCC), w i is q. j. In one example, whether to apply the chroma fusion method and / or how to apply the chroma fusion method and its weights can be signaled using at least one syntax element (SE). i. In one example, the syntax element can be signaled in the SPS / PPS / APS / picture header / strip header, etc. ii. In one example, a first SE can be signaled to indicate whether chroma fusion with multiple weights is applied. iii. In one example, a set of SEs can be signaled to indicate how the weights are predefined. 1) In one example, each element in a set of SEs indicates a different w i 。 k. In one example, whether to use and / or how to use one or more sets of weights in the above methods can be signaled in the bitstream or derived using codec information. i. In one example, one or more group indices indicating which set of weights to use can be signaled in the bitstream. 1) In one example, a single group index can be used to indicate a set of weights used for two chroma components. 2) In one example, the indication of a set of weights for two chroma components can be signaled separately. ii. In one example, the group index can be derived using codec information. 1) In one example, a method based on template matching can be used. 2) In one example, the group indices for two chroma components can be derived together. 3) In one example, the group indices for two chroma components can be derived separately. 2. Instead of using a constant median depending on the bit depth to derive a linear model for chroma fusion, it is proposed that an adaptive median can be used to derive a model for chroma fusion. a. In one example, the model can be linear or non - linear. b. In one example, the adaptive median can be derived using codec information. i. In one example, the codec information can refer to the reconstructed luma samples, and / or neighboring reconstructed luma samples, and / or neighboring reconstructed chroma samples, and / or predicted chroma samples. ii. In one example, the adaptive median can be derived from codec tools such as LMCS or LIC. iii. The median can be derived from different coded signals using different calculation methods. 1) In one example, the median value may be the average of one or more types of signals in a block region. 2) In one example, the median value may be the median of one or more types of signals in a block region. 3) In one example, the type of signal may be a template value. 4) In one example, the type of signal may be a prediction of luminance / chrominance. 5) In one example, the type of signal may be a reconstruction of luminance. iv. The median value may be derived from different methods depending on coding / decoding information, coding / decoding context, and coding / decoding background. 1) In one example, the median value may be derived based on coding / decoding position information. a) In one example, the median value may be calculated from regions of the left block and the upper block of the current block. b) In one example, the median value may be calculated from the region of the left block of the current block. c) In one example, the median value may be calculated from the region of the upper block of the current block. c. In one example, the adaptive median value may be signaled in the bitstream. i. In one example, the adaptive median value may be signaled at the block level / sequence level / group of pictures level / picture level / strip level / slice group level, such as in the coding / decoding structures of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. ii. In one example, a first syntax element may be signaled to indicate whether chrominance fusion with a median value is applied. iii. In one example, a second syntax element may be signaled to indicate which level of the median value is applied or not applied. 1) In one example, the second syntax element may be signaled only when the first syntax indicates that chrominance fusion with a median value is applied. iv. In one example, a single SE may be signaled, and the single SE indicates whether chrominance fusion is used and which level of the median value is used. v. The (multiple) SEs may be coded using fixed-length coding, exponential-Golomb coding, truncation (unary) coding, etc. vi. The (multiple) SEs may be coded using at least one context in arithmetic coding. vii. The (multiple) SEs may be bypass-coded. viii. The (multiple) SEs can be signaled only when chroma fusion with a median value is allowed to be used. ix. The (multiple) SEs can be signaled in an SEI or VUI message. d. In one example, the same median value can be used for two chroma components. i. Alternatively, two separate median values can be used for two chroma components. e. There can be multiple median values applied to the chroma fusion method. i. In one example, whether to apply the chroma fusion method and / or how to apply the chroma fusion method and its multiple median values can be signaled using at least one syntax element (SE). 1) In one example, the syntax element can be signaled in the SPS / PPS / APS / picture header / strip header / CTU / CU, etc. 2) In one example, a first syntax element can be signaled to indicate whether chroma fusion with multiple median values is applied. 3) In one example, a second syntax element can be signaled to indicate which set of median values is applied or not applied. a) In one example, the second syntax element can be signaled only when the first syntax indicates that chroma fusion with multiple median values is applied. 4) In one example, a single SE can be signaled, and the single SE indicates whether chroma fusion is used and which set of median values is used. 5) The (multiple) SEs can be encoded and decoded using fixed-length coding / decoding, exponential-Golomb coding / decoding, truncation (unary) coding / decoding, etc. 6) The (multiple) SEs can be encoded and decoded using at least one context in arithmetic coding. 7) The (multiple) SEs can be bypass-coded. 8) The (multiple) SEs can be signaled only when chroma fusion with multiple weights is allowed to be used. 9) The (multiple) SEs can be signaled in an SEI or VUI message. 3. In addition to using the reconstructed luma samples and the predicted chroma samples derived from non-LM modes for chroma fusion, gradients, and / or non-downsampled / downsampled luma samples, neighboring samples are also proposed for chroma fusion. a. In one example, a linear model or a non-linear model can be used for chroma fusion. b. In one example, neighboring reconstructed luma samples can be used for chroma fusion. c. In one example, adjacent reconstructed chrominance samples can be used for chrominance fusion. d. In one example, position information can be used. e. In one example, gradients can be used in the model for chrominance fusion. i. In one example, gradients can be calculated using the reconstructed luma samples. 1) In one example, non-downsampled luma samples can be used. 2) In one example, downsampled luma samples can be used. ii. In one example, gradients can be calculated using the predicted chrominance samples. iii. In one example, how to calculate the gradients can depend on the color format and / or color components. iv. In one example, different types of gradient information can be applied to chrominance fusion. 1) In one example, chrominance fusion can use only the horizontal gradient Gx. 2) In one example, chrominance fusion can use only the vertical gradient Gy. 3) In one example, chrominance fusion can use both the horizontal gradient G x and the vertical gradient G y both. v. In one example, different gradient calculation methods can be applied to chrominance fusion. 1) In one example, the linear sum and linear difference of the signals on the horizontal axis can be used to calculate G x , while the linear sum and linear difference of the signals on the vertical axis can be used to calculate G y . 2) In one example, the multiplication and division of the signals on the horizontal axis can be used to calculate G x , while the multiplication and division of the signals on the vertical axis can be used to calculate G y . 3) In one example, the derivation method of the signals on the horizontal axis can be used to calculate G x , while the derivation method of the signals on the vertical axis can be used to calculate G y . vi. In one example, different gradient calculation shapes can be applied to chrominance fusion. 1) In one example, a square block covering the current sample can be used to calculate the gradients G x and G y . a) In one example, as Figure 60 shown, G x and Gy It can be derived from the adjacent luminances using the following equation. G x = (2W + NW + SW) – (2E + NE + SE) G y = (2N + NW + NE) – (2S + SW + SE) 2) In one example, the diamond block covering the current sample point can be used to calculate the gradients G x and G y . a) In one example, as Figure 61 shown, G x and G y can be derived from the adjacent luminances using the following equation. G x = (2W + NW + SW + W f ) – (2E + NE + SE + E f ) G y = (2N + NW + NE + N f ) – (2S + SW + SE + S f ) vii. In one example, flexible modification of the terms in gradient calculation can be applied to chroma fusion. 1) In one example, when calculating the gradients G x and G y , some terms can be replaced. 2) In one example, when calculating the gradients G x and G y , some additional terms can be added. 3) In one example, when calculating the gradients G x and G y , some terms can be combined. f. In one example, multiple downsampling filters can be applied to chroma fusion. i. In one example, chroma fusion can use a 6-tap filter to downsample luminance samples. 1) In one example, chroma fusion can use a 6-tap filter with coefficients 1 - 2 - 1 - 1 - 2 - 1 to downsample luminance samples. ii. In one example, chroma fusion can use a 3-tap filter to downsample luminance samples. 1) In one example, chroma fusion can use a 3-tap filter with coefficients 1 - 2 - 1 to downsample luminance samples. iii. In one example, whether to apply the chroma fusion method and / or how to apply the chroma fusion method and its downsampling filter can be signaled using at least one syntax element (SE). 1) In one example, the syntax element can be signaled in the SPS / PPS / APS / picture header / strip header / CTU / CU, etc. 2) In one example, a first syntax element can be signaled to indicate whether chroma fusion with a downsampling filter is applied. 3) In one example, a second syntax element can be signaled to indicate which set of the downsampling filter is applied or not applied. a) In one example, the second syntax element can be signaled only if the first syntax indicates that chroma fusion with a downsampling filter is applied. 4) In one example, a single SE can be signaled, and the single SE indicates whether chroma fusion is used and which set of the downsampling filter is used. 5) The (multiple) SEs can be coded using fixed-length coding / decoding, exponential-Golomb coding / decoding, truncation (unary) coding / decoding, etc. 6) The (multiple) SEs can be coded using at least one context in arithmetic coding. 7) The (multiple) SEs can be bypass-coded. 8) The (multiple) SEs can be signaled only if chroma fusion with multiple weights is allowed to be used. 9) The (multiple) SEs can be signaled in the SEI or VUI message. g. In one example, a non-downsampling filter can be applied to chroma fusion depending on the coding information, coding context, or coding background. i. In one example, a non-downsampling filter can be applied to chroma fusion according to the video content. 1) In one example, a non-downsampling filter can be applied to chroma fusion in the video units of SCC. ii. In one example, multiplication and division of the signals on the horizontal axis can be used to calculate G x , and multiplication and division of the signals on the vertical axis can be used to calculate G y . iii. In one example, a derivation method of the signals on the horizontal axis can be used to calculate G x , and a derivation method of the signals on the vertical axis can be used to calculate G y . h. In one example, whether to apply the chroma fusion method and / or how to apply the chroma fusion method and its downsampling filter can be signaled using at least one high-level syntax element (HLS). i. In one example, the syntax element can be signaled in the SPS / PPS / APS / picture header / strip header, etc. ii. In one example, the first HLS element can be signaled to indicate whether chroma fusion with a downsampling filter is applied. iii. In one example, a set of HLS elements can be signaled to indicate which downsampling filters and non-downsampling filters are selected. 1) In one example, each element in a set of HLS indicates a different downsampling filter f i 。 4. When decompressing and compressing a video unit using a prediction tool, the above method can also be applied to the fusion method for the luminance component in both the decoder and the encoder. General aspects 5. Whether to apply and / or how to apply the method disclosed above can be signaled at the block level / sequence level / group of pictures level / picture level / strip level / slice group level, such as in the coding / decoding structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 6. Whether to apply and / or how to apply the method disclosed above can depend on the coded / decoded information, such as block size, color format, single / double tree splitting, color component, strip type / picture type. 7. The proposed method disclosed herein can be used in other coding / decoding tools that require chroma fusion. 5. Embodiments

[0124] More details of embodiments of the present disclosure will be described below, which are related to component fusion for video coding / decoding. The embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be construed in a narrow sense. In addition, these embodiments can be applied individually or combined in any way.

[0125] As used herein, the term "block" may represent a color component, a sub-picture, a picture, a strip, a slice, a coding tree unit (CTU), a CTU row, a group of CTUs, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), a sub-block of a video block, a sub-region within a video block, a video processing unit including a plurality of samples / pixels, etc. The block may be rectangular or non-rectangular.

[0126] Figure 62 FIG. 6200 is a flow chart of a method for video processing according to some embodiments of the present disclosure. Method 6200 may be implemented during the conversion between a current video block of a video and a bitstream of the video. As shown in Figure 62 FIG. 6200 begins at 6202, where multiple sets of weights are obtained. In some embodiments, the multiple sets of weights may be predetermined. Alternatively, at least one set of weights of the multiple sets of weights may be determined based on a non-linear model.

[0127] At 6204, a target prediction for a first color component of the current video block is determined based on the multiple sets of weights and multiple candidate predictions for the first color component. For example, the target prediction for the color component may be generated by fusing the multiple candidate predictions with one or more sets of weights of the multiple sets of weights. This process may also be referred to as "prediction fusion" or "component fusion". As used herein, the term "fusion" may refer to mixing or combining more than one prediction to obtain a single prediction, and the term "fusion" may be used interchangeably with the terms "mixing" and "integration".

[0128] In some embodiments, the multiple candidate predictions may be determined based on different prediction schemes, such as a non-LM mode, an MMLM-LM mode, etc. In some embodiments, the first color component may be a chrominance component, such as a chrominance red (Cr) component or a chrominance blue (Cb) component. In this case, the above prediction fusion may also be referred to as chrominance fusion. Alternatively, the first color component may be a luminance component. In this case, the above prediction fusion may also be referred to as luminance fusion.

[0129] At 6206, the conversion is performed based on the target prediction. In some embodiments, the conversion may include encoding the current video block into the bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream. It should be understood that the above description and / or examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.

[0130] In view of the above, the prediction fusion for color components is performed based on multiple sets of weights. Compared with the conventional solutions that only use a single set of predetermined weights, the proposed method can advantageously improve the coding and decoding quality, especially for blocks or frames with different signal and texture distributions.

[0131] In some embodiments, the multiple sets of weights may be predetermined. In one example, information related to at least one of the following may be indicated in the bitstream: whether to use multiple sets of weights, or a set of weights used among the multiple sets of weights. In another example, the information may be determined based on the coding and decoding information of the current video block or the coding and decoding information of the blocks coded before the current video block. Additionally or alternatively, the number of sets in the multiple sets of weights may depend on the coding and decoding information of the current video block or the coding and decoding information of the blocks coded before the current video block.

[0132] By way of example and not limitation, the coding and decoding information may include block dimension, block size, block depth, slice type, picture type, segmentation tree type, temporal layer identifier, and / or quantization parameter (QP).

[0133] In some embodiments, the target prediction may be determined based on a weighted sum of multiple candidate predictions. For example, the target prediction may be determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, w i represents the weight for the i-th candidate prediction, n represents one less than the number of candidate predictions among the multiple candidate predictions, and midvalue represents an offset.

[0134] Alternatively, the target prediction may be determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, w i represents the weight for the i-th candidate prediction, n represents one less than the number of candidate predictions among the multiple candidate predictions, b represents an offset, and k represents the number of right shifts. In one example, the value of the weight w i may be predefined. Alternatively, the value of the weight w i may depend on the context of the codec used to code and decode the current video block.

[0135] In some additional embodiments, the target prediction may be determined based on a non-linear fusion scheme for fusing multiple candidate predictions. For example, a convolutional neural network may be used in the non-linear fusion scheme.

[0136] For example, the target prediction can be determined as follows: C(p i ) = conv2(w i , p i ) + b i , i = 0, 1, …, n B = M(C(p1), C(p2), …, C(p n )) where B represents the target prediction, M() represents the fusion scheme, p i represents the i-th candidate prediction among multiple candidate predictions, C(p i ) represents the intermediate result corresponding to the i-th candidate prediction, conv2() represents a two-dimensional convolution, w i represents the weight of the convolutional layer for the i-th candidate prediction, b i is the bias of the convolutional layer, and n represents the number of candidate predictions among multiple candidate predictions minus one.

[0137] In some embodiments, a rectified linear unit (ReLU) can be used as the activation function for the convolutional neural network. By way of example and not limitation, the rectified linear unit can be as follows: where R(p i ) represents the rectified linear unit corresponding to the i-th candidate prediction.

[0138] In some embodiments, one of the multiple candidate predictions can include a prediction of samples for a first color component, a prediction of samples for a second color component different from the first color component, or a prediction of samples after filtering or rescaling. For example, if the first color component is a chrominance component, the second color component can be a luminance component. If the first color component is the Cr component, the second color component can be the Cb component or the luminance component. If the first color component is the Cb component, the second color component can be the Cr component or the luminance component. If the first color component is the luminance component, the second color component can be the chrominance component.

[0139] In some embodiments, the number of candidate predictions among multiple candidate predictions can be 2. In this case, one of the multiple candidate predictions can be determined based on a linear model (LM) mode, and the other candidate prediction among multiple candidate predictions can be determined based on a non-LM mode. For example, the target prediction can be determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among multiple candidate predictions, and w iRepresents the weight for the i-th candidate prediction. For example, the weight w i value can be predefined. Alternatively, the weight w i value can depend on the context of the codec used to encode and decode the current video block.

[0140] In some embodiments, at least one set of weights among multiple sets of weights can be indicated in the bitstream. In one example, at least one set of weights can be indicated at the block level, sequence level, group of pictures level, picture level, slice level, and / or slice group level. In another example, at least one set of weights can be indicated in one of the following: the coding structure of a coding tree unit (CTU), the coding structure of a coding unit (CU), the coding structure of a transform unit (TU), the coding structure of a prediction unit (PU), the coding structure of a coding tree block (CTB), the coding structure of a coding block (CB), the coding structure of a transform block (TB), the coding structure of a prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.

[0141] In some embodiments, a first syntax element indicating whether prediction fusion with multiple weights is applied can be included in the bitstream. Additionally or alternatively, a second syntax element indicating a set of weights among multiple sets of weights that is applied can be included in the bitstream. In some alternative embodiments, if the first syntax element indicates that prediction fusion with multiple weights is applied, then the second syntax element is included in the bitstream.

[0142] In some further embodiments, the bitstream can include a single syntax element indicating whether prediction fusion with multiple weights is applied and a set of weights among multiple sets of weights that is applied.

[0143] In some embodiments, the syntax element can be coded using fixed-length coding, exponential Golomb (EG) coding, truncated unary coding, and / or unary coding. Alternatively, the syntax element can be coded in arithmetic coding using at least one context, or the syntax element can be bypassed.

[0144] In some embodiments, if prediction fusion with multiple weights is allowed to be used, the syntax element associated with the prediction fusion can be indicated in the bitstream. In some embodiments, the syntax element can be included in a supplementary enhancement information (SEI) message or a video usability information (VUI) message.

[0145] In some embodiments, at least one set of weights may be indicated in a predictive manner. Additionally, the weights in at least one set of weights may be indicated in a predictive manner. In some alternative embodiments, for a current video block, multiple sets of weights may not be signaled, and multiple sets of weights may be reused from blocks that were coded before the current video block.

[0146] In some embodiments, at least one set of weights among multiple sets of weights may be determined based on the reconstructed samples of a second color component different from a first color component, neighboring reconstructed samples of the second color component, neighboring reconstructed samples of the first color component, predicted samples of the first color component, a least mean squares (LMS) scheme, an LDL decomposition scheme, and / or a Cholesky decomposition scheme.

[0147] In some embodiments, multiple sets of weights may depend on at least one of a color format or a color component. For example, multiple sets of weights may be the same for multiple color components. Alternatively, multiple sets of weights may be different for multiple color components.

[0148] In some embodiments, how to obtain multiple sets of weights may be different for different color formats. By way of example and not limitation, if the color format for a current video block is a 4:2:0 color format or a 4:2:2 color format, multiple sets of weights may be determined based on the downsampled reconstructed samples of a second color component different from a first color component. If the color format for a current video block is a 4:4:4 color format, multiple sets of weights may be determined based on the non-downsampled reconstructed samples of the second color component.

[0149] In some embodiments, multiple sets of weights may be different for different coding information, different coding contexts, or different coding backgrounds. For example, multiple sets of weights may depend on the block size of a sub-picture, the block size of a slice, the block size of a strip, the block size of a CTU, the block size of a CU, the block size of a PU, the block size of a TU, the block size of a CTB, the block size of a CB, the block size of a PB, the block size of a TB, etc.

[0150] For example, for a non-LM intra prediction mode, if the block size is a first size (such as 256×256, etc.), the first weight among multiple sets of weights may be equal to a first value, and if the block size is a second size (such as 128×128, etc.), the first weight among multiple sets of weights may be equal to a second value.

[0151] In some embodiments, multiple sets of weights may depend on the video content of the current video block. By way of example and not limitation, for the LM intra prediction mode, if the type of the video content is a natural sequence, the second weight in the multiple sets of weights may be equal to a first value, and if the type of the video content is screen content, the second weight in the multiple sets of weights may be equal to a second value.

[0152] In some embodiments, the bitstream may include at least one syntax element indicating at least one of the following: whether prediction fusion is applied, how prediction fusion is applied, or weights for prediction fusion. In one example, the at least one syntax element may be indicated in the SPS, PPS, APS, picture header, slice header, etc.

[0153] In some embodiments, the at least one syntax element may include a syntax element indicating whether prediction fusion is applied. Additionally or alternatively, the at least one syntax element may include a set of syntax elements indicating how the weights are predetermined. For example, each syntax element in the set of syntax elements indicates a weight.

[0154] In some embodiments, at least one of the following may be indicated in the bitstream or determined based on the coding information of the current video block or the coding information of the blocks coded before the current video block: whether multiple sets of weights are used, or how multiple sets of weights are used.

[0155] In some embodiments, the bitstream may include one or more group indices indicating a set of weights that may be used in the multiple sets of weights. In one example, a single group index may be used to indicate a set of weights for multiple color components. In another example, the indication of a set of weights for multiple color components may be signaled separately.

[0156] In some embodiments, the one or more group indices may be determined based on the coding information of the current video block or the coding information of the blocks coded before the current video block. Alternatively, the one or more group indices may be determined based on a template matching-based scheme. In some additional embodiments, the group indices for multiple color components may be determined together. Alternatively, the group indices for multiple color components may be determined separately.

[0157] In some embodiments, the target prediction for a first color component may be determined based on multiple sets of weights, multiple candidate predictions, and at least one offset. By way of example and not limitation, one of the at least one offset may be represented as a median or a midpoint. The at least one offset may not depend on the bit depth.

[0158] In some embodiments, at least one offset may be adaptive and used to determine a model for determining a target prediction. For example, the model may be linear or non-linear. For example, at least one offset may be determined based on the codec information of the current video block or the codec information of the blocks that have been coded before the current video block. By way of example and not limitation, the codec information may include the reconstructed samples of a second color component that is different from a first color component, the neighboring reconstructed samples of the second color component, the neighboring reconstructed samples of the first color component, and / or the predicted samples of the first color component. Compared with the conventional solutions using fixed offsets, with the aid of the adaptive offset, the proposed method can advantageously improve the coding and decoding quality, especially for blocks or frames with different signal and texture distributions.

[0159] In some embodiments, at least one offset may be determined from codec tools such as luminance mapping and chrominance scaling (LMCS) or local illumination compensation (LIC).

[0160] In some embodiments, at least one offset may be determined from different coded signals using different calculation methods. In one example, at least one offset may be determined as the average value of one or more types of signals in a block region. In another example, at least one offset may be determined as the median value of one or more types of signals in a block region. By way of example and not limitation, one or more types of signals may include template values, the prediction of a second color component that is different from a first color component, the prediction of the first color component, and / or the reconstruction of the second color component.

[0161] In some embodiments, at least one offset may be determined from different schemes depending on at least one of the codec information, the codec context, or the codec background. In some additional embodiments, at least one offset may be determined based on the codec position information. For example, at least one offset may be determined based on the left block of the current video block and / or the upper block of the current video block.

[0162] In some embodiments, at least one offset may be indicated in the bitstream. In one example, at least one offset may be indicated at a block level, a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level. Alternatively, at least one offset may be indicated in the coding structure of a coding tree unit (CTU), the coding structure of a coding unit (CU), the coding structure of a transform unit (TU), the coding structure of a prediction unit (PU), the coding structure of a coding tree block (CTB), the coding structure of a coding block (CB), the coding structure of a transform block (TB), the coding structure of a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a slice header, or a slice group header.

[0163] In some embodiments, a third syntax element indicating whether prediction fusion with at least one offset is applied may be included in the bitstream. Additionally or alternatively, a fourth syntax element indicating the level of at least one offset applied may be included in the bitstream. In some further embodiments, if the third syntax element indicates that prediction fusion with at least one offset is applied, the fourth syntax element is included in the bitstream.

[0164] In some alternative embodiments, the bitstream may include a single syntax element indicating whether prediction fusion with at least one offset is applied and the level of at least one offset applied.

[0165] In some embodiments, a syntax element may be coded using fixed-length coding, exponential Golomb (EG) coding, truncated unary coding, or unary coding. Alternatively, a syntax element may be coded using at least one context in arithmetic coding. In some further embodiments, a syntax element may be bypass-coded.

[0166] In some embodiments, if prediction fusion with at least one offset is allowed to be used, syntax elements associated with the prediction fusion may be indicated in the bitstream. In some embodiments, the syntax elements may be included in a supplementary enhancement information (SEI) message or a video usability information (VUI) message.

[0167] In some embodiments, the same offset may be used for multiple color components. Alternatively, different offsets may be used for multiple first color components.

[0168] In some embodiments, at least one offset may include multiple offsets. In such a case, the bitstream may include at least one syntax element indicating at least one of the following: whether to apply prediction fusion with multiple offsets, how to apply prediction fusion with multiple offsets, or the offsets for prediction fusion. For example, the at least one syntax element may be indicated in one of the following: SPS, PPS, APS, picture header, slice header, CTU, or CU.

[0169] In some embodiments, the gradient of the sample values, the non-downsampled samples of a second color component different from the first color component, the downsampled samples of the second color component, neighboring samples, linear models, non-linear models, the neighboring reconstructed samples of the second color component, the neighboring reconstructed samples of the first color component, and / or the position information of the current video block may be used to determine the target prediction for the first color component.

[0170] In some embodiments, the gradient may be used in a model for determining the target prediction for the first color component. For example, the gradient may be determined based on the reconstructed samples of the second color component, the non-downsampled samples of the second color component, the downsampled samples of the second color component, the predicted samples of the first color component, etc. In some additional embodiments, how to determine the gradient may depend on at least one of the color format or color component of the current video block.

[0171] In some embodiments, the gradient may include different types of gradients. For example, the gradient may include at least one of a horizontal gradient or a vertical gradient.

[0172] In some embodiments, the gradient may be determined based on different schemes. In one example, the horizontal gradient may be determined based on the linear sum, linear difference, multiplication, or division of the signals along the horizontal direction. Additionally or alternatively, the vertical gradient may be determined based on the linear sum, linear difference, multiplication, or division of the signals along the vertical direction.

[0173] In some additional embodiments, the horizontal gradient may be determined based on a scheme for deriving the signal along the horizontal direction. Additionally or alternatively, the vertical gradient may be determined based on a scheme for deriving the signal along the vertical direction.

[0174] In some embodiments, the gradient may be determined based on different computational shapes, such as the square shape shown in Figure 60 or the diamond shape shown in Figure 61 For example, the gradient may be determined based on a square block or a diamond block covering the first sample in the current video block. By way of example and not limitation, the gradient may be determined based on the reconstructed samples of the second color component adjacent to the first sample, which is shown in Section 4 above.

[0175] In some embodiments, at least one term in the equation for determining the gradient can be modified. For example, at least one term in the equation can be replaced. Alternatively, at least one additional term can be added to the equation. In some further embodiments, multiple terms in the equation can be combined.

[0176] In some embodiments, multiple downsampling filters can be allowed for determining the target prediction for the first color component. For example, the multiple downsampling filters can include a 6-tap filter or a 3-tap filter. By way of example and not limitation, the 6-tap filter can have coefficients 1 - 2 - 1 - 1 - 2 - 1. The 3-tap filter can have coefficients 1 - 2 - 1.

[0177] In some embodiments, the bitstream can include at least one syntax element indicating at least one of the following: whether prediction fusion with a downsampling filter is applied, how prediction fusion with a downsampling filter is applied, or one or more downsampling filters for prediction fusion. In one example, the at least one syntax element can be included in a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a picture header, a slice header, a coding tree unit (CTU), or a coding unit (CU).

[0178] In some embodiments, the at least one syntax element can include a fifth syntax element indicating whether prediction fusion with a downsampling filter is applied. Additionally or alternatively, the at least one syntax element can include a sixth syntax element indicating one or more downsampling filters for prediction fusion. In some further embodiments, if the fifth syntax element indicates that prediction fusion with a downsampling filter is applied, the at least one syntax element can further include the sixth syntax element.

[0179] In some alternative embodiments, the at least one syntax element can include a single syntax element indicating whether prediction fusion with a downsampling filter is applied and one or more downsampling filters for prediction fusion.

[0180] In some embodiments, the at least one syntax element can be decoded using fixed - length coding, exponential Golomb (EG) coding, truncated unary coding, or unary coding. Alternatively, the at least one syntax element can be decoded using at least one context in arithmetic coding. In some embodiments, the at least one syntax element can be bypass - decoded.

[0181] In some embodiments, if predictive fusion with a downsampling filter is allowed to be used, the bitstream may include at least one syntax element indicating at least one of the following: whether predictive fusion with a downsampling filter is applied, how predictive fusion with a downsampling filter is applied, or one or more downsampling filters for predictive fusion. In some embodiments, the at least one syntax element may be included in an SEI message or a VUI message.

[0182] In some embodiments, whether a non-downsampling filter is used to determine a target prediction for a first color component may depend on the coding information, coding context, coding background, and / or video content of the current video block. For example, if the type of the video content is screen content, a non-downsampling filter may be used to determine the target prediction.

[0183] In some embodiments, the bitstream may include at least one high-level syntax (HLS) element indicating at least one of the following: whether predictive fusion with a downsampling filter is applied, how predictive fusion with a downsampling filter is applied, or one or more downsampling filters for predictive fusion. In one example, the at least one HLS element may be included in one of the following: a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a picture header, or a slice header.

[0184] In some embodiments, the at least one HLS element may include a first HLS element indicating whether predictive fusion with a downsampling filter is applied. Additionally or alternatively, the at least one HLS element may include a set of HLS elements indicating at least one downsampling filter or non-downsampling filter that may be selected. In some embodiments, each HLS element in the set of HLS elements indicates a downsampling filter.

[0185] In some embodiments, whether the method is applied and / or how the method is applied may be indicated at the block level, sequence level, group of pictures level, picture level, slice level, or slice group level. Alternatively, whether the method is applied and / or how the method is applied may be indicated in the coding structure of a CTU, CU, TU, PU, CTB, CB, TB, PB, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.

[0186] In some embodiments, whether to apply the method and / or how to apply the method may depend on the information of the warp decoding of the current video block. For example, the information of the warp decoding may include block size, color format, single-tree segmentation, double-tree segmentation, color component, stripe type, picture type, etc.

[0187] In some embodiments, the method may also be applicable to the codec tools that require prediction fusion.

[0188] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a video processing device for a video. In this method, multiple sets of weights are obtained. A target prediction for a first color component of a current video block of the video is determined based on the multiple sets of weights and multiple candidate predictions for the first color component. In addition, the bitstream is generated based on the target prediction.

[0189] According to still some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, multiple sets of weights are obtained. A target prediction for a first color component of a current video block of the video is determined based on the multiple sets of weights and multiple candidate predictions for the first color component. In addition, the bitstream is generated based on the target prediction and is stored in a non-transitory computer-readable recording medium.

[0190] Embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.

[0191] Item 1. A method for video processing, including: obtaining multiple sets of weights for the conversion between a current video block of a video and the bitstream of the video; determining a target prediction for a first color component of the current video block based on the multiple sets of weights and multiple candidate predictions for the first color component; and performing the conversion based on the target prediction.

[0192] Item 2. The method according to Item 1, wherein the first color component includes a chrominance component or a luminance component.

[0193] Item 3. The method according to any one of Items 1 to 2, wherein the multiple candidate predictions are determined based on different prediction schemes, the multiple sets of weights are predetermined, or at least one set of weights among the multiple sets of weights is determined based on a non-linear model.

[0194] Item 4. The method according to any one of Items 1 to 3, wherein the multiple sets of weights are predetermined.

[0195] Item 5. The method according to Item 4, wherein information related to at least one of the following is indicated in the bitstream: whether to use the multiple sets of weights, or a set of weights used among the multiple sets of weights.

[0196] Item 6. The method according to Item 4, wherein information related to at least one of the following is determined based on the coding information of the current video block or the coding information of the blocks coded before the current video block: whether to use the multiple sets of weights, or a set of weights used among the multiple sets of weights.

[0197] Item 7. The method according to any one of Items 1 to 6, wherein the number of groups in the multiple sets of weights depends on the coding information of the current video block or the coding information of the blocks coded before the current video block.

[0198] Item 8. The method according to any one of Items 6 to 7, wherein the coding information includes at least one of the following: block dimension, block size, block depth, stripe type, picture type, segmentation tree type, temporal layer identifier, or quantization parameter.

[0199] Item 9. The method according to any one of Items 1 to 8, wherein the target prediction is determined based on a weighted sum of the multiple candidate predictions.

[0200] Item 10. The method according to Item 9, wherein the target prediction is determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, w i represents the weight for the i-th candidate prediction, n represents one less than the number of candidate predictions among the multiple candidate predictions, and midvalue represents an offset.

[0201] Item 11. The method according to Item 9, wherein the target prediction is determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, w i represents the weight for the i-th candidate prediction, n represents one less than the number of candidate predictions among the multiple candidate predictions, b represents an offset, and k represents the number of right shifts.

[0202] Item 12. The method according to any one of Items 10 to 11, wherein the value of the weight w i is predefined.

[0203] Item 13. The method according to any one of Items 10 to 11, wherein the weight w i has a value that depends on the context of the codec used to encode and decode the current video block.

[0204] Item 14. The method according to any one of Items 1 to 8, wherein the target prediction is determined based on a non-linear fusion scheme for fusing the plurality of candidate predictions.

[0205] Item 15. The method according to Item 14, wherein a convolutional neural network is used in the non-linear fusion scheme.

[0206] Item 16. The method according to Item 15, wherein the target prediction is determined as follows: C(p i ) = conv2(w i , p i ) + b i , i = 0, 1, …, n B = M(C(p1), C(p2), …, C(p n )) where B represents the target prediction, M() represents the fusion scheme, p i represents the i-th candidate prediction among the plurality of candidate predictions, C(p i ) represents the intermediate result corresponding to the i-th candidate prediction, conv2() represents two-dimensional convolution, w i represents the weight of the convolutional layer for the i-th candidate prediction, b i is the bias of the convolutional layer, and n represents one less than the number of candidate predictions among the plurality of candidate predictions.

[0207] Item 17. The method according to Item 16, wherein a rectified linear unit (ReLU) is used as the activation function for the convolutional neural network.

[0208] Item 18. The method according to Item 17, wherein the rectified linear unit is as follows: where R(p i ) represents the rectified linear unit corresponding to the i-th candidate prediction.

[0209] Item 19. The method according to any one of Items 1 to 18, wherein one of the plurality of candidate predictions includes one of the following: a prediction of a sample point for the first color component, a prediction of a sample point for a second color component different from the first color component, or a prediction of a sample point after filtering or rescaling.

[0210] Item 20. The method according to any one of Items 1 to 9, wherein the number of candidate predictions among the multiple candidate predictions is 2.

[0211] Item 21. The method according to Item 20, wherein one candidate prediction among the multiple candidate predictions is determined based on a linear model (LM) mode, and another candidate prediction among the multiple candidate predictions is determined based on a non-LM mode.

[0212] Item 22. The method according to any one of Items 20 to 21, wherein the target prediction is determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, and w i represents the weight for the i-th candidate prediction.

[0213] Item 23. The method according to Item 22, wherein the value of the weight w i is predefined.

[0214] Item 24. The method according to any one of Items 22 to 23, wherein the value of the weight w i depends on the context of the codec used to encode and decode the current video block.

[0215] Item 25. The method according to any one of Items 1 to 24, wherein at least one set of weights among the multiple sets of weights is indicated in the bitstream.

[0216] Item 26. The method according to Item 25, wherein the at least one set of weights is indicated at one of the following: block level, sequence level, group of pictures level, picture level, strip level, or slice group level.

[0217] Item 27. The method according to Item 25, wherein the at least one set of weights is indicated in one of the following: the coding and decoding structure of a coding tree unit (CTU), the coding and decoding structure of a coding unit (CU), the coding and decoding structure of a transform unit (TU), the coding and decoding structure of a prediction unit (PU), the coding and decoding structure of a coding tree block (CTB), the coding and decoding structure of a coding block (CB), the coding and decoding structure of a transform block (TB), the coding and decoding structure of a prediction block (PB), sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.

[0218] Item 28. The method according to any one of Items 1 to 27, wherein a first syntax element indicating whether a prediction fusion with multiple weights is applied is included in the bitstream.

[0219] Item 29. The method according to any one of Items 1 to 28, wherein a second syntax element indicating a set of weights applied among the multiple sets of weights is included in the bitstream.

[0220] Item 30. The method according to Item 28, wherein if the first syntax element indicates that the prediction fusion with the multiple weights is applied, a second syntax element indicating a set of weights applied among the multiple sets of weights is included in the bitstream.

[0221] Item 31. The method according to any one of Items 1 to 27, wherein the bitstream includes a single syntax element indicating whether a prediction fusion with multiple weights is applied and a set of weights applied among the multiple sets of weights.

[0222] Item 32. The method according to any one of Items 28 to 31, wherein the syntax element is coded and decoded using one of the following: fixed-length coding and decoding, exponential Golomb (EG) coding and decoding, truncated unary coding and decoding, or unary coding and decoding.

[0223] Item 33. The method according to any one of Items 28 to 31, wherein the syntax element is coded and decoded using at least one context in arithmetic coding.

[0224] Item 34. The method according to any one of Items 28 to 31, wherein the syntax element is bypass coded and decoded.

[0225] Item 35. The method according to any one of Items 28 to 34, wherein if the prediction fusion with the multiple weights is allowed to be used, the syntax element associated with the prediction fusion is indicated in the bitstream.

[0226] Item 36. The method according to any one of Items 28 to 31, wherein the syntax element is included in a supplementary enhancement information (SEI) message or a video usability information (VUI) message.

[0227] Item 37. The method according to any one of Items 25 to 27, wherein the at least one set of weights is indicated in a predictive manner.

[0228] Item 38. The method according to Item 38, wherein the weights in the at least one set of weights are indicated in a predictive manner.

[0229] Item 39. The method according to any one of Items 1 to 24, wherein for the current video block, the multiple sets of weights are not signaled, and the multiple sets of weights are reused from blocks that have been encoded and decoded before the current video block.

[0230] Item 40. The method according to any one of Items 1 to 24, wherein at least one set of weights among the multiple sets of weights is determined based on at least one of the following: reconstructed samples of a second color component different from the first color component, neighboring reconstructed samples of the second color component, neighboring reconstructed samples of the first color component, predicted samples of the first color component, a least mean squares (LMS) scheme, an LDL decomposition scheme, or a Cholesky decomposition scheme.

[0231] Item 41. The method according to any one of Items 1 to 24, wherein the multiple sets of weights depend on at least one of a color format or a color component.

[0232] Item 42. The method according to any one of Items 1 to 24, wherein the multiple sets of weights are the same for multiple color components, or the multiple sets of weights are different for multiple color components.

[0233] Item 43. The method according to Item 41, wherein for different color formats, how to obtain the multiple sets of weights is different.

[0234] Item 44. The method according to Item 43, wherein if the color format for the current video block is a 4:2:0 color format or a 4:2:2 color format, the multiple sets of weights are determined based on downsampled reconstructed samples of a second color component different from the first color component, or if the color format for the current video block is a 4:4:4 color format, the multiple sets of weights are determined based on non-downsampled reconstructed samples of the second color component.

[0235] Item 45. The method according to any one of Items 1 to 24, wherein the multiple sets of weights are different for different encoding and decoding information, different encoding and decoding contexts, or different encoding and decoding backgrounds.

[0236] Item 46. The method according to any one of Items 1 to 24, wherein the multiple sets of weights depend on at least one of the following: the block size of a sub-picture, the block size of a slice, the block size of a strip, the block size of a coding tree unit (CTU), the block size of a coding unit (CU), the block size of a prediction unit (PU), the block size of a transform unit (TU), the block size of a coding tree block (CTB), the block size of a coding block (CB), the block size of a prediction block (PB), or the block size of a transform block (TB).

[0237] Item 47. The method according to Item 46, wherein for a non-LM intra prediction mode, if the block size is a first size, the first weight in the multiple sets of weights is equal to a first value, and if the block size is a second size, the first weight in the multiple sets of weights is equal to a second value.

[0238] Item 48. The method according to any one of Items 1 to 24, wherein the multiple sets of weights depend on the video content of the current video block.

[0239] Item 49. The method according to Item 48, wherein for an LM intra prediction mode, if the type of the video content is a natural sequence, the second weight in the multiple sets of weights is equal to a first value, and if the type of the video content is screen content, the second weight in the multiple sets of weights is equal to a second value.

[0240] Item 50. The method according to any one of Items 1 to 49, wherein the bitstream includes at least one syntax element indicating at least one of the following: whether prediction fusion is applied, how the prediction fusion is applied, or the weights for the prediction fusion.

[0241] Item 51. The method according to Item 50, wherein the at least one syntax element is indicated in one of the following: SPS, PPS, APS, picture header, or slice header.

[0242] Item 52. The method according to any one of Items 50 to 51, wherein the at least one syntax element includes a syntax element indicating whether the prediction fusion is applied.

[0243] Item 53. The method according to any one of Items 50 to 52, wherein the at least one syntax element includes a set of syntax elements indicating how the weights are predetermined.

[0244] Item 54. The method according to Item 53, wherein each syntax element in the set of syntax elements indicates a weight.

[0245] Item 55. The method according to any one of Items 1 to 54, wherein at least one of the following is indicated in the bitstream or determined based on the codec information of the current video block or the codec information of the blocks encoded before the current video block: whether the multiple sets of weights are used, or how the multiple sets of weights are used.

[0246] Item 56. The method according to any one of Items 1 to 54, wherein the bitstream includes one or more group indices indicating a set of weights used in the multiple sets of weights.

[0247] Item 57. The method according to Item 56, wherein a single group index is used to indicate the set of weights used for multiple color components, or wherein the indication of the set of weights used for multiple color components is signaled separately.

[0248] Item 58. The method according to Item 56, wherein the one or more group indexes are determined based on the coding and decoding information of the current video block or the coding and decoding information of the blocks coded and decoded before the current video block, or the one or more group indexes are determined based on a template matching-based scheme, or the group indexes for multiple color components are determined together, or the group indexes for multiple color components are determined separately.

[0249] Item 59. The method according to any one of Items 1 to 58, wherein the target prediction for the first color component is determined based on the multiple sets of weights, the multiple candidate predictions, and at least one offset.

[0250] Item 60. The method according to Item 59, wherein the at least one offset is adaptive and is used to determine a model for determining the target prediction.

[0251] Item 61. The method according to Item 60, wherein the model is linear or non-linear.

[0252] Item 62. The method according to any one of Items 59 to 61, wherein the at least one offset is determined based on the coding and decoding information of the current video block or the coding and decoding information of the blocks coded and decoded before the current video block.

[0253] Item 63. The method according to Item 62, wherein the coding and decoding information includes at least one of the following: reconstructed samples of a second color component different from the first color component, neighboring reconstructed samples of the second color component, neighboring reconstructed samples of the first color component, or predicted samples of the first color component.

[0254] Item 64. The method according to any one of Items 59 to 61, wherein the at least one offset is determined from coding and decoding tools.

[0255] Item 65. The method according to Item 64, wherein the coding and decoding tools include luminance mapping and chrominance scaling (LMCS) or local illumination compensation (LIC).

[0256] Item 66. The method according to any one of Items 59 to 61, wherein the at least one offset is determined from different coding and decoding signals using different calculation methods.

[0257] Item 67. The method according to any one of Items 59 to 61, wherein the at least one offset is determined as an average value of one or more types of signals in a block region, or wherein the at least one offset is determined as a median value of one or more types of signals in the block region.

[0258] Item 68. The method according to Item 67, wherein the one or more types of signals include at least one of the following: a template value, a prediction of a second color component different from the first color component, a prediction of the first color component, or a reconstruction of the second color component.

[0259] Item 69. The method according to any one of Items 59 to 61, wherein the at least one offset is determined from different schemes depending on at least one of encoding / decoding information, encoding / decoding context, or encoding / decoding background.

[0260] Item 70. The method according to any one of Items 59 to 61, wherein the at least one offset is determined based on encoding / decoding position information.

[0261] Item 71. The method according to Item 70, wherein the at least one offset is determined based on at least one of the following: a left block of the current video block, or an upper block of the current video block.

[0262] Item 72. The method according to any one of Items 59 to 71, wherein the at least one offset is indicated in the bitstream.

[0263] Item 73. The method according to Item 72, wherein the at least one offset is indicated at one of the following: block level, sequence level, picture group level, picture level, strip level, or slice group level.

[0264] Item 74. The method according to Item 72, wherein the at least one offset is indicated in one of the following: an encoding / decoding structure of a coding tree unit (CTU), an encoding / decoding structure of a coding unit (CU), an encoding / decoding structure of a transform unit (TU), an encoding / decoding structure of a prediction unit (PU), an encoding / decoding structure of a coding tree block (CTB), an encoding / decoding structure of a coding block (CB), an encoding / decoding structure of a transform block (TB), an encoding / decoding structure of a prediction block (PB), a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptive parameter set (APS), a strip header, or a slice group header.

[0265] Item 75. The method according to any one of Items 59 to 74, wherein a third syntax element indicating whether prediction fusion with the at least one offset is applied is included in the bitstream.

[0266] Item 76. The method according to any one of Items 59 to 75, wherein a fourth syntax element indicating the level of the at least one applied offset is included in the bitstream.

[0267] Item 77. The method according to Item 75, wherein if the third syntax element indicates that the prediction merge with the at least one offset is applied, a fourth syntax element indicating the level of the at least one applied offset is included in the bitstream.

[0268] Item 78. The method according to any one of Items 59 to 74, wherein the bitstream includes a single syntax element indicating whether the prediction merge with the at least one offset is applied and the level of the at least one applied offset.

[0269] Item 79. The method according to any one of Items 75 to 78, wherein the syntax element is coded and decoded using one of the following: fixed-length coding and decoding, exponential Golomb (EG) coding and decoding, truncated unary coding and decoding, or unary coding and decoding.

[0270] Item 80. The method according to any one of Items 75 to 78, wherein the syntax element is coded and decoded using at least one context in arithmetic coding.

[0271] Item 81. The method according to any one of Items 75 to 78, wherein the syntax element is bypass coded and decoded.

[0272] Item 82. The method according to any one of Items 75 to 81, wherein if the prediction merge with the at least one offset is allowed to be used, the syntax element associated with the prediction merge is indicated in the bitstream.

[0273] Item 83. The method according to any one of Items 75 to 82, wherein the syntax element is included in an SEI message or a VUI message.

[0274] Item 84. The method according to any one of Items 59 to 83, wherein the same offset is used for multiple color components, or different offsets are used for the multiple first color components.

[0275] Item 85. The method according to any one of Items 59 to 74, wherein the at least one offset includes a plurality of offsets.

[0276] Item 86. The method according to Item 85, wherein the bitstream includes at least one syntax element indicating at least one of the following: whether the prediction merge with the plurality of offsets is applied, how the prediction merge with the plurality of offsets is applied, or the offsets used for the prediction merge.

[0277] Item 87. The method according to item 86, wherein the at least one syntax element is indicated in one of the following: SPS, PPS, APS, picture header, slice header, CTU, or CU.

[0278] Item 88. The method according to any one of items 1 to 87, wherein at least one of the following is used to determine the target prediction for the first color component: gradient of sample values, non-downsampled samples of a second color component different from the first color component, downsampled samples of the second color component, neighboring samples, linear model, non-linear model, neighboring reconstructed samples of the second color component, neighboring reconstructed samples of the first color component, position information of the current video block.

[0279] Item 89. The method according to item 88, wherein the gradient is used in a model for determining the target prediction for the first color component.

[0280] Item 90. The method according to any one of items 88 to 89, wherein the gradient is determined based on at least one of the following: reconstructed samples of the second color component, non-downsampled samples of the second color component, downsampled samples of the second color component, or predicted samples of the first color component.

[0281] Item 91. The method according to any one of items 88 to 90, wherein how the gradient is determined depends on at least one of the color format or color component of the current video block.

[0282] Item 92. The method according to any one of items 88 to 91, wherein the gradient includes different types of gradients.

[0283] Item 93. The method according to item 92, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient.

[0284] Item 94. The method according to any one of items 88 to 93, wherein the gradient is determined based on different schemes.

[0285] Item 95. The method according to item 93, wherein the horizontal gradient is determined based on a linear sum, linear difference, multiplication, or division of signals along the horizontal direction, or the vertical gradient is determined based on a linear sum, linear difference, multiplication, or division of signals along the vertical direction.

[0286] Item 96. The method according to item 93, wherein the horizontal gradient is determined based on a scheme for deriving a signal along the horizontal direction, or the vertical gradient is determined based on a scheme for deriving a signal along the vertical direction.

[0287] Item 97. The method according to any one of items 88 to 96, wherein the gradient is determined based on different calculation shapes.

[0288] Item 98. The method according to any one of items 88 to 96, wherein the gradient is determined based on a square block or a diamond block covering the first samples in the current video block.

[0289] Item 99. The method according to item 98, wherein the gradient is determined based on the reconstructed samples of the second color component adjacent to the first samples.

[0290] Item 100. The method according to any one of items 94 to 99, wherein at least one term in the equation for determining the gradient is modified, at least one term in the equation is replaced, at least one additional term is added to the equation, or multiple terms in the equation are combined.

[0291] Item 101. The method according to any one of items 1 to 87, wherein multiple downsampling filters are allowed to be used to determine the target prediction for the first color component.

[0292] Item 102. The method according to item 101, wherein the multiple downsampling filters include a 6-tap filter or a 3-tap filter.

[0293] Item 103. The method according to item 102, wherein the 6-tap filter has coefficients 1 - 2 - 1 - 1 - 2 - 1, or the 3-tap filter has coefficients 1 - 2 - 1.

[0294] Item 104. The method according to any one of items 101 to 103, wherein the bitstream includes at least one syntax element indicating at least one of the following: whether to apply prediction fusion with a downsampling filter, how to apply the prediction fusion with a downsampling filter, or one or more downsampling filters for the prediction fusion.

[0295] Item 105. The method according to item 104, wherein the at least one syntax element is included in one of the following: a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a picture header, a slice header, a coding tree unit (CTU), or a coding unit (CU).

[0296] Item 106. The method according to any one of Items 104 to 105, wherein the at least one syntax element includes a fifth syntax element indicating whether to apply the prediction fusion with a downsampling filter.

[0297] Item 107. The method according to any one of Items 104 to 106, wherein the at least one syntax element includes a sixth syntax element indicating one or more downsampling filters for the prediction fusion.

[0298] Item 108. The method according to Item 106, wherein if the fifth syntax element indicates that the prediction fusion with a downsampling filter is applied, the at least one syntax element further includes a sixth syntax element indicating one or more downsampling filters for the prediction fusion.

[0299] Item 109. The method according to any one of Items 104 to 105, wherein the at least one syntax element includes a single syntax element indicating whether to apply the prediction fusion with a downsampling filter and one or more downsampling filters for the prediction fusion.

[0300] Item 110. The method according to any one of Items 104 to 109, wherein the at least one syntax element is encoded and decoded using one of the following: fixed-length encoding and decoding, exponential Golomb (EG) encoding and decoding, truncated unary encoding, or unary encoding.

[0301] Item 111. The method according to any one of Items 104 to 110, wherein the at least one syntax element is encoded and decoded using at least one context in arithmetic encoding and decoding.

[0302] Item 112. The method according to any one of Items 104 to 110, wherein the at least one syntax element is bypass encoded and decoded.

[0303] Item 113. The method according to any one of Items 101 to 103, wherein if the prediction fusion with a downsampling filter is allowed to be used, the bitstream includes at least one syntax element indicating at least one of the following: whether to apply the prediction fusion with a downsampling filter, how to apply the prediction fusion with a downsampling filter, or one or more downsampling filters for the prediction fusion.

[0304] Item 114. The method according to any one of Items 104 to 113, wherein the at least one syntax element is included in an SEI message or a VUI message.

[0305] Item 115. The method according to any one of Items 1 to 87, wherein whether a non-downsampling filter is used to determine the target prediction for the first color component depends on at least one of the coding information, coding context, coding background, or video content of the current video block.

[0306] Item 116. The method according to Item 115, wherein if the type of the video content is screen content, the non-downsampling filter is used to determine the target prediction.

[0307] Item 117. The method according to any one of Items 1 to 116, wherein the bitstream includes at least one high-level syntax (HLS) element indicating at least one of: whether prediction fusion with a downsampling filter is applied, how the prediction fusion with the downsampling filter is applied, or one or more downsampling filters for the prediction fusion.

[0308] Item 118. The method according to Item 117, wherein the at least one HLS element is included in one of: a sequence parameter set (SPS), a picture parameter set (PPS), an adaptive parameter set (APS), a picture header, or a slice header.

[0309] Item 119. The method according to any one of Items 117 to 118, wherein the at least one HLS element includes a first HLS element indicating whether the prediction fusion with a downsampling filter is applied.

[0310] Item 120. The method according to any one of Items 117 to 119, wherein the at least one HLS element includes a set of HLS elements indicating at least one selected downsampling filter or non-downsampling filter.

[0311] Item 121. The method according to Item 120, wherein each HLS element in the set of HLS elements indicates a downsampling filter.

[0312] Item 122. The method according to any one of Items 1 to 121, wherein whether the method is applied and / or how the method is applied is indicated at one of: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

[0313] Item 123. The method according to any one of Items 1 to 121, wherein whether to apply the method and / or how to apply the method is indicated in one of the following: the encoding / decoding structure of the CTU, the encoding / decoding structure of the CU, the encoding / decoding structure of the TU, the encoding / decoding structure of the PU, the encoding / decoding structure of the CTB, the encoding / decoding structure of the CB, the encoding / decoding structure of the TB, the encoding / decoding structure of the PB, the sequence header, the picture header, the sequence parameter set (SPS), the video parameter set (VPS), the dependency parameter set (DPS), the decoding capability information (DCI), the picture parameter set (PPS), the adaptive parameter set (APS), the slice header, or the slice group header.

[0314] Item 124. The method according to any one of Items 1 to 123, wherein whether to apply the method and / or how to apply the method depends on the encoded / decoded information of the current video block.

[0315] Item 125. The method according to Item 124, wherein the encoded / decoded information includes at least one of the following: block size, color format, single-tree segmentation, double-tree segmentation, color component, slice type, or picture type.

[0316] Item 126. The method according to any one of Items 1 to 125, wherein the method is applicable to encoding / decoding tools that require prediction fusion.

[0317] Item 127. The method according to any one of Items 1 to 126, wherein the conversion includes encoding the current video block into the bitstream.

[0318] Item 128. The method according to any one of Items 1 to 126, wherein the conversion includes decoding the current video block from the bitstream.

[0319] Item 129. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Items 1 to 128.

[0320] Item 130. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Items 1 to 128.

[0321] Item 131. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by a video processing apparatus for a video, wherein the method includes: obtaining multiple sets of weights; determining a target prediction for a first color component of the current video block of the video based on the multiple sets of weights and multiple candidate predictions for the first color component; and generating the bitstream based on the target prediction.

[0322] Article 132. A method for storing a bitstream of a video, comprising: obtaining multiple sets of weights; determining a target prediction for a first color component of a current video block of the video based on the multiple sets of weights and multiple candidate predictions for the first color component; generating the bitstream based on the target prediction; and storing the bitstream in a non-transitory computer-readable recording medium. Example device

[0323] Figure 63 A block diagram of a computing device 6300 in which various embodiments of the present disclosure can be implemented is shown. The computing device 6300 can be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or can be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).

[0324] It should be understood that Figure 63 the computing device 6300 shown is for illustrative purposes only and does not imply any limitation on the functionality and scope of the embodiments of the present disclosure in any way.

[0325] As Figure 63 shown, the computing device 6300 includes a general computing device 6300. The computing device 6300 can include at least one or more processors or processing units 6310, a memory 6320, a storage unit 6330, one or more communication units 6340, one or more input devices 6350, and one or more output devices 6360.

[0326] In some embodiments, the computing device 6300 can be implemented as any user terminal or server terminal having computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / video camera, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 6300 can support any type of interface to the user (such as a "wearable" circuitry, etc.).

[0327] The processing unit 6310 can be a physical processor or a virtual processor, and can implement various processes based on the programs stored in the memory 6320. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 6300. The processing unit 6310 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.

[0328] The computing device 6300 generally includes various computer storage media. Such media can be any media accessible by the computing device 6300, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 6320 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 6330 can be any removable or non-removable media, and can include machine-readable media, such as a memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 6300.

[0329] The computing device 6300 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 63 it, a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk can be provided. In this case, each drive can be connected to a bus (not shown) via one or more data media interfaces.

[0330] The communication unit 6340 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 6300 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Therefore, the computing device 6300 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.

[0331] The input device 6350 can be one or more of various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and so on. The output device 6360 can be one or more of various output devices, such as a display, a speaker, a printer, and so on. With the aid of the communication unit 6340, the computing device 6300 can also communicate with one or more external devices (not shown), such as a storage device and a display device, the computing device 6300 can also communicate with one or more devices that enable a user to interact with the computing device 6300, or if needed, the computing device 6300 can also communicate with any device (such as a network card, a modem, etc.) that enables the computing device 6300 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0332] In some embodiments, some or all components of the computing device 6300 can also be arranged in a cloud computing architecture instead of being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which will not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.

[0333] In an embodiment of the present disclosure, the computing device 6300 can be used to implement video encoding / decoding. The memory 6320 can include one or more video codec modules 6325 having one or more program instructions. These modules are accessible and executable by the processing unit 6310 to perform the functions of various embodiments described herein.

[0334] In an example embodiment of performing video encoding, an input device 6350 may receive video data as an input 6370 to be encoded. The video data may be processed, for example, by a video codec module 6325 to generate an encoded bitstream. The encoded bitstream may be provided as an output 6380 via an output device 6360.

[0335] In an example embodiment of performing video decoding, an input device 6350 may receive the encoded bitstream as an input 6370. The encoded bitstream may be processed, for example, by a video codec module 6325 to generate decoded video data. The decoded video data may be provided as an output 6380 via an output device 6360.

[0336] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.

Claims

1. A method for video processing, comprising: obtaining multiple sets of weights for the conversion between the current video block of a video and the bitstream of the video; determining a target prediction for the first color component based on the multiple sets of weights and multiple candidate predictions for the first color component of the current video block; and performing the conversion based on the target prediction.

2. The method according to claim 1, wherein the first color component comprises a chrominance component or a luminance component.

3. The method according to any one of claims 1 to 2, wherein the multiple candidate predictions are determined based on different prediction schemes, the multiple sets of weights are predetermined, or at least one set of weights among the multiple sets of weights is determined based on a non - linear model.

4. The method according to any one of claims 1 to 3, wherein the multiple sets of weights are predetermined.

5. The method according to claim 4, wherein information related to at least one of the following is indicated in the bitstream: whether to use the multiple sets of weights, or one set of weights among the multiple sets of weights to be used.

6. The method according to claim 4, wherein information related to at least one of the following is determined based on the coding - decoding information of the current video block or the coding - decoding information of the blocks coded - decoded before the current video block: whether to use the multiple sets of weights, or one set of weights among the multiple sets of weights to be used.

7. The method according to any one of claims 1 to 6, wherein the number of sets in the multiple sets of weights depends on the coding - decoding information of the current video block or the coding - decoding information of the blocks coded - decoded before the current video block.

8. The method according to any one of claims 6 to 7, wherein the coding - decoding information comprises at least one of the following: block dimension, block size, block depth, slice type, picture type, segmentation tree type, temporal layer identifier, or quantization parameter.

9. The method according to any one of claims 1 to 8, wherein the target prediction is determined based on a weighted sum of the multiple candidate predictions.

10. The method according to claim 9, wherein the target prediction is determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, w i represents the weight for the i-th candidate prediction, n represents one less than the number of candidate predictions among the multiple candidate predictions, and midvalue represents an offset.

11. The method according to claim 9, wherein the target prediction is determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, w i represents the weight for the i-th candidate prediction, n represents one less than the number of candidate predictions among the multiple candidate predictions, b represents an offset, and k represents the number of right shifts.

12. The method according to any one of claims 10 to 11, wherein the weight w i has a predefined value.

13. The method according to any one of claims 10 to 11, wherein the weight w i value depends on the context of the codec used to encode and decode the current video block.

14. The method according to any one of claims 1 to 8, wherein the target prediction is determined based on a non - linear fusion scheme for fusing the multiple candidate predictions.

15. The method according to claim 14, wherein a convolutional neural network is used in the non - linear fusion scheme.

16. The method according to claim 15, wherein the target prediction is determined as follows: C(p i ) = conv2(w i , p i ) + b i , i = 0, 1, …, n B = M(C(p1), C(p2), …, C(p n )) where B represents the target prediction, M() represents the fusion scheme, and p i represents the i-th candidate prediction among the multiple candidate predictions, and C(p i ) represents the intermediate result corresponding to the i-th candidate prediction, conv2() represents a two-dimensional convolution, and w i represents the weight of the convolutional layer for the i-th candidate prediction, b i is the bias of the convolutional layer, and n represents one less than the number of candidate predictions among the multiple candidate predictions.

17. The method according to claim 16, wherein a rectified linear unit (ReLU) is used as an activation function for the convolutional neural network.

18. The method according to claim 17, wherein the rectified linear unit is as follows: where R(p i ) represents the rectified linear unit corresponding to the i-th candidate prediction.

19. The method according to any one of claims 1 to 18, wherein one of the multiple candidate predictions comprises one of the following: a prediction of a sample point for the first color component, a prediction of a sample point for a second color component different from the first color component, or Prediction of samples after filtering or resizing.

20. The method according to any one of claims 1 to 9, wherein the number of candidate predictions among the plurality of candidate predictions is 2.

21. The method according to claim 20, wherein one candidate prediction among the plurality of candidate predictions is determined based on a linear model (LM) mode, and the other candidate prediction among the plurality of candidate predictions is determined based on a non-LM mode.

22. The method according to any one of claims 20 to 21, wherein the target prediction is determined as follows: where B represents the target prediction, p i represents the i-th candidate prediction among the multiple candidate predictions, and w i represents the weight for the i-th candidate prediction.

23. The method according to claim 22, wherein the weight w i has a predefined value.

24. The method according to any one of claims 22 to 23, wherein the weight w i value depends on the context of the codec used to encode and decode the current video block.

25. The method according to any one of claims 1 to 24, wherein at least one set of weights among the plurality of sets of weights is indicated in the bitstream.

26. The method according to claim 25, wherein the at least one set of weights is indicated at one of the following: block level, sequence level, group of pictures level, picture level, slice level, or slice group level.

27. The method according to claim 25, wherein the at least one set of weights is indicated in one of the following: coding tree unit (CTU) coding structure, coding unit (CU) coding structure, transformation unit (TU) coding structure, prediction unit (PU) coding structure, coding tree block (CTB) coding structure, coding block (CB) coding structure, transformation block (TB) coding structure, prediction block (PB) coding structure, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.

28. The method according to any one of claims 1 to 27, wherein a first syntax element indicating whether prediction fusion with multiple weights is applied is included in the bitstream.

29. The method according to any one of claims 1 to 28, wherein a second syntax element indicating a set of weights among the plurality of sets of weights that is applied is included in the bitstream.

30. The method according to claim 28, wherein if the first syntax element indicates that prediction fusion with the multiple weights is applied, a second syntax element indicating a set of weights among the plurality of sets of weights that is applied is included in the bitstream.

31. The method according to any one of claims 1 to 27, wherein the bitstream includes a single syntax element indicating whether prediction fusion with multiple weights is applied and a set of weights among the plurality of sets of weights that is applied.

32. The method according to any one of claims 28 to 31, wherein the syntax element is decoded using one of the following: fixed-length coding, exponential Golomb (EG) coding, truncated unary coding, or unary coding.

33. The method according to any one of claims 28 to 31, wherein the syntax element is decoded using at least one context in arithmetic coding.

34. The method according to any one of claims 28 to 31, wherein the syntax element is bypass decoded.

35. The method according to any one of claims 28 to 34, wherein if the prediction fusion with the plurality of weights is allowed to be used, the syntax element associated with the prediction fusion is indicated in the bitstream.

36. The method according to any one of claims 28 to 31, wherein the syntax element is included in a supplementary enhancement information (SEI) message or a video usability information (VUI) message.

37. The method according to any one of claims 25 to 27, wherein the at least one set of weights is indicated in a predictive manner.

38. The method according to claim 38, wherein the weights in the at least one set of weights are indicated in a predictive manner.

39. The method according to any one of claims 1 to 24, wherein for the current video block, the multiple sets of weights are not signaled, and the multiple sets of weights are reused from blocks encoded and decoded before the current video block.

40. The method according to any one of claims 1 to 24, wherein at least one set of the multiple sets of weights is determined based on at least one of the following: The reconstructed samples of a second color component different from the first color component, The neighboring reconstructed samples of the second color component, The neighboring reconstructed samples of the first color component, The predicted samples of the first color component, The least mean square (LMS) scheme, The LDL decomposition scheme, or The Cholesky decomposition scheme.

41. The method according to any one of claims 1 to 24, wherein the multiple sets of weights depend on at least one of the color format or the color component.

42. The method according to any one of claims 1 to 24, wherein the multiple sets of weights are the same for multiple color components, or The multiple sets of weights are different for multiple color components.

43. The method according to claim 41, wherein for different color formats, how to obtain the multiple sets of weights is different.

44. The method according to claim 43, wherein if the color format for the current video block is a 4:2:0 color format or a 4:2:2 color format, the multiple sets of weights are determined based on the downsampled reconstructed samples of a second color component different from the first color component, or If the color format for the current video block is a 4:4:4 color format, the multiple sets of weights are determined based on the non-downsampled reconstructed samples of the second color component.

45. The method according to any one of claims 1 to 24, wherein the multiple sets of weights are different for different coding information, different coding contexts, or different coding backgrounds.

46. The method according to any one of claims 1 to 24, wherein the multiple sets of weights depend on at least one of the following: The block size of a sub-picture, The block size of a slice, The block size of a strip, The block size of a coding tree unit (CTU), The block size of a coding unit (CU), The block size of a prediction unit (PU), The block size of a transform unit (TU), The block size of a coding tree block (CTB), The block size of a coding block (CB), The block size of a prediction block (PB), or The block size of a transform block (TB).

47. The method according to claim 46, wherein for non-LM intra prediction modes, if the block size is a first size, the first weight in the plurality of sets of weights is equal to a first value, and if the block size is a second size, the first weight in the plurality of sets of weights is equal to a second value.

48. The method according to any one of claims 1 to 24, wherein the plurality of sets of weights depends on the video content of the current video block.

49. The method according to claim 48, wherein for the LM intra prediction mode, if the type of the video content is a natural sequence, the second weight in the plurality of sets of weights is equal to a first value, and if the type of the video content is screen content, the second weight in the plurality of sets of weights is equal to a second value.

50. The method according to any one of claims 1 to 49, wherein the bitstream includes at least one syntax element indicating at least one of the following: whether prediction fusion is applied, how the prediction fusion is applied, or weights for the prediction fusion.

51. The method according to claim 50, wherein the at least one syntax element is indicated in one of the following: SPS, PPS, APS, picture header, or slice header.

52. The method according to any one of claims 50 to 51, wherein the at least one syntax element includes a syntax element indicating whether the prediction fusion is applied.

53. The method according to any one of claims 50 to 52, wherein the at least one syntax element includes a set of syntax elements indicating how the weights are predetermined.

54. The method according to claim 53, wherein each syntax element in the set of syntax elements indicates a weight.

55. The method according to any one of claims 1 to 54, wherein at least one of the following is indicated in the bitstream or determined based on the codec information of the current video block or the codec information of the blocks decoded before the current video block: whether the plurality of sets of weights are used, or how the plurality of sets of weights are used.

56. The method according to any one of claims 1 to 54, wherein the bitstream includes one or more group indices indicating a set of weights used in the plurality of sets of weights.

57. The method according to claim 56, wherein a single group index is used to indicate the set of weights used for multiple color components, or wherein the indication of the set of weights used for multiple color components is signaled separately.

58. The method according to claim 56, wherein the one or more group indices are determined based on the codec information of the current video block or the codec information of the blocks decoded before the current video block, or the one or more group indices are determined based on a template matching-based scheme, or the group indices for multiple color components are determined together, or the group indices for multiple color components are determined separately.

59. The method according to any one of claims 1 to 58, wherein the target prediction for the first color component is determined based on the plurality of sets of weights, the plurality of candidate predictions, and at least one offset.

60. The method according to claim 59, wherein the at least one offset is adaptive and is used to determine a model for determining the target prediction.

61. The method according to claim 60, wherein the model is linear or non-linear.

62. The method according to any one of claims 59 to 61, wherein the at least one offset is determined based on the codec information of the current video block or the codec information of the blocks coded before the current video block.

63. The method according to claim 62, wherein the codec information includes at least one of the following: Reconstructed samples of a second color component different from the first color component, Reconstructed neighboring samples of the second color component, Reconstructed neighboring samples of the first color component, or Predicted samples of the first color component.

64. The method according to any one of claims 59 to 61, wherein the at least one offset is determined from codec tools.

65. The method according to claim 64, wherein the codec tools include luminance mapping and chroma scaling (LMCS) or local illumination compensation (LIC).

66. The method according to any one of claims 59 to 61, wherein the at least one offset is determined from different coded signals using different calculation methods.

67. The method according to any one of claims 59 to 61, wherein the at least one offset is determined as the average of one or more types of signals in a block region, or wherein the at least one offset is determined as the median of one or more types of signals in the block region.

68. The method according to claim 67, wherein the one or more types of signals include at least one of the following: Template values, Predictions of a second color component different from the first color component, Predictions of the first color component, or Reconstructions of the second color component.

69. The method according to any one of claims 59 to 61, wherein the at least one offset is determined from different schemes depending on at least one of codec information, codec context, or codec background.

70. The method according to any one of claims 59 to 61, wherein the at least one offset is determined based on codec position information.

71. The method according to claim 70, wherein the at least one offset is determined based on at least one of the following: The left block of the current video block, or The upper block of the current video block.

72. The method according to any one of claims 59 to 71, wherein the at least one offset is indicated in the bitstream.

73. The method according to claim 72, wherein the at least one offset is indicated at one of the following: Block level, Sequence level, Group of pictures level, Picture level, Slice level, or Slice group level.

74. The method according to claim 72, wherein the at least one offset is indicated in one of the following: The codec structure of a coding tree unit (CTU), The codec structure of a coding unit (CU), Coding and decoding structure of a transform unit (TU) Coding and decoding structure of a prediction unit (PU) Coding and decoding structure of a coding tree block (CTB) Coding and decoding structure of a coding block (CB) Coding and decoding structure of a transform block (TB) Coding and decoding structure of a prediction block (PB) Sequence header Picture header Sequence parameter set (SPS) Video parameter set (VPS) Dependency parameter set (DPS) Decoding capability information (DCI) Picture parameter set (PPS) Adaptive parameter set (APS) Slice header, or Picture group header 75. The method according to any one of claims 59 to 74, wherein a third syntax element indicating whether prediction fusion with the at least one offset is applied is included in the bitstream.

76. The method according to any one of claims 59 to 75, wherein a fourth syntax element indicating the level of the at least one offset applied is included in the bitstream.

77. The method according to claim 75, wherein if the third syntax element indicates that the prediction fusion with the at least one offset is applied, a fourth syntax element indicating the level of the at least one offset applied is included in the bitstream.

78. The method according to any one of claims 59 to 74, wherein the bitstream includes a single syntax element indicating whether prediction fusion with the at least one offset is applied and the level of the at least one offset applied.

79. The method according to any one of claims 75 to 78, wherein the syntax element is coded and decoded using one of the following: Fixed-length coding and decoding Exponential Golomb (EG) coding and decoding Truncated unary coding and decoding, or Unary coding and decoding 80. The method according to any one of claims 75 to 78, wherein the syntax element is coded and decoded using at least one context in arithmetic coding.

81. The method according to any one of claims 75 to 78, wherein the syntax element is bypass-coded.

82. The method according to any one of claims 75 to 81, wherein if the prediction fusion with the at least one offset is allowed to be used, the syntax element associated with the prediction fusion is indicated in the bitstream.

83. The method according to any one of claims 75 to 82, wherein the syntax element is included in an SEI message or a VUI message.

84. The method according to any one of claims 59 to 83, wherein the same offset is used for multiple color components, or Different offsets are used for the multiple first color components.

85. The method according to any one of claims 59 to 74, wherein the at least one offset includes a plurality of offsets.

86. The method according to claim 85, wherein the bitstream includes at least one syntax element indicating at least one of the following: Whether prediction fusion with the plurality of offsets is applied How to apply the prediction fusion with the plurality of offsets, or The offsets for the prediction fusion.

87. The method according to claim 86, wherein the at least one syntax element is indicated in one of the following: SPS PPS APS Picture header Strip header CTU, or CU 88. The method according to any one of claims 1 to 87, wherein at least one of the following is used to determine the target prediction for the first color component: Gradient of sample values Upsampled samples of a second color component different from the first color component Downsampled samples of the second color component Neighboring samples Linear model Nonlinear model Reconstructed neighboring samples of the second color component Reconstructed neighboring samples of the first color component Position information of the current video block 89. The method according to claim 88, wherein the gradient is used in a model for determining the target prediction for the first color component.

90. The method according to any one of claims 88 to 89, wherein the gradient is determined based on at least one of the following: Reconstructed samples of the second color component Upsampled samples of the second color component Downsampled samples of the second color component, or Predicted samples of the first color component 91. The method according to any one of claims 88 to 90, wherein how the gradient is determined depends on at least one of the color format or color component of the current video block.

92. The method according to any one of claims 88 to 91, wherein the gradient includes different types of gradients.

93. The method according to claim 92, wherein the gradient includes at least one of a horizontal gradient or a vertical gradient.

94. The method according to any one of claims 88 to 93, wherein the gradient is determined based on different schemes.

95. The method according to claim 93, wherein the horizontal gradient is determined based on a linear sum, linear difference, multiplication, or division of signals along the horizontal direction, or the vertical gradient is determined based on a linear sum, linear difference, multiplication, or division of signals along the vertical direction.

96. The method according to claim 93, wherein the horizontal gradient is determined based on a scheme for deriving signals along the horizontal direction, or the vertical gradient is determined based on a scheme for deriving signals along the vertical direction.

97. The method according to any one of claims 88 to 96, wherein the gradient is determined based on different calculation shapes.

98. The method according to any one of claims 88 to 96, wherein the gradient is determined based on a square block or a diamond block covering the first sample in the current video block.

99. The method according to claim 98, wherein the gradient is determined based on the reconstructed samples of the second color component adjacent to the first sample.

100. The method according to any one of claims 94 to 99, wherein at least one of the equations for determining the gradient is modified, at least one of the equations is replaced, at least one additional term is added to the equation, or multiple terms in the equation are combined.

101. The method according to any one of claims 1 to 87, wherein a plurality of downsampling filters are allowed to be used to determine the target prediction for the first color component.

102. The method according to claim 101, wherein the plurality of downsampling filters includes a 6 - tap filter or a 3 - tap filter.

103. The method according to claim 102, wherein the 6 - tap filter has coefficients 1 - 2 - 1 - 1 - 2 - 1, or the 3 - tap filter has coefficients 1 - 2 - 1.

104. The method according to any one of claims 101 to 103, wherein the bitstream includes at least one syntax element indicating at least one of the following: Whether prediction fusion with a downsampling filter is applied, How to apply the prediction fusion with a downsampling filter, or One or more downsampling filters for the prediction fusion.

105. The method according to claim 104, wherein the at least one syntax element is included in one of the following: Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Picture header, Slice header, Coding Tree Unit (CTU), or Coding Unit (CU).

106. The method according to any one of claims 104 to 105, wherein the at least one syntax element includes a fifth syntax element indicating whether prediction fusion with a downsampling filter is applied.

107. The method according to any one of claims 104 to 106, wherein the at least one syntax element includes a sixth syntax element indicating one or more downsampling filters for the prediction fusion.

108. The method according to claim 106, wherein if the fifth syntax element indicates that prediction fusion with a downsampling filter is applied, the at least one syntax element further includes a sixth syntax element indicating one or more downsampling filters for the prediction fusion.

109. The method according to any one of claims 104 to 105, wherein the at least one syntax element includes a single syntax element indicating whether prediction fusion with a downsampling filter is applied and one or more downsampling filters for the prediction fusion.

110. The method according to any one of claims 104 to 109, wherein the at least one syntax element is coded and decoded using one of the following: Fixed - length coding and decoding, Exponential Golomb (EG) coding and decoding, Truncated unary coding and decoding, or Unary coding and decoding.

111. The method according to any one of claims 104 to 110, wherein the at least one syntax element is coded and decoded using at least one context in arithmetic coding.

112. The method according to any one of claims 104 to 110, wherein the at least one syntax element is bypass - coded.

113. The method according to any one of claims 101 to 103, wherein if prediction fusion with a downsampling filter is allowed to be used, the bitstream includes at least one syntax element indicating at least one of the following: Whether to apply predictive fusion with a downsampling filter, How to apply the predictive fusion with a downsampling filter, or One or more downsampling filters for the predictive fusion.

114. The method according to any one of claims 104 to 113, wherein the at least one syntax element is included in an SEI message or a VUI message.

115. The method according to any one of claims 1 to 87, wherein whether a non-downsampling filter is used to determine the target prediction for the first color component depends on at least one of the coding information, coding context, coding background, or video content of the current video block.

116. The method according to claim 115, wherein if the type of the video content is screen content, the non-downsampling filter is used to determine the target prediction.

117. The method according to any one of claims 1 to 116, wherein the bitstream includes at least one high-level syntax (HLS) element indicating at least one of the following: Whether to apply predictive fusion with a downsampling filter, How to apply the predictive fusion with a downsampling filter, or One or more downsampling filters for the predictive fusion.

118. The method according to claim 117, wherein the at least one HLS element is included in one of the following: Sequence parameter set (SPS), Picture parameter set (PPS), Adaptive parameter set (APS), Picture header, or Slice header.

119. The method according to any one of claims 117 to 118, wherein the at least one HLS element includes a first HLS element indicating whether to apply the predictive fusion with a downsampling filter.

120. The method according to any one of claims 117 to 119, wherein the at least one HLS element includes a set of HLS elements indicating at least one selected downsampling filter or non-downsampling filter.

121. The method according to claim 120, wherein each HLS element in the set of HLS elements indicates a downsampling filter.

122. The method according to any one of claims 1 to 121, wherein whether to apply the method and / or how to apply the method is indicated at one of the following: Block level, Sequence level, Group of pictures level, Picture level, Slice level, or Slice group level.

123. The method according to any one of claims 1 to 121, wherein whether to apply the method and / or how to apply the method is indicated in one of the following: Coding structure of CTU, Coding structure of CU, Coding structure of TU, Coding structure of PU, Coding structure of CTB, Coding structure of CB, Coding structure of TB, Coding structure of PB, Sequence header, Picture header, Sequence parameter set (SPS), Video parameter set (VPS), Dependency parameter set (DPS), Decoding capability information (DCI), Picture parameter set (PPS), Adaptive parameter set (APS), Slice header, or Slice group header.

124. The method according to any one of claims 1 to 123, wherein whether to apply the method and / or how to apply the method depends on the warp-decoded information of the current video block.

125. The method according to claim 124, wherein the warp-decoded information comprises at least one of the following: block size, color format, single-tree segmentation, dual-tree segmentation, color component, slice type, or picture type.

126. The method according to any one of claims 1 to 125, wherein the method is applicable to codec tools that require prediction fusion.

127. The method according to any one of claims 1 to 126, wherein the transformation comprises encoding the current video block into the bitstream.

128. The method according to any one of claims 1 to 126, wherein the transformation comprises decoding the current video block from the bitstream.

129. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 128.

130. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 128.

131. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by a video processing apparatus for a video, wherein the method comprises: obtaining multiple sets of weights; determining a target prediction for a first color component of the current video block of the video based on the multiple sets of weights and multiple candidate predictions for the first color component; and generating the bitstream based on the target prediction.

132. A method for storing a bitstream of a video, comprising: obtaining multiple sets of weights; determining a target prediction for a first color component of the current video block of the video based on the multiple sets of weights and multiple candidate predictions for the first color component; generating the bitstream based on the target prediction; and storing the bitstream in a non-transitory computer-readable recording medium.