Method and device for video processing and medium

By performing multiple sub-segments on video blocks and combining intra-prediction modes such as angle prediction, intra-block copying, and template matching, the problem of improving encoding and decoding efficiency and quality in existing video encoding and decoding technologies is solved, achieving more efficient encoding and decoding results.

CN120937357APending Publication Date: 2025-11-11DOUYIN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480025450.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-14
Filing Date
2024-04-12
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have room for improvement in terms of efficiency and quality, especially when processing blocks with different signal and texture distributions, traditional methods are difficult to effectively improve encoding and decoding efficiency.

Method used

The encoding and decoding efficiency and quality are improved by employing multiple sub-segments based on the current video block and at least one intra-frame prediction scheme, including angle prediction mode, intra-frame block copy mode, template matching mode or cross-component prediction mode.

Benefits of technology

By combining sub-segmentation and multiple intra-frame prediction modes, the efficiency and quality of video encoding and decoding are significantly improved, especially when blocks have different signal and texture distributions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120937357A_ABST
    Figure CN120937357A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises: for a conversion between a current video block of a video and a bitstream of the video, determining a prediction for the current video block based on a plurality of sub-partitions of the current video block and at least one intra prediction scheme, the current video block being a chroma block, the plurality of sub-partitions being obtained by partitioning the current video block, and the intra prediction scheme being obtained by partitioning the current video block; and the at least one intra prediction scheme comprises at least one of: an angle prediction mode, an intra block copy (IBC) mode, a template matching mode, or a cross-component prediction mode; and performing the conversion based on the prediction for the current video block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to block segmentation. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: a conversion between a current video block and a bitstream of the video; determining a prediction for the current video block based on multiple sub-segments of the current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: an angle prediction mode, an intra-block copy (IBC) mode, a template matching mode, or a cross-component prediction mode; and performing the conversion based on the prediction for the current video block.

[0005] Based on the method according to the first aspect of this disclosure, the sub-segmentation scheme is allowed to be used in combination with angle prediction mode, intra-block copy (IBC) mode, template matching mode, and / or cross-component prediction mode for chroma blocks. Compared with conventional solutions, the proposed method can advantageously improve encoding / decoding efficiency and quality, especially when blocks have different signal and texture distributions.

[0006] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of this disclosure.

[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: determining a prediction for a current video block based on multiple sub-segments of a current video block and at least one intra-frame prediction scheme, wherein the current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-frame prediction scheme includes at least one of the following: an angle prediction mode, an intra-block copy (IBC) mode, a template matching mode, or a cross-component prediction mode; and generating a bitstream based on the prediction for the current video block.

[0009] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: determining a prediction for a current video block based on multiple sub-segments of a current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: angular prediction mode, intra-block copy (IBC) mode, template matching mode, or cross-component prediction mode; generating a bitstream based on the prediction for the current video block; and storing the bitstream in a non-transitory computer-readable recording medium.

[0010] The purpose of this invention is to present, in a simplified form, the concept choices further described below in the detailed embodiments. This invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0011] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0012] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown; Figure 5 An example of an encoder block diagram is shown; Figure 6 67 intra-frame prediction modes are shown; Figure 7 Reference samples for wide-angle intra-frame prediction are shown; Figure 8A It shows that in the direction beyond The problem of discontinuity in certain situations; Figure 8B A luminance block is shown for deriving the direct block vector; Figure 9 The derivation is shown The location of the sample points; Figure 10A Four Sobel-based gradient modes for the gradient linear model (GLM) are shown. Figure 10B An example of classifying neighboring samples into two groups is shown; Figure 10C The effect of the slope adjustment parameter is shown; Figure 11A This is a schematic diagram showing the definition of the sample points used by the PDPC applied to the upper right diagonal mode; Figure 11B This is a schematic diagram showing the definition of the sample points used by the PDPC applied to the lower left diagonal mode; Figure 11C This is a schematic diagram illustrating the definition of the sample points used by the PDPC applied to the adjacent diagonal upper right pattern; Figure 11D This is a schematic diagram illustrating the definition of the sample points used by the PDPC applied to the adjacent diagonal lower left mode; Figure 12 A gradient method for non-vertical / non-horizontal modes is shown; Figure 13 The nScale values ​​for nTbH and mode number are shown; for all In cases where gradient methods are used; Figure 14 The flowchart shows the current PDPC (left) and the proposed PDPC (right). Figure 15 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown. Figure 16 An example of the proposed intra-frame reference mapping is shown; Figure 17 An example of four reference rows adjacent to the prediction block is shown; Figure 18A This is a schematic diagram showing examples of 4×8 and 8×4 CU sub-segments; Figure 18BThis is a schematic diagram showing examples of sub-segments of the CU other than 4×8, 8×4 and 4×4; Figure 19A The matrix-weighted intra-frame prediction process is illustrated. Figure 19B The intra-frame template matching search area used is shown; Figure 20 The target sample, template sample, and reference sample of the template used in DIMD are shown. Figure 21 The proposed intra-frame block decoding process is shown; Figure 22 The HoG calculation from a template with a width of 3 pixels is shown; Figure 23 The prediction fusion is shown by weighted averaging of two HoG modes and a plane; Figure 24 The neighbor reconstructed sample points used for the DIMD chromaticity mode are shown; Figure 25 The MMVD search point is shown; Figure 26 An illustration of the symmetric MVD mode is shown; Figure 27 The extended CU region used in BDOF is shown; Figure 28 An affine motion model based on control points is shown; Figure 29 The affine MVF for each sub-block is shown; Figure 30 The location of the inherited affine motion prediction value is shown; Figure 31 This demonstrates the inheritance of control point motion vectors; Figure 32 The locations of candidate positions for constructing the affine Merge pattern are shown; Figure 33 A diagram illustrating the motion vectors used for the proposed combination method is shown. Figure 34 The sub-block MV VSB and pixels are shown. ; Figure 35A and Figure 35B The SbTMVP procedure in VVC is shown; Figure 36 Local lighting compensation is shown; Figure 37 This indicates that no downsampling was performed on the short side; Figure 38 This shows the refinement of motion vectors on the decoding side; Figure 39The diamond-shaped area in the search region is shown; Figure 40 The locations of the spatial merge candidates are shown; Figure 41 The candidate pairs considered for redundancy checks of spatial merge candidates are shown. Figure 42 An illustration of motion vector scaling for temporal Merge candidates is shown; Figure 43 Candidate positions for temporal Merge candidates are shown. ; Figure 44 The VVC spatial neighboring blocks of the current block are shown; Figure 45 A diagram illustrating the virtual block in the i-th round of search is shown; Figure 46 An example of GPM partitioning grouped at the same angle is shown; Figure 47 The unidirectional prediction MV selection for geometric segmentation patterns is shown; Figure 48 The bending weights using the geometric segmentation pattern are shown. An example of generation; Figure 49 The spatial neighbor block used to derive spatial merge candidates is shown; Figure 50 This shows that template matching is performed on the search area surrounding the initial MV; Figure 51 A diagram of a sub-block of the OBMC application is shown; Figure 52 The location, type, and transformation type of the SBT are shown; Figure 53 The neighboring sample points used to calculate SAD are shown; Figure 54 The neighboring samples used to calculate SAD for sub-CU level motion information are shown; Figure 55 The sorting process is shown; Figure 56 The reordering process in the encoder is shown; Figure 57 The reordering process in the decoder is shown; Figure 58 The spatial portion of the convolution filter is shown; Figure 59 The reference region (and its padding) used to derive the filter coefficients is shown. Figure 60 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 61 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 62 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 63 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 64 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 65 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 66 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 67 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 68 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 69 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 70 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 71 A schematic diagram of example segmentation patterns according to some embodiments of the present disclosure is shown; Figure 72 The location of the inversion prediction pattern based on sub-segmentation is shown; Figure 73 The location of the co-position block vector derivation based on sub-segmentation is shown; Figure 74 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 75 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0013] Throughout all the accompanying figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Implementation

[0014] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0015] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0016] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0017] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0019] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0020] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0021] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0022] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0023] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0024] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0025] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0026] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0027] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0028] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.

[0029] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0030] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0031] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0032] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0033] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0034] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images containing multiple reference video blocks in lists 0 and 1, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0035] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0036] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0037] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0038] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0039] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0040] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0041] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0042] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0043] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0044] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0045] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0046] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0047] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0048] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0049] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0050] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0051] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0052] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.

[0053] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "strip" can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0054] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0055] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0056] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.

[0057] 1. Brief Overview This disclosure relates to video encoding and decoding technology. Specifically, it relates to an encoding and decoding tool that segments multiple components of a frame into blocks of different substructures and shapes to achieve better encoding and decoding efficiency. It can be applied to existing video encoding and decoding standards such as HEVC or VVC. It is also applicable to future video encoding and decoding standards or video codecs.

[0058] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.

[0059] 2.1. Color Space and Chromaticity Downsampling A color space, also known as a color model (or color system), is an abstract mathematical model that simply describes the range of colors as tuples of numbers, typically 3 or 4 values ​​or color components (such as RGB). Essentially, a color space is a development of coordinate systems and subspaces.

[0060] For video compression, the most commonly used color spaces are YCbCr and RGB.

[0061] YCbCr, Y′CbCr, or Y Pb / Cb Pr / Cr, also written as YCBCR or Y'CBCR, is a color space family used as part of the color image pipeline in video and digital photography systems. Y′ is the luminance component, and CB and CR are the blue and red difference chromaticity components. Y′ (with an apostrophe) is distinguished from Y, which is luminance, meaning that light intensity is non-linearly encoded based on gamma-corrected RGB primary colors.

[0062] Chromaticity downsampling is a practice of encoding images by applying a lower resolution to chromaticity information compared to luminance information. It takes advantage of the fact that the human visual system is less sensitive to color differences than to luminance differences. 2.1.1. 4:4:4 Each of the three Y'CbCr components has the same sampling rate, therefore there is no chromaticity downsampling. This scheme is sometimes used in high-end film scanners and film post-production. 2.1.2. 4:2:2 The two chroma components are sampled at half the luminance sampling rate: the horizontal chroma resolution is halved, while the vertical chroma resolution remains unchanged. This reduces the bandwidth of the uncompressed video signal by one-third with almost no visual difference. An example of the nominal vertical and horizontal positions of the 4:2:2 color format is shown in the VVC working draft. Figure 4 It is depicted in the middle. Figure 4 The nominal vertical and horizontal positions of the 4:2:2 luminance and chrominance samples in the image are shown. 2.1.3. 4:2:0 In 4:2:0, the horizontal sampling is doubled compared to 4:1:1, but the vertical resolution is halved because the Cb and Cr channels are sampled only on each alternating row. Therefore, the data rate is the same. Cb and Cr are each downsampled horizontally and vertically by a factor of 2. There are three variations of the 4:2:0 scheme with different horizontal and vertical positions.

[0066] • In MPEG-2, Cb and Cr are set horizontally in the same position. Cb and Cr are set between pixels in the vertical direction (with intervals).

[0067] • In JPEG / JFIF, H.261, and MPEG-1, Cb and Cr are set at intervals, in the middle of alternating luminance samples.

[0068] • In a 4:2:0 DV, Cb and Cr are co-located in the horizontal direction. In the vertical direction, they are co-located on alternating rows.

[0069] Table 1. SubWidthC and SubHeightC values ​​derived from chroma_format_idc and separate_colour_plane_flag.

[0070] 2.2. Encoding Flow of a Typical Video Codec Figure 5 An example of an encoder block diagram is shown. Figure 5An example of a VVC encoder block diagram is shown, comprising three loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Compensation (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding an offset and by applying a Finite Impulse Response (FIR) filter, respectively, and by utilizing the encoded / decoded side information through signal transmission offset and filter coefficients. ALF is located in the final processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.

[0071] 2.3. Intra-mode encoding and decoding with 67 intra-prediction modes Figure 6 Sixty-seven intra-frame prediction modes are shown. This is to capture arbitrary edge directions presented in natural video, such as... Figure 6 As shown, the number of oriented intra-prediction modes has been expanded from 33 used in HEVC to 65, while the planar and DC modes remain unchanged. These denser oriented intra-prediction modes are applicable to all block sizes and both luma and chroma intra-prediction.

[0072] In HEVC, each intra-codec block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is needed to generate intra-prediction values ​​using DC mode. In VVC, blocks can have rectangular shapes, which generally requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks.

[0073] 2.3.1. Wide-angle intra-frame prediction Although 67 modes are defined in VVC, the precise prediction direction for a given intra-prediction mode index depends on the block shape. Regular angular intra-prediction directions are defined clockwise from 45 degrees to -135 degrees. In VVC, for non-square blocks, several regular angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are transmitted via signaling using the original mode index, which is then remapped to the wide-angle mode index after resolution. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding / decoding method remains unchanged.

[0074] Figure 7 Reference samples are shown for wide-angle intra-frame prediction. To support these prediction directions, such as... Figure 7 As shown, a top reference with a length of 2W+1 and a left reference with a length of 2H+1 are defined.

[0075] The number of modes replaced in the wide-angle directional mode depends on the block aspect ratio. The replaced intra-prediction modes are shown in Table 2.

[0076] Table 2 Intra-prediction modes replaced by wide-angle mode

[0077] Figure 8A It shows that in the direction beyond The problem of discontinuity in certain situations. For example, in Figure 8A As shown, in the case of wide-angle intra-frame prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap. The negative impact of wide-angle mode. If the wide-angle mode represents a non-fractional offset. There are 8 wide-angle modes that satisfy this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted through these modes, the samples in the reference cache are directly copied without applying any interpolation. With this modification, the number of samples required for smoothing is reduced. In addition, it aligns the design of non-fractional modes in regular prediction modes with that of wide-angle modes.

[0078] In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, with the number of entries expanded from 35 to 67 to align with the expansion of intra-prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra-prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​in the mapping table entries to more accurately convert the prediction angle for chroma blocks.

[0079] 2.4. Intra-prediction mode encoding and decoding for chroma components For the chroma component of the intra-frame prediction unit (PU), the encoder selects the optimal chroma prediction mode from five modes: planar, DC, horizontal, vertical, and a direct copy of the intra-frame prediction mode for the luma component. The mapping between the intra-frame prediction direction for chroma and the intra-frame prediction mode number is shown in Table 3.

[0080] When the intra-prediction mode number of the chroma component is 4, the intra-prediction direction of the luma component is used for the generation of intra-prediction samples of the chroma component. When the intra-prediction mode number of the chroma component is not 4 and is the same as the intra-prediction mode number of the luma component, the intra-prediction direction 66 is used for the generation of intra-prediction samples of the chroma component.

[0081] 2.5. Inter-frame prediction For each inter-frame prediction CU, motion parameters consist of a motion vector, a reference picture index, a reference picture list usage index, and additional information required for the new encoding / decoding features of the VVC generated from the inter-frame prediction samples. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and does not have significant residual coefficients, encoded motion vector increments, or reference picture indices. A Merge mode is specified, whereby motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates and additional scheduling introduced in the VVC. The Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where for each CU, the motion vector, the corresponding reference picture index for each reference picture list, the reference picture list usage flag, and other required information are explicitly transmitted via signaling.

[0082] 2.6. Intra-Block Copy (IBC) Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well known to significantly improve the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer precise. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-encoded CUs are considered a third prediction mode, distinct from intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.

[0083] On the encoder side, hash-based motion estimation for IBC is performed. The encoder performs RD checks on blocks with a width or height no greater than 16 luminance samples. For non-Merge mode, block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a local search based on block matching is performed.

[0084] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4×4 sub-blocks. For the larger current block, a hash key match with a reference block is determined when all hash keys of all 4×4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the lowest cost is selected.

[0085] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.

[0086] At the CU level, the IBC mode is transmitted via signaling using a flag, and it can be transmitted via signaling as IBCAMVP mode or IBC skip / Merge mode, as follows: - IBC Skip / Merge Mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and paired candidates.

[0087] - IBC AMVP Mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as prediction values, one from the left nearest neighbor and one from the top nearest neighbor (if IBC encoded and decoded). When either nearest neighbor is unavailable, the default block vector is used as the prediction value. A flag is transmitted via signaling to indicate the index of the block vector prediction value.

[0088] 2.6.1 Direct block vectors for chroma blocks Direct block vectors are used for chroma blocks in a dual-tree stripe. When the chroma dual-tree is activated, a flag is transmitted via signaling to indicate whether the chroma blocks are encoded or decoded using IBC mode. Figure 8B If one of the luma blocks in the five locations shown is encoded or decoded in IBC or intraTMP mode, its block vector is scaled and used as the block vector for the chroma block. Template matching is used to perform the block vector scaling.

[0089] 2.7. Prediction using a cross-component linear model To reduce cross-component redundancy, a cross-component linear model (CCLM) prediction mode is used in VVC. For this CCLM prediction mode, chroma samples are predicted based on reconstructed luminance samples from the same CU using the following linear model: (2-1) in This represents the predicted chromaticity samples in the CU, and This represents the reconstructed brightness sample points from the same CU (Cubic Unit).

[0090] The CCLM parameters (α and β) are derived using at most four neighboring chroma samples and their corresponding downsampled luminance samples. Assuming the current chroma block dimension is W×H, then W' and H' are set as follows: – When applying the LM pattern, W' = W, H' = H; – When applying LM_T mode, W' = W + H; – When applying LM_L mode, H' = H + W.

[0091] The upper neighboring positions are denoted as S[0, -1]…S[W'-1, -1], and the left neighboring positions are denoted as S[-1, 0]…S[-1, H'-1]. Then, four sample points are selected as follows: – When LM mode is applied and both the upper neighbor sample and the left neighbor sample are available, S[W' / 4, -1], S[3 W' / 4,-1],S[-1,H' / 4],S[-1,3 H' / 4]; – When applying LM_T mode or when only the upper neighboring sample is available, S[W' / 8, -1], S[3 W' / 8,-1],S[5 W' / 8,-1],S[7 W' / 8,-1]; – When applying LM_L mode or when only the left neighboring sample is available, S[-1, H' / 8], S[-1, 3 H' / 8],S[-1,5 H' / 8],S[-1,7 H' / 8].

[0092] Four neighboring brightness samples at the selected location are downsampled and compared four times to find the two larger values: and two smaller values: Their corresponding chromaticity sample values ​​are represented as .but It is deduced as:

[0093] Finally, the linear model parameters α and β are obtained according to the following equation.

[0094]

[0095] Figure 9 This shows an example of the positions of the left and top samples, as well as the sample of the current block, involved in CCLM mode. Figure 9 The locations of the sample points used to derive α and β are shown.

[0096] The division operation for calculating the parameter α is implemented using a lookup table. To reduce the memory required to store this table, diff The value (the difference between the maximum and minimum values) and the parameter α are represented using exponential notation. For example, diff The 4-bit significant part and the exponent are approximated. Therefore, for the 16 values ​​of the significant number, the table for 1 / diff is reduced to 16 elements, as follows:

[0097] This will help reduce computational complexity and the memory size required to store the necessary tables.

[0098] In addition to the fact that the top and left templates can be used together to calculate the coefficients of the linear model, they can also be used alternately in the other two LM modes (called LM_T and LM_L modes).

[0099] In LM_T mode, only the upper template is used to calculate the linear model coefficients. To obtain more samples, the upper template is expanded to (W+H) samples. In LM_L mode, only the left template is used to calculate the linear model coefficients. To obtain more samples, the left template is expanded to (H+W) samples.

[0100] In LM mode, the left and top templates are used to calculate the coefficients of the linear model.

[0101] To match the chroma sample locations of a 4:2:0 video sequence, two types of downsampling filters are applied to the luminance samples to achieve a 2:1 downsampling ratio in both the horizontal and vertical directions. The selection of the downsampling filters is specified by the SPS level flag. The two downsampling filters are as follows, corresponding to "Type 0" and "Type 2" content, respectively.

[0102]

[0103] Note that when the upper reference line is at the CTU boundary, only one luminance line (the general line buffer in intra-frame prediction) is used to generate downsampled luminance samples.

[0104] This parameter calculation is performed as part of the decoding process, not just as part of the encoder search operation. Therefore, no syntax is used to pass the α and β values ​​to the decoder.

[0105] For chroma intra-mode encoding and decoding, a total of eight intra-modes are permitted. These modes include five regular intra-modes and three cross-component linear model modes (LM, LM_T, and LM_L). The chroma modes are shown in Table 3 through signal transmission and derivation processes. Chroma mode encoding and decoding directly depends on the intra-prediction mode of the corresponding luma block. Since separate block partitioning structures are enabled for luma and chroma components in I-strips, one chroma block can correspond to multiple luma blocks. Therefore, for chroma DM mode, the intra-prediction mode of the corresponding luma block covering the center position of the current chroma block is directly inherited.

[0106] Table 3. Derivation of chromaticity prediction mode from luminance mode when CCLM is enabled.

[0107] As shown in Table 4, a single binarization table is used regardless of the value of sps_cclm_enabled_flag.

[0108] Table 4. Unified Binarization Table for Colorimetric Prediction Modes

[0109] In Table 4, the first bit indicates whether it is in normal mode (0) or LM mode (1). If it is LM mode, the next bit indicates whether it is LM_CHROMA (0). If it is not LM_CHROMA, the next bit indicates whether it is LM_L (0) or LM_T (1). In this case, when sps_cclm_enabled_flag is 0, the first bit of the binarization table for the corresponding intra_chroma_pred_mode can be discarded before entropy encoding / decoding. Or, in other words, the first bit is presumed to be 0 and therefore not encoded / decoded. This single binarization table is used for both cases where sps_cclm_enabled_flag equals 0 and 1. The first two bits in Table 4 are context-encoded using their own context model, and the remaining bits are bypassed.

[0110] Furthermore, to reduce luma-chroma latency in dual-tree systems, when a 64×64 luma codec tree node is split using Not Split (and ISP is not used for 64×64 CU) or QT, the chroma CU in a 32×32 / 32×16 chroma codec tree node is allowed to use CCLM in the following manner: – If the 32×32 chroma node is not partitioned or is partitioned by QT, then all chroma CUs in the 32×32 node can use CCLM.

[0111] – If a 32×32 chroma node is divided by a horizontal BT, and the 32×16 child node is not divided or is divided by a vertical BT, then all chroma CUs in the 32×16 chroma node can use CCLM.

[0112] CCLM is not allowed for the chroma CU under all other luma and chroma codec tree partitioning conditions.

[0113] 2.8. Gradient Linear Model (GLM) Compared to CCLM, GLM uses the gradient of luminance samples to derive the linear model, rather than downsampling luminance values. In other words, in the CCLM process, the gradient... The low-pass filter brightness samples were replaced. Other aspects of the CCLM design (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged.

[0114]

[0115] Figure 10A It shows It can be computed using one of four Sobel-based gradient modes.

[0116] For signaling, when CCLM mode is enabled for the current CU, two flags are transmitted separately for the Cb / Cr components to indicate whether GLM is enabled for that component; if GLM is enabled for a component, a syntax element is further transmitted via signaling to select... Figure 10A One of the four gradient modes is used for gradient calculation.

[0117] 2.9. Multimodel Linear Model (MMLM) Using MMLM, there can be more than one linear model between luma and chroma samples in the CU. In this method, the neighboring luma and chroma samples of the current block are classified into several groups, and each group is used as a training set to derive the linear model (i.e., specific α and β are derived for a specific group). In addition, the samples of the current luma block are also classified according to the same rules as the classification of the neighboring luma samples.

[0118] Neighboring samples can be classified into M groups, where M is 2 or 3. In addition to the original LM mode, the MMLM method with M=2 and M=3 is designed with two additional chromaticity prediction modes, called MMLM2 and MMLM3. The encoder selects the optimal mode during the RDO process and transmits that mode via signal transmission.

[0119] When M equals 2 Figure 10BAn example of classifying neighboring samples into two groups is shown. The threshold is calculated as the average value of the neighboring reconstructed brightness samples. Neighboring samples with Rec'L[x,y] <= the threshold are classified into group 1; while neighboring samples with Rec'L[x,y] > the threshold are classified into group 2. Similar to CCLM, MMLM has three modes: MMLM, MMLM_T, and MMLM_L. The two models are derived as follows.

[0120] (2-8) Figure 10B An example of classifying neighboring samples into two groups is shown.

[0121] The threshold is the average of the neighboring samples in the brightness reconstruction. The linear model for each category is derived using the least mean square (LMS) method (if enabled) or the minimum / maximum method of VVC.

[0122] 2.10. Slope Adjustment of CCLM CCLM uses a two-parameter model to map luminance values ​​to chrominance values. The slope parameter "a" and the bias parameter "b" define the mapping as follows:

[0123] The slope parameter "u" is adjusted via signal transmission to update the model in the following form:

[0124] in

[0125] Using this selection, the mapping function revolves around the brightness value. The points are tilted or rotated. The average value of the reference brightness samples used in model creation is used as... This is to provide meaningful modifications to the model. The process is in... Figure 10C The effect of the slope adjustment parameter "u" is shown in the figure. Figure 10C The left subplot shows the model created using the current CCLM, and Figure 10C The right-hand subgraph in the figure shows the updated model as proposed.

[0126] Implementation The slope adjustment parameter is provided as an integer between -4 and 4 (inclusive) and is transmitted via signal in the bitstream. The unit of the slope adjustment parameter is 1 / 8 of the chroma sample value per luminance sample value (for 10-bit content).

[0127] The CCLM model (“LM_CHROMA_IDX” and “MMLM_CHROMA_IDX”) can be adjusted to use reference samples from both the top and left sides of the block, but not for “one-sided” mode. This choice is based on a trade-off between encoding / decoding efficiency and complexity.

[0128] When slope adjustment is applied to a multi-mode CCLM model, both models can be adjusted, so at most two slope updates are transmitted via signal for a single chroma block.

[0129] Encoder method The encoder method performs a SATD-based search for the optimal slope update for Cr and a similar SATD-based search for Cb. If either results in a non-zero slope adjustment parameter, the combined slope adjustment pair (SATD-based update for Cr, SATD-based update for Cb) is included in the list of RD checks for TU.

[0130] 2.11. Improvements to Color Blending This method proposes a novel chroma fusion approach where weights are derived from the neighboring templates of both the reconstructed luminance samples and the predicted chroma samples obtained by applying a non-LM mode. The derivation is based on the LDL decomposition method used in CCCM.

[0131] The color blending method is as follows:

[0132] in This represents the final predicted chromaticity sample points. This represents the reconstructed brightness sample points. This represents the predicted chromaticity samples obtained by applying a non-LM mode. They are labeled as... of This represents the average pixel value of the reference region. Model parameters. It is derived from the LDL decomposition method used in CCCM, based on the same adjacent sample points (two rows).

[0133] As shown in Table 5, there are three methods for color blending.

[0134] Table 5. Color blending modes in the method

[0135] 2.12. Location-dependent intra-frame prediction combination In VVC, the results of intra-prediction for DC, planar, and several angle modes are further modified using the Position-Dependent Intra-Prediction Combination (PDPC) method. PDPC is an intra-prediction method that calls a combination of boundary reference samples and HEVC-style intra-prediction with filtered boundary reference samples. PDPC is applied to the following intra-prediction modes without signal transmission: planar, DC, intra-prediction angles less than or equal to horizontal, and intra-prediction angles greater than or equal to vertical and less than or equal to 80 degrees. PDPC is not applied if the current block is in BDPCM mode or the MRL index is greater than 0.

[0136] Based on Equation 2-8 below, predict the sample points. Predicted using a linear combination of intra-frame prediction modes (DC, plane, angle) and reference samples:

[0137] in These represent the locations of the current sample points ( x , y Reference points at the top and left boundaries of ).

[0138] If PDPC is applied to DC, planar, horizontal, and vertical intra-frame modes, no additional boundary filters are required, as are those required in the case of HEVC DC mode boundary filters or horizontal / vertical mode edge filters. The PDPC process is the same for DC and planar modes. For angular modes, if the current angular mode is HOR_IDX or VER_IDX, the left or top reference samples are not used, respectively. PDPC weights and scaling factors depend on the prediction mode and block size. PDPC is applied to blocks with both width and height greater than or equal to 4.

[0139] Figure 11A and 8B Reference samples of PDPC applied to various prediction modes are shown. R x,-1 and R -1,y Definition of ). Predicted sample points. pred ( x’ , y’ () located within the prediction block x’ , y’ ( ) at this location. As an example, refer to the sample point. R x,-1 coordinates x Given by the following formula: x = x’ + y’ +1, and refer to the sample points. R -1,y coordinates ySimilarly, it is given by the following formula: y = x’ + y’ +1 (for diagonal mode). For other angle modes, refer to the sample points. R x,-1 and R -1,y It can be located at a fractional sample point position. In this case, the sample value at the nearest integer sample point position is used.

[0140] Figure 11A-11D The definition of the sample points used by the PDPC applied to the intra-frame modes of diagonal and adjacent angles is shown. Figure 11A The top-right diagonal pattern is shown. Figure 11B The bottom left diagonal pattern is shown. Figure 11C The adjacent diagonal top right pattern is shown. Figure 11D The adjacent diagonal bottom left pattern is shown.

[0141] 2.13. Gradient PDPC like Figure 12 As shown, the gradient-based method is extended for non-vertical / non-horizontal modes. Here, the gradient is calculated as r(-1, y) – r(-1 + d, -1), where d is the horizontal displacement depending on the angular direction. Several points need to be noted here.

[0142] The gradient term r(-1, y) – r(-1+ d, -1) needs to be computed once for each row because it does not depend on the x position.

[0143] The calculation of d is already part of the original intra-frame prediction process that can be reused, so there is no need to calculate d separately. Therefore, d has 1 / 32 pixel precision.

[0144] When d is at the fractional position, we use a two-tap (linear) filter. That is, if dPos is a displacement with a precision of 1 / 32 pixel, dInt is the integer part (rounded down) (dPos >> 5), and dFract is the fractional part with a precision of 1 / 32 pixel (dPos & 31), then r(-1+d) is calculated as: r(-1+d) = (32 – dFrac) r(-1+dInt) + dFrac r(-1+dInt+1).

[0145] As explained in section a, the two-tap filter is performed once for each row (if necessary).

[0146] Finally, the predicted signal is calculated:

[0147] in ,and This is the same as the vertical / horizontal mode. In short, the same process is applied as with the vertical / horizontal mode (in fact, d=0 indicates the vertical / horizontal mode). Figure 9 A gradient method for non-vertical / non-horizontal modes is shown.

[0148] Secondly, when (nScale < 0) or when PDPC cannot be applied due to the unavailability of secondary reference samples, the gradient-based method for non-vertical / non-horizontal modes is activated. We have already... Figure 13 The nScale values ​​for TB size and angle mode are shown to better visualize the gradient method being used. Additionally, flowcharts of the current PDPC and the proposed PDPC are presented.

[0149] Figure 13 The nScale values ​​for nTbH and mode number are shown; for all cases where nScale < 0, the gradient method is used. Figure 14 The flowchart shows the current PDPC (left) and the proposed PDPC (right).

[0150] 2.14. Secondary MPM A secondary MPM list is introduced. The existing primary MPM (PMPM) list consists of 6 entries, and the secondary MPM (SMPM) list includes 16 entries. First, a general MPM list with 22 entries is constructed. Then, the first 6 entries from this general MPM list are included in the PMPM list, and the remaining entries form the SMPM list. The first entry in the general MPM list is a flat pattern. (Example: ...) Figure 15 As shown, the remaining entries consist of the intra-frame modes of the left (L), top (A), bottom left (BL), top right (AR), and top left (AL) neighboring blocks, the directional mode with an offset added from the first two available directional modes of the neighboring blocks, and the default mode.

[0151] If the CU block is vertically oriented, the order of the neighboring blocks is A, L, BL, AR, AL; otherwise, it is L, A, BL, AR, AL. Figure 15 The neighboring blocks (L, A, BL, AR, AL) used in the derivation of the general MPM list are shown.

[0152] The PMPM flag is parsed first. If it is equal to 1, the PMPM index is parsed to determine which entry in the PMPM list is selected. Otherwise, the SPMPM flag is parsed to determine whether to parse the SMPM index or the remaining patterns.

[0153] 2.15. 6-Tap Intraframe Interpolation Filter To improve prediction accuracy, a 6-tap cubic interpolation filter is proposed to replace the 4-tap interpolation filter. The filter coefficients are derived based on the same polynomial regression model, but the polynomial order is 6.

[0154] The filter coefficients are listed below: {0,0, 256,0,0,0}, / / 0 / 32 position {0,-4, 253,9,-2,0}, / / position 1 / 32 {1,-7, 249,17,-4,0}, / / 2 / 32 position {1, -10, 245,25,-6,1}, / / position 3 / 32 {1, -13, 241,34,-8,1}, / / position 4 / 32 {2, -16, 235,44, -10,1}, / / position 5 / 32 {2, -18, 229,53, -12,2}, / / position 6 / 32 {2, -20, 223,63, -14,2}, / / 7 / 32 position {2, -22, 217,72, -15,2}, / / 8 / 32 position {3, -23, 209,82, -17,2}, / / position 9 / 32 {3, -24, 202,92, -19,2}, / / Position 10 / 32 {3, -25, 194, 101, -20,3}, / / Position 11 / 32 {3, -25, 185, 111, -21,3}, / / Position 12 / 32 {3, -26, 178, 121, -23,3}, / / Position 13 / 32 {3, -25, 168, 131, -24,3}, / / Position 14 / 32 {3, -25, 159, 141, -25,3}, / / Position 15 / 32 {3, -25, 150, 150, -25,3} / / Half-pixel position The reference samples used for interpolation come from reconstructed samples or, as in HEVC, filled samples, so there is no need to check the availability of reference samples.

[0155] A 4-tap cubic interpolation filter is proposed to replace the nearest-nearest-round operation in deriving the extended intra-frame reference samples. For example, in... Figure 16 As shown in the example, a four-tap interpolation filter is used to derive the value of the reference sample P, whereas in JEM-3.0 or HM, P is directly set to X1.

[0156] Figure 16 An example of the proposed intra-frame reference mapping is shown.

[0157] 2.16. Multi-reference line (MRL) intra-frame prediction Multi-reference line (MRL) intra-prediction uses more reference lines for intra-prediction. Figure 17 The example depicts four reference lines, where the samples for segments A and F are not obtained from reconstructed neighboring samples, but are instead filled with the nearest samples from segments B and E, respectively. HEVC intra-frame image prediction uses the nearest reference line (i.e., reference line 0). In MRL, two additional lines (reference line 1 and reference line 2) are used.

[0158] The index (mrl_idx) of the selected reference line is transmitted via signaling and used to generate the intra-prediction value. For reference line indices greater than 0, only the additional reference line mode is included in the MPM list, and only the MPM index is transmitted via signaling without the remaining modes. The reference line index is transmitted via signaling before the intra-prediction modes, and in the case of transmitting a non-zero reference line index via signaling, the planar mode is excluded from the intra-prediction modes. Figure 17 An example of four reference rows adjacent to the prediction block is shown.

[0159] MRL is disabled for the first row of blocks within the CTU to prevent the use of extended reference samples outside the current CTU row. Additionally, PDPC is disabled when additional rows are used. For MRL mode, the derivation of the DC value in the intra-frame prediction mode for a non-zero reference row index is aligned with the derivation for reference row index 0. MRL requires storing three luma reference rows with CTU neighbors to generate predictions. The Cross-Component Linear Model (CCLM) tool also requires three neighboring luma reference rows for its downsampling filters. The definition of MRL using the same three rows is aligned with CCLM to reduce storage requirements for the decoder.

[0160] 2.17. Intra-Frame Sub-Segmentation (ISP) Intra-frame sub-segmentation (ISP) divides the luma intra-prediction block vertically or horizontally into 2 or 4 sub-segments based on the block size. For example, the minimum block size for ISP is 4×8 (or 8×4). If the block size is larger than 4×8 (or 8×4), the corresponding block is divided into 4 sub-segments. It has already been noted that... ( )and ( ) ISP blocks may be affected VDPUs pose potential problems. For example, in the case of a single tree. CU has Brightness TB and two corresponding Chromaticity TB. If the CU uses an ISP, the luminance TB will be divided into four. TB (horizontal partitioning is possible only), each TB is less than However, in the current ISP design, chroma blocks are not divided. Therefore, both chroma components will have a value greater than [missing value]. Block size. Similarly, using ISP. CU can cause a similar situation. Therefore, these two situations are for... The decoder pipeline is problematic. For this reason, the CU size usable with an ISP is limited to a maximum. . Figure 18A and Figure 18B Examples of two possibilities are shown. All sub-segments satisfy the condition of having at least 16 samples.

[0161] In ISP, the dependency of 1×N / 2×N sub-block prediction on the reconstructed values ​​of previously decoded 1×N / 2×N sub-blocks of the codec block is not allowed, making the minimum prediction width for a sub-block 4 samples. For example, an 8×N (N>4) codec block encoded and decoded using an ISP with vertical partitioning is divided into two prediction regions of size 4×N and four transforms of size 2×N. Furthermore, a 4×N codec block encoded and decoded using an ISP with vertical partitioning is predicted using the entire 4×N block; four 1×N transforms are used. Although 1×N and 2×N transform sizes are allowed, it is asserted that the transforms of these blocks within a 4×N region can be performed in parallel. For example, when a 4×N prediction region contains four 1×N transforms, there are no transforms in the horizontal direction; the transform in the vertical direction can be performed as a single 4×N transform in the vertical direction. Similarly, when a 4×N prediction region contains two 2×N transform blocks, the transform operations of the two 2×N blocks in each direction (horizontal and vertical) can be performed in parallel. Therefore, processing these smaller blocks does not increase latency compared to processing intra-frame blocks of regular 4×4 encoding and decoding.

[0162] Figure 18A and Figure 15 B shows a sub-segmentation that depends on the block size. Figure 18A Examples of sub-segments for 4×8 and 8×4 CUs are shown. Figure 18B Examples of sub-segments of the CU other than 4×8, 8×4, and 4×4 are shown.

[0163] Table 6 Entropy Encoding / Decoding Coefficient Group Dimensions

[0164] For each subsegment, reconstructed samples are obtained by adding the residual signal to the prediction signal. Here, the residual signal is generated through processes such as entropy decoding, inverse quantization, and inverse transform. Therefore, the reconstructed sample values ​​of each subsegment can be used to generate the prediction for the next subsegment, and each subsegment is processed repeatedly. Furthermore, the first subsegment to be processed is the one containing the upper left sample of the CU, and then processing continues downwards (horizontal division) or to the right (vertical division). As a result, the reference samples used to generate the subsegment prediction signal are only located to the left and top of the row. All subsegments share the same intra-frame mode. The following is a summary of the interactions between the ISP and other codec tools.

[0165] – Multiple Reference Line (MRL): If a block has an MRL index other than 0, the ISP encoding / decoding mode will be presumed to be 0, and therefore the ISP mode information will not be sent to the decoder.

[0166] – Entropy Encoding Coefficient Group Size: The size of the entropy encoding sub-blocks has been modified so that they have 16 samples in all possible cases, as shown in Table 6. Note that the new size only affects blocks generated by the ISP with one dimension less than 4 samples. In all other cases, the coefficient group remains unchanged. Dimension.

[0167] – CBF encoding / decoding: Assume that at least one sub-segment has a non-zero CBF. Therefore, if It is the number of sub-segments, and the first... If the number of sub-segments has already generated zero CBF, then the number of sub-segments is... The CBF of each sub-segment is presumed to be 1.

[0168] – Transform size limitation: All ISP transforms with a length greater than 16 points use DCT-II.

[0169] – MTS Flag: If the CU uses ISP codec mode, the MTS CU flag will be set to 0, and it will not be sent to the decoder. Therefore, the encoder will not perform RD tests for each different available transform in the resulting sub-segment. Instead, the transform selection for ISP mode will be fixed and selected based on the intra-frame mode utilized, the processing order, and the block size. Therefore, no signaling is required. For example, suppose... Each is for w×h The horizontal and vertical transformations selected for sub-segmentation, where w It is the width, and h It is the height. The transformation is then selected according to the following rules:

[0170] In ISP mode, all 67 intra-prediction modes are allowed. PDPC is also applied if the corresponding width and height are at least 4 sample lengths. Furthermore, the conditions for reference sample filtering (reference smoothing) and intra-interpolation filter selection no longer exist, and the cubic (DCT-IF) filter is always applied to fractional position interpolation in ISP mode.

[0171] 2.18. Matrix-Weighted Intra-Frame Prediction (MIP) Matrix-weighted intra-prediction (MIP) is a new intra-prediction technique added to VVC. To predict samples of a rectangular block of width W and height H, MIP uses H reconstructed neighboring boundary samples from the left row of the block and W reconstructed neighboring boundary samples from the top row of the block as input. If reconstructed samples are unavailable, they are generated in the same manner as regular intra-prediction. The generation of the predicted signal is based on three steps: averaging, matrix-vector multiplication, and linear interpolation, as follows: Figure 19A As shown. Figure 19A The matrix-weighted intra-frame prediction process is illustrated.

[0172] 2.18.1. Calculate the average of neighboring samples. In the boundary sampling points, either four or eight sampling points are selected by averaging based on the block size and shape. Specifically, the boundary is input by averaging neighboring boundary sampling points according to predefined rules depending on the block size. and Shrinking to a smaller boundary Then, the two narrowing boundaries spliced ​​to the reduced boundary vector Therefore, for shapes of For blocks of any shape, the size of the reduced boundary vector is 4, while for all other block shapes, the size of the reduced boundary vector is 8. If If it refers to MIP mode, then the splicing is defined as follows:

[0173] 2.18.2. Matrix Multiplication Using averaged samples as input, matrix-vector multiplication is performed, followed by the addition of an offset. The result is a scaled-down predicted signal on a downsampled set of samples from the original block. From the scaled-down input vector... In the middle, the reduced prediction signal The reduced prediction signal is generated with a width of And the height is The signal on the downsampled block. Here, and Defined as:

[0174] Reduced prediction signal It is calculated by multiplying the matrix and vector and adding an offset:

[0175] Here, if Then the matrix have The matrix has 4 rows and 4 columns, and in all other cases... It has 8 columns. The size is Vectors. Matrices and offset vector From the set One of them is retrieved. Index Defined as follows:

[0176] Here, each coefficient of matrix A is represented with 8 bits of precision. (Set) Composed of 16 matrices and 16 offset vectors The set consists of matrices, each with 16 rows and 4 columns, and each offset vector with a size of 16. The matrices and offset vectors in this set are used for a set of matrices with a size of 16. The block. Set Composed of 8 matrices and 8 offset vectors Composition, each matrix has It has 8 rows and 8 columns, and each offset vector has a size of 16. (Set) Composed of 6 matrices and 6 offset vectors of size 64 The matrix is ​​composed of 64 rows and 8 columns.

[0177] 2.18.3. Interpolation The predicted signals at the remaining locations are generated from the predicted signals on the downsampled set through linear interpolation, which is a single-step linear interpolation in each direction. The interpolation is first performed in the horizontal direction and then in the vertical direction, regardless of the block shape or block size.

[0178] 2.18.4. Signaling in MIP mode and coordination with other codec tools For each codec unit (CU) in intra-frame mode, a flag indicating whether MIP mode will be applied is sent. If MIP mode will be applied, then MIP mode... It is transmitted via signal. For MIP mode, a transpose flag determines whether the mode is transposed. And determine which matrix's MIP pattern ID will be used for a given MIP pattern. The following is derived.

[0179]

[0180] The MIP codec mode coordinates with other codec tools by taking into account the following aspects: – Enable LFNST for MIPs on large blocks. Here, the LFNST transform for planar mode is used.

[0181] – The reference sample derivation for MIP is performed in exactly the same way as the reference sample derivation for conventional intra-prediction modes.

[0182] – For the upsampling step used in MIP prediction, the original reference sample is used instead of the downsampling reference sample.

[0183] - The clipping is applied before upsampling, not after upsampling.

[0184] – Regardless of the maximum transform size, MIP is allowed up to 64×64.

[0185] For sizeId=0, the number of MIP patterns is 32; for sizeId=1, the number of MIP patterns is 16; and for sizeId=2, the number of MIP patterns is 12.

[0186] 2.18.5. Intra-frame template matching Intra-Template Matching Prediction (IntraTMP) is a special intra-prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signal transmission, and the same prediction operation is performed on the decoder side.

[0187] The prediction signal is obtained by comparing the L-shaped causal nearest neighbors of the current block with... Figure 19B It is generated by matching another block in a predefined search region, which consists of the following parts: R1: Current CTU; R2: Top left CTU; R3: Above CTU; R4: Left CTU.

[0188] The sum of absolute differences (SAD) is used as the cost function.

[0189] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.

[0190] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons. That is:

[0191] in' ' is a constant that controls the trade-off between gain and complexity. In practice, ' 'Equals 5.'

[0192] Enable intra-frame template matching for CUs with width and height dimensions less than or equal to 64. The maximum CU size for intra-frame template matching is configurable.

[0193] When DIMD is not used in the current CU, the intra-template matching prediction mode is transmitted at the CU level via a dedicated flag.

[0194] 2.19. Derivation of Intra-Frame Mode on the Decoder Side In JEM-2.0, the number of intra-frame modes has been expanded from 35 in HEVC to 67, and these modes are derived at the encoder and explicitly signaled to the decoder. A significant amount of overhead is incurred in intra-frame mode encoding and decoding in JEM-2.0. For example, in all intra-frame codec configurations, intra-frame mode signaling overhead can reach 5–10% of the total bit rate. This paper proposes a decoder-side intra-frame mode derivation method to reduce intra-frame mode encoding and decoding overhead while maintaining prediction accuracy.

[0195] To reduce the overhead of intra-frame mode signaling, this paper proposes a decoder-side intra-frame mode derivation (DIMD) method. In the proposed method, instead of explicitly transmitting the intra-frame mode via signaling, this information is derived from the neighboring reconstructed samples of the current block at both the encoder and decoder. The intra-frame mode derived via DIMD is used in two ways: 1) For a 2N×2N CU, when the corresponding CU level DIMD flag is enabled, the DIMD mode is used as the intra-frame mode for intra-frame prediction. 2) For N×N CU, the DIMD mode is used to replace a candidate in the existing MPM list to improve the efficiency of intra-mode encoding and decoding.

[0196] 2.19.1. Template-based intra-frame mode derivation like Figure 20 As shown, the target represents the current block (block size N) for which the intra-prediction mode is to be estimated. The template (made by...) Figure 20 The pattern area indicator in the template specifies a set of reconstructed samples used to derive the intra-frame pattern. The template size is represented as the number of samples extending above and to the left of the target block within the template, i.e., L. In the current implementation, a template size of 2 (i.e., L) is used for 4×4 and 8×8 blocks. ), and for blocks of 16×16 and larger, use a template size of 4 (i.e., Template reference (from) Figure 20 The dashed area (indicated by the region) refers to a set of neighboring samples from above and to the left of the template, as defined by JEM-2.0. Unlike template samples, which always come from the reconstructed region, the template's reference samples may not have been reconstructed when encoding / decoding the target block. In this case, JEM-2.0's existing reference sample replacement algorithm is used to replace unavailable reference samples with available ones.

[0197] Figure 20 The target sample, template sample, and template reference sample used in DIMD are shown.

[0198] For each intra-prediction mode, DIMD calculates the absolute difference (SAD) between the reconstructed template samples and the predicted samples obtained from the reference samples of the template. The intra-prediction mode that produces the minimum SAD is selected as the final intra-prediction mode for the target block.

[0199] 2.19.2. DIMD for intra-frame 2N×2N CUs For an intra-frame 2N×2N CU, DIMD is used as an additional intra-frame mode, which is adaptively selected by comparing the DIMD intra-frame mode with the optimal normal intra-frame mode (i.e., explicitly transmitted via signaling). For each intra-frame 2N×2N CU, a flag is transmitted via signaling to indicate the use of DIMD. If the flag is 1, the intra-frame mode derived from DIMD is used to predict the CU; otherwise, DIMD is not applied, and the intra-frame mode explicitly transmitted via signaling in the bitstream is used to predict the CU. When DIMD is enabled, the chroma component always reuses the same intra-frame mode derived for the luma component, i.e., the DM mode.

[0200] Furthermore, for each CU encoded and decoded by DIMD, blocks within the CU can adaptively choose to derive their intra-frame modes at either the PU level or the TU level. Specifically, when the DIMD flag is 1, another CU-level DIMD control flag is signaled to indicate the level at which DIMD is performed. If the flag is 0, it means that DIMD is performed at the PU level, and all TUs in the PU use the same derived intra-frame mode for their intra-frame prediction; otherwise (i.e., the DIMD control flag is 1), it means that DIMD is performed at the TU level, and each TU in the PU derives its own intra-frame mode.

[0201] Furthermore, when DIMD is enabled, the number of angular directions increases to 129, while the DC and planar modes remain the same. To accommodate the increased granularity of the angular intra-frame modes, the precision of the intra-frame interpolation filtering for DIMD-encoded CUs increases from 1 / 32 pixel to 1 / 64 pixel. Additionally, to use the derived intra-frame modes of DIMD-encoded CUs as MPM candidates for neighboring intra-frame blocks, these 129 directions of the DIMD-encoded CUs are converted to "normal" intra-frame modes (i.e., 65 angular intra-frame directions) before being used as MPMs.

[0202] 2.19.3. DIMD for intra-frame N×N CUs In the proposed method, the intra-frame mode of the intra-N×N CU is always transmitted via signaling. However, to improve the efficiency of intra-frame mode encoding and decoding, the intra-frame mode derived from DIMD is used as an MPM candidate to predict the intra-frame mode of the four PUs in the CU. To avoid increasing the overhead of MPM index signaling, the DIMD candidate is always placed at the first position in the MPM list, and the last existing MPM candidate is removed. Furthermore, a deduplication operation is performed so that if a DIMD candidate is redundant, it will not be added to the MPM list.

[0203] 2.19.4. Intra-frame mode search algorithm for DIMD To reduce encoding / decoding complexity, a direct and fast intra-frame mode search algorithm is used for DIMD. First, an initial estimation process is performed to provide a good starting point for the intra-frame mode search. Specifically, this is achieved by selecting from the allowed intra-frame modes... N A fixed pattern is used to create an initial candidate list. Then, the SAD (Shortest Aspect Ratio) is calculated for all candidate intra-modes, and the one that minimizes the SAD is selected as the starting intra-mode. To achieve a good complexity / performance tradeoff, the initial candidate list consists of 11 intra-modes, including DC, planar, and every fourth of the 33 intra-direction angles defined in HEVC, i.e., intra-modes 0, 1, 2, 6, 10…30, 34.

[0204] If the initial intra-mode is DC or planar, it is used as the DIMD mode. Otherwise, based on the initial intra-mode, a refinement process is then applied, where the optimal intra-mode is identified through an iterative search. This works by comparing the SAD values ​​of three intra-modes separated by a given search interval at each iteration and maintaining the intra-mode that minimizes the SAD. The search interval is then halved, and the selected intra-mode from the previous iteration becomes the center intra-mode for the current iteration. For the current DIMD implementation with 129 angular intra-mode directions, a maximum of four iterations are used in the refinement process to find the optimal DIMD intra-mode.

[0205] 2.20. Derivation of the decoder-side intra-frame mode by calculating the gradient of neighboring samples. Three angular patterns are selected from the gradient histograms (HoGs) calculated based on the neighboring pixels of the current block. Once the three patterns are selected, their predictions are calculated normally, and then their weighted average is used as the final prediction for the block. To determine the weights, the corresponding magnitudes in the HoG are used for each of the three patterns. The DIMD pattern is used as an alternative prediction pattern and is always checked in the FullRD pattern.

[0206] The current version of DIMD has modified several aspects of signaling, HoG computation, and prediction fusion. The aim of these modifications is to improve encoding / decoding performance and address the complexity issues raised during the last meeting (i.e., throughput for 4x4 blocks). The following sections describe the modifications for each aspect.

[0207] 2.20.1. Signaling Figure 21 The order of the parsing flags / indexes integrated with the proposed DIMD in VTM5 is shown.

[0208] As can be seen, the DIMD flag of the block is first parsed using a single CABAC context, which is initialized to the default value of 154.

[0209] If flag == 0, the parsing continues normally.

[0210] Otherwise (if flag == 1), only the ISP index is resolved, and the following flags / indexes are presumed to be zero: BDPCM flag, MIP flag, and MRL index. In this case, the entire IPM resolution is also skipped.

[0211] During the resolution phase, when a regular non-DIMD block queries the IPM of its DIMD nearest neighbor, the pattern PLANA_IDX is used as the virtual IPM for the DIMD block.

[0212] Figure 21 The proposed intra-frame block decoding process is shown.

[0213] 2.20.2. Texture Analysis DIMD's texture analysis includes gradient histogram (HoG) calculation (such as... Figure 22 (As shown). HoG calculations are performed by applying horizontal and vertical Sobel filters to pixels in a 3-width stencil surrounding the block. Unless the stencil pixels above fall into a different CTU, they are not used in texture analysis.

[0214] Once calculated, the IPM corresponding to the two highest histogram bars is selected for the block.

[0215] In previous versions, all pixels in the middle row of the template participated in the HoG calculation. However, the current version improves the throughput of this process by applying the Sobel filter more sparsely across the 4x4 block. For this purpose, only one pixel from the left and one pixel from the top are used. This is as follows: Figure 22 As shown.

[0216] In addition to reducing the number of operations used for gradient calculation, this property also simplifies the selection of the two best modes from the HoG, since the resulting HoG cannot have more than two non-zero amplitudes. Figure 22 The HoG calculation is shown from a template with a width of 3 pixels.

[0217] 2.20.3. Predictive Fusion The current version of this method also uses a fusion of three predictions for each block. However, the choice of prediction mode differs, and it utilizes the combined assumption intra-prediction method proposed in [2], where the planar mode is considered to be used in combination with other modes when computing intra-prediction candidates. In the current version, the two IPMs corresponding to the two highest HoG bars are combined with the planar mode.

[0218] Predictive fusion is applied as a weighted average of the three predictions above. For this purpose, the weight of the plane is fixed at 21 / 64 (~1 / 3). Then, the remaining weight of 43 / 64 (~2 / 3) is shared between the two HoG IPMs, proportional to the magnitude of their HoG bars. Figure 23 The process was visualized. Figure 23 The prediction fusion is shown by weighted averaging of two HoG modes and a plane.

[0219] 2.21. DIMD Chroma Mode The DIMD chromaticity model uses the DIMD derivation method based on, for example, Figure 24The reconstructed Y, Cb, and Cr samples from the second nearest neighbor row and column, as shown, derive the chroma intra-prediction mode for the current block. Specifically, the horizontal and vertical gradients are calculated for each co-located reconstructed luminance sample and the reconstructed Cb and Cr samples of the current chroma block to construct the HoG. The intra-prediction mode with the largest histogram amplitude value is then used to perform chroma intra-prediction for the current chroma block. Figure 24 The neighbor reconstructed samples used for the DIMD chromaticity mode are shown.

[0220] When the intra-prediction mode derived from the DIMD chroma mode is the same as the intra-prediction mode derived from the DM mode, the intra-prediction mode with the second largest histogram amplitude value is used as the DIMD chroma mode. A CU level flag is transmitted via signaling to indicate whether the proposed DIMD chroma mode is applied.

[0221] 2.22. Template-based Intra-Frame Mode Derivation (TIMD) This paper presents a template-based intra-frame mode derivation (TIMD) method using MPM, where TIMD modes are derived from MPM using neighboring templates. TIMD modes are used as an additional intra-frame prediction method for CU.

[0222] 2.22.1. Derivation of TIMD Mode For each intra-prediction mode in the MPM, the SATD between the predicted and reconstructed samples of the template is calculated. The intra-prediction mode with the minimum SATD is selected as the TIMD mode and used for intra-prediction of the current CU. The Position-Related Intra-Prediction Combination (PDPC) is included in the derivation of the TIMD mode.

[0223] 2.22.2. TIMD Signaling The proposed method is enabled / disabled via signaling flags in the Sequence Parameter Set (SPS). When the flag is true, a CU-level flag is signaled to indicate whether the proposed TIMD method is used. The TIMD flag is signaled immediately after the MIP flag. If the TIMD flag is true, the remaining syntax elements related to the luma intra-prediction mode (including MRL, ISP, and the normal resolution phase of the luma intra-prediction mode) are skipped.

[0224] 2.22.3. Interaction with the new codec tools A planar DIMD method with predictive fusion is integrated into EE2. When the EE2 DIMD flag is true, the proposed TIMD flag is not transmitted through the signal and is set to false.

[0225] Similar to PDPC, gradient PDPC is also included in the derivation of TIMD modes.

[0226] When the secondary MPM is enabled, both the primary and secondary MPMs are used to derive the TIMD pattern.

[0227] The 6-tap interpolation filter is not used in the derivation of TIMD modes.

[0228] 2.22.4. Modifications to the MPM list construction in the derivation of the TIMD pattern During the construction of the MPM list, the intra-prediction modes of neighboring blocks are derived as planes when they are inter-coded. To improve the accuracy of the MPM list, the propagated intra-prediction modes are derived using motion vectors and reference images when neighboring blocks are inter-coded and are used in the construction of the MPM list. This modification is only applied to the derivation of TIMD modes.

[0229] 2.22.5. TIMD with Fusion Instead of selecting only one mode with the minimum SATD cost, this paper proposes selecting the top two modes with the minimum SATD cost for intra-modes derived using the TIMD method, then fusing them using weights, and using such weighted intra-prediction to encode and decode the current CU.

[0230] The costs of the two selected patterns are compared to a threshold, and a cost factor of 2 is applied during the test, as shown below:

[0231] If the condition is true, fusion is applied; otherwise, only mode1 is used.

[0232] The weights of the patterns are calculated from their SATD costs as follows:

[0233] 2.23. Merge Schema with MVD (MMVD) In addition to the Merge mode (where implicitly derived motion information is directly used for generating prediction samples for the current CU), the Merge mode with motion vector difference (MMVD) is introduced into the VVC. The MMVD flag is transmitted via signaling immediately after the regular Merge flag is sent to indicate whether the MMVD mode is used for the CU.

[0234] In MMVD, after a Merge candidate is selected, its MVD information transmitted via signals is further refined. This further information includes a Merge candidate flag, an index specifying the motion amplitude, and an index indicating the motion direction. In MMVD mode, one of the first two candidates in the Merge list is selected as the MV basis. The MMVD candidate flag is transmitted via signals to specify which candidate to use between the first and second Merge candidates.

[0235] Figure 25 The MVD search point is shown. The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. For example... Figure 25 As shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 7.

[0236] Table 7 – Relationship between Distance Index and Predefined Offset

[0237] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 8. It is important to note that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is a unidirectional or bidirectional prediction MV, and both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the symbol in Table 8 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV, and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), and the POC difference in list 0 is greater than the POC difference in list 1, the symbol in Table 8 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Otherwise, if the POC difference in list 1 is greater than the POC difference in list 0, the symbol in Table 8 specifies the sign of the MV offset added to the list 1 MV component of the starting MV, and the sign of the list 0 MV has the opposite value.

[0238] MVD is scaled based on the POC difference in each direction. If the POC differences in the two lists are the same, no scaling is needed. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, the MVD of list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb, as follows. Figure 26 As shown. If the POC difference of L1 is greater than that of L0, then the MVD of list 0 is scaled in the same way. If the initial MV is unidirectionally predicted, then the MVD is added to the available MV.

[0239] Table 8 – Symbols of MV Offsets Specifyed by Direction Index

[0240] 2.24. Symmetric MVD Encoding and Decoding In VVC, in addition to the normal one-way and two-way predictive MVD signaling, a symmetric MVD mode for two-way predictive MVD signaling is also applied. In the symmetric MVD mode, the reference image indices of both list-0 and list-1, as well as the motion information of list-1's MVD, are not transmitted through signals but are derived.

[0241] The decoding process of the symmetric MVD mode is as follows: 1. At the strip level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to 0.

[0242] Otherwise, if the most recent reference image in list-0 and the most recent reference image in list-1 form a forward and backward reference pair or a backward and forward reference pair, then BiDirPredFlag is set to 1, and both the list-0 and list-1 reference images are short-term reference images. Otherwise, BiDirPredFlag is set to 0.

[0243] 2. At the CU level, if the CU is bidirectional predictive codec and BiDirPredFlag equals 1, the symmetry mode flag, which indicates whether the symmetry mode is used, is explicitly transmitted via signaling.

[0244] When the symmetry mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly transmitted via signaling. The reference indices of list-0 and list-1 are each set to equal a pair of reference images. MVD1 is set to equal to (-MVD0). The final motion vector is shown in the following formula.

[0245]

[0246] Figure 26 The symmetric MVD model is illustrated. In the encoder, symmetric MVD motion estimation begins with an initial MV evaluation. A set of initial MV candidates includes MVs obtained from a one-way prediction search, MVs obtained from a two-way prediction search, and MVs from an AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.

[0247] 2.25. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) was included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version, requiring far fewer computations, especially in terms of the number of multiplications and the size of the multipliers.

[0248] BDOF is used to refine the bidirectional prediction signal of the CU at the 4×4 sub-block level. BDOF is applied to the CU if all of the following conditions are met: - CU is encoded and decoded using a "true" bidirectional prediction mode, that is, one of the two reference images is displayed before the current image in the order of display, and the other is displayed after the current image in the order of display.

[0249] - The distances (i.e., the difference in point of view) from the two reference images to the current image are the same.

[0250] - Both reference images are short-term reference images.

[0251] - CU is not encoded or decoded using affine mode or SbTMVP Merge mode.

[0252] - The CU has more than 64 luminance samples.

[0253] - Both the CU height and CU width are greater than or equal to 8 luminance samples.

[0254] - The BCW weight index indicates equal weights.

[0255] - WP is not enabled for the current CU.

[0256] - CIIP mode is not used for the current CU.

[0257] BDOF is applied only to the luminance component. As the name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of an object is smooth. Motion refinement is performed for each 4×4 sub-block. The difference between the L0 and L1 predicted samples is calculated by minimizing the difference. Motion refinement is then used to adjust the bidirectional predicted sample values ​​in the 4x4 sub-blocks. The following steps are applied during the BDOF process.

[0258] First, the horizontal and vertical gradients of the two predicted signals. It is calculated by directly calculating the difference between two neighboring sample points, i.e.

[0259] in It is a list , Coordinates of the predicted signal in The sample value at the location, and shift1 is calculated based on the luminance bit depth bitDepth as shift1 = max(6, bitDepth-6).

[0260] Then, the autocorrelation and cross-correlation of the gradients. Calculated as

[0261] in

[0262] in It is a 6x6 window surrounding a 4x4 sub-block, and and The values ​​are set to min(1, bitDepth - 11) and min(4, bitDepth - 8) respectively.

[0263] Then the motion is refined. The following formula is derived using cross-correlation and autocorrelation terms:

[0264] in . It is a floor function, and .

[0265] Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4×4 sub-block:

[0266] Finally, the BDOF samples of CU are calculated by adjusting the bidirectional prediction samples as follows:

[0267] These values ​​were chosen such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.

[0268] To derive the gradient values, the list outside the current CU boundary is used. ( Some predicted samples in ) It needs to be generated. For example... Figure 25As shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region (white area) are generated by directly taking reference samples at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and a normal 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values ​​are used only for gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled from their nearest neighbors (i.e., repeated).

[0269] Figure 27 The extended CU region used in BDOF is shown.

[0270] When the width and / or height of a CU is greater than 16 luminance samples, it will be divided into sub-blocks with a width and / or height equal to 16 luminance samples, and the sub-block boundaries will be considered as CU boundaries in the BDOF process. The maximum cell size for the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. The BDOF process is not applied to the sub-block when the SAD between the initial L0 predicted samples and the L1 predicted samples is less than a threshold. The threshold is set to equal to (8... W (H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and L1 prediction samples calculated during the DVMR process is reused here.

[0271] Bidirectional optical flow (BDOF) is disabled if BCW is enabled for the current block, meaning the BCW weight index indicates unequal weights. Similarly, BDOF is disabled if WP is enabled for the current block, meaning either luma_weight_lx_flag is 1 for either of the two reference images. BDOF is also disabled when the CU is encoded and decoded in symmetric MVD or CIIP mode.

[0272] 2.26. Intra-frame and Inter-frame Joint Prediction (CIIP) 2.27. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 28 As shown, the affine motion field of a block is described by motion information from two control point motion vectors (4 parameters) or three control point motion vectors (6 parameters).

[0273] Figure 28 An affine motion model based on control points is shown. Figure 28 In the figure, both the 4-parameter affine model and the 6-parameter affine model are shown.

[0274] For a 4-parameter affine motion model, the location of the sample point in the block ( x, y The motion vector at point () is derived as:

[0275] For a 6-parameter affine motion model, the location of the sample points in the block ( x, y The motion vector at point () is derived as:

[0276] in It is the motion vector of the top left control point. It is the motion vector of the upper right control point, and It is the motion vector of the lower left control point.

[0277] To simplify motion compensation prediction, block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation, such as... Figure 29 As shown, the values ​​are rounded to 1 / 16 fractional precision. A motion-compensated interpolation filter is then applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks. Figure 29 The affine MVF for each sub-block is shown.

[0278] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.

[0279] 2.27.1. Affine Merge Prediction The AF_MERGE mode can be applied to CUs with a width and height greater than or equal to 8. In this mode, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and the one to be used for the current CU is indicated by a signal transmission index. The following three types of CPVM candidates are used to form the affine merge candidate list: – Affine Merge candidates inferred from the CPMV of neighboring CUs.

[0280] – A constructive affine Merge candidate CPMVP derived using translational MV of neighboring CUs.

[0281] – Zero MV.

[0282] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 30 As shown. For the left-hand prediction, the scan order is A0->A1, and for the top-hand prediction, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between candidates from two inherited sides. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. As shown, if the neighboring lower-left block A is encoded and decoded in affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU containing block A are... It is obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are based on Computed. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated according to... Calculated. Figure 30 The location of the inherited affine motion prediction value is shown. Figure 31 The inheritance of control point motion vectors is shown.

[0283] The constructed affine candidate is achieved by combining the translational motion information of each control point's neighbors. The motion information of the control points is derived from... Figure 31 The spatial and temporal nearest neighbors shown are derived. CPMV k (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. If the TMVP is available, it is used as CPMV4.

[0284] After obtaining the motion signatures (MVs) of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, { CPMV1, CPMV2}, { CPMV1, CPMV3}.

[0285] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different. Figure 32 The locations of candidate positions for constructing the affine Merge pattern are shown.

[0286] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list.

[0287] 2.27.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height both greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated by sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.

[0288] – A constructive affine AMVP candidate CPMVP derived using the translational MV of neighboring CUs.

[0289] – Translation MV from the neighboring CU.

[0290] – Zero MV.

[0291] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AVMP candidates, only affine CUs with the same reference picture as those in the current block are considered. No deduplication process is applied when inserting inherited affine motion predictions into the candidate list.

[0292] The constructed AMVP candidate is from Figure 15 The specified spatial nearest neighbor derivation is shown. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is checked. The first block in the checking order to be inter-coded and has the same reference picture as in the current CU is used. Only one. When the current CU is encoded and decoded using a 4-parameter affine mode, and When all three CPMVs are available, they are added as candidates in the affine AMVP list. If the current CU is encoded / decoded using a 6-parameter affine mode and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.

[0293] If, after checking the inherited affine AMVP candidates and the constructed AMVP candidates, the affine AMVP list still has fewer than 2 candidates, then They will be added sequentially as translation MVs to predict all control point MVs of the current CU (when available). Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.

[0294] 2.27.3. Affine Motion Information Storage In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is used only to generate inherited CPMVPs in the affine Merge mode and affine AMVP mode for the most recently encoded / decoded CU. Subblock MVs derived from the CPMV are used for motion compensation, MV derivation of the Merge / AMVP list for translation MVs, and deblocking.

[0295] To avoid image row caching for additional CPMVs, affine motion data inheritance from CUs of the upper CTU is treated differently from inheritance from regular neighboring CUs. If a candidate CU for affine motion data inheritance is in the upper CTU row, the lower left and lower right sub-block MVs in the row cache, instead of the CPMV, are used for affine MVP derivation. Thus, the CPMV is only stored in the local cache. If the candidate CU is a 6-parameter affine codec, the affine model is downgraded to a 4-parameter model. Figure 16 As shown, along the top CTU boundary, the motion vectors of the lower left and lower right sub-blocks of the CU are used for the affine inheritance of the CU in the bottom CTU. Figure 33 A diagram illustrating the motion vectors used for the proposed combination method is shown.

[0296] 2.27.4. Refinement of Optical Flow Prediction for Affine Modes Compared to pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of prediction accuracy. To achieve finer-grained motion compensation, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth used for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the brightness prediction samples are refined by adding differences derived from the optical flow equation. PROF is described in the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate sub-block predictions. .

[0297] Step 2) Calculate the spatial gradient of the sub-block prediction at each sample location using a 3-tap filter [-1, 0, 1]. The gradient calculation is exactly the same as the gradient calculation in BDOF.

[0298]

[0299] This is used to control the precision of the gradient. The sub-block (i.e., 4x4) prediction is expanded by one sample point on each side for gradient computation. To avoid additional memory bandwidth and additional interpolation computation, those expanded samples on the expanded boundaries are copied from the nearest integer pixel position in the reference image.

[0300] Step 3) Brightness prediction refinement is calculated using the following optical flow equation.

[0301]

[0302] in It refers to the location of the sample points. The calculated sample MV (by (representation) and sample points The differences between the MVs of the sub-blocks of the same sub-block, such as Figure 34 As shown. It is quantized in units of 1 / 32 brightness sample precision. Figure 34 The sub-block MV V is shown SB and pixels (Indicated by the gray arrow).

[0303] Since the affine model parameters and the sample point positions relative to the sub-block center do not change from one sub-block to another, therefore It can be computed for the first sub-block and reused for other sub-blocks in the same CU. Let... and From the sample point location To the center of the sub-block Horizontal and vertical offsets It can be derived through the following equation.

[0304]

[0305] To maintain accuracy, the center of the sub-block Calculated as ( ( W SB - 1) / 2, (H) SB - 1) / 2), where W SBand H SB These are the width and height of the sub-block, respectively.

[0306] For a 4-parameter affine model,

[0307] For a 6-parameter affine model,

[0308] in These are the motion vectors of the top left, top right, and bottom left control points. and These are the width and height of the CU.

[0309] Step 4) Finally, refine the brightness prediction. Added to sub-block prediction Final prediction I’ It is generated as the following equation.

[0310]

[0311] PROF is not applied in two cases for affine codecs: 1) all control points MV are the same, which indicates that the CU only has translational motion; 2) the affine motion parameters are greater than the specified limits, because the sub-block-based affine MC is downgraded to the CU-based MC to avoid large memory access bandwidth requirements.

[0312] Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) if the CU is not the root block and its parent block does not choose an affine mode as its optimal mode, then PROF is not applied because the probability of the current CU choosing an affine mode as its optimal mode is low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold, and the current image is not a low-latency image, then PROF is not applied because the improvement introduced by PROF is small in this case. Thus, affine motion estimation using PROF can be accelerated.

[0313] 2.28. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-image to improve motion vector prediction and merge patterns for the current CU in the current image. The same co-image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: – TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level; – TMVP obtains temporal motion vectors from co-op blocks in the co-op image (the co-op block is the lower right or center block relative to the current CU), and SbTMVP applies motion displacement before obtaining temporal motion information from the co-op image, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks from the current CU.

[0314] The SbTVMP process is as follows: Figure 18A As shown. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, it checks... Figure 18A The spatial nearest neighbor A1 is selected. If A1 has a motion vector that uses a co-located image as its reference image, then that motion vector is chosen as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0).

[0315] In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to the position of the current block. Figure 18B The corresponding image shown obtains motion information (motion vectors and reference indices) at the sub-CU level. Figure 18B The example assumes the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) in the co-location image is used to derive the motion information for the sub-CU. After the motion information of the co-location sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU. Figure 35A and Figure 35B The SbTMVP procedure in VVC is shown. Figure 35A The spatial neighbor block used by ATVMP is shown. Figure 35B The paper demonstrates how to derive the motion field of a sub-CU by applying motion displacements from spatial neighbors and scaling motion information from the corresponding co-located sub-CU.

[0316] In VVC, a sub-block-based Merge list containing a combination of SbTVMP candidates and affine Merge candidates is used for signaling in sub-block-based Merge mode. SbTVMP mode is enabled / disabled via the Sequence Parameter Set (SPS) flag. If SbTVMP mode is enabled, the SbTVMP prediction is added as the first entry in the sub-block-based Merge candidate list, followed by the affine Merge candidate. The size of the sub-block-based Merge list is transmitted via signaling in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5.

[0317] The sub-CU size used in SbTMVP is fixed at 8x8, and like the affine Merge pattern, the SbTMVP pattern only applies to CUs with a width and height greater than or equal to 8.

[0318] The encoding logic for the additional SbTMVP Merge candidate is the same as that for other Merge candidates, that is, for each CU in the P-strip or B-strip, an additional RD check is performed to determine whether to use the SbTMVP candidate.

[0319] 2.29. Adaptive Motion Vector Resolution (AMVR) In HEVC, when `use_integer_mv_flag` in the strip header is equal to 0, the motion vector difference (MVD) (between the CU's motion vector and the predicted motion vector) is transmitted through the signal in units of quarter-luminance samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the CU's MVD to be encoded and decoded with different precisions. Depending on the current CU's mode (normal AMVP mode or affine AVMP mode), the current CU's MVD can be adaptively selected as follows: – Normal AMVP mode: quarter brightness sample, half brightness sample, integer brightness sample or four brightness sample.

[0320] – Affine AMVP mode: quarter luminance sample, integer luminance sample, or 1 / 16 luminance sample.

[0321] If the current CU has at least one non-zero MVD component, the MVD resolution indication at the CU level is conditionally transmitted via signal transmission. If all MVD components (i.e., both the horizontal and vertical MVD for reference lists L0 and L1) are zero, the quarter-spot luminance sample MVD resolution is presumed.

[0322] For a CU with at least one non-zero MVD component, a first flag is transmitted to indicate whether quarter-luminance sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required, and quarter-luminance sample MVD precision is used for the current CU. Otherwise, a second flag is transmitted to indicate that half-luminance samples or other MVD precision (integer or quad-luminance samples) is used for the normal AMVP CU. In the case of half-luminance samples, a 6-tap interpolation filter is used instead of the default 8-tap interpolation filter at the half-luminance sample location. Otherwise, a third flag is transmitted to indicate whether integer or quad-luminance sample MVD precision is used for the normal AMVP CU. In the case of an affine AMVP CU, the second flag is used to indicate whether integer or 1 / 16 luminance sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luminance sample, half-luminance sample, integer, or quad-luminance sample), the CU's motion vector predictions are rounded to the same precision as the MVD before being added to it. The predicted motion vector values ​​are rounded to zero (i.e., negative motion vector prediction values ​​are rounded to positive infinity, and positive motion vector prediction values ​​are rounded to negative infinity).

[0323] The encoder uses RD checks to determine the motion vector resolution for the current CU. To avoid always performing CU-level RD checks four times for every MVD resolution, in VTM11, RD checks for MVD accuracy other than quarter-luminance samples are only conditionally invoked. For normal AVMP mode, the RD cost for quarter-luminance sample MVD accuracy and the RD cost for integer luminance sample MVD accuracy are first calculated. Then, the RD cost for integer luminance sample MVD accuracy is compared with the RD cost for quarter-luminance sample MVD accuracy to determine if it is necessary to further check the RD cost for four-luminance sample MVD accuracy. When the RD cost for quarter-luminance sample MVD accuracy is much smaller than the RD cost for integer luminance sample MVD accuracy, the RD check for four-luminance sample MVD accuracy is skipped. Then, if the RD cost for integer luminance sample MVD accuracy is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for half-luminance sample MVD accuracy is skipped. For affine AMVP mode, if the affine inter-frame mode is not selected after checking the rate-distortion cost of affine Merge / Skip mode, Merge / Skip mode, quarter-lumen sample MVD precision normal AMVP mode, and quarter-lumen sample MVD precision affine AMVP mode, then the 1 / 16-lumen sample MVD precision and 1-pixel MVD precision affine inter-frame modes are not checked. Furthermore, the affine parameters obtained in the quarter-lumen sample MVD precision affine inter-frame mode are used as the starting search point in the 1 / 16-lumen sample and quarter-lumen sample MVD precision affine inter-frame modes.

[0324] 2.30. Bidirectional prediction with CU-level weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference images and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals.

[0325]

[0326] Five weights are allowed in weighted average two-way forecasting. For each bidirectional prediction CU, the weights w are determined in one of two ways: 1) for non-merge CUs, the weight index is transmitted via signal after the motion vector difference; 2) for merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW is applied only to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency images, all 5 weights are used. For non-low-latency images, only 3 weights (w∈{3,4,5}) are used.

[0327] At the encoder, fast search algorithms are applied to find weight indices without significantly increasing encoder complexity. These algorithms are summarized below. When combined with AMVR, if the current image is a low-latency image, weights with unequal precision for 1-pixel motion vectors and 4-pixel motion vectors are only conditionally checked.

[0328] - When combined with an affine pattern, the affine ME will be executed for unequal weights if and only if the affine pattern is selected as the current best pattern.

[0329] - When the two reference images in bidirectional prediction are the same, unequal weights are only checked conditionally.

[0330] - Under certain conditions, unequal weights are not searched, depending on the POC distance between the current image and its reference image, the QP of the encoding / decoding algorithm, and the temporal level.

[0331] The BCW weight index is encoded using one context codec bit and a subsequent bypass codec bit. The first context codec bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is transmitted via bypass codec to indicate which unequal weight is used.

[0332] Weighted Prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with fade-in and fade-out effects. Support for WP has also been added to the VVC standard. WP allows weighted parameters (weights and offsets) to be transmitted via signaling for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the corresponding weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if the CU uses WP, the BCW weight index is not transmitted via signaling, and w is presumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge patterns and inherited affine Merge patterns. For constructed affine Merge patterns, affine motion information is constructed based on motion information from up to 3 blocks. The BCW index of the CU using the constructed affine Merge pattern is simply set to be equal to the BCW index of the first control point MV.

[0333] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is encoded and decoded in CIIP mode, the BCW index of the current CU is set to 2, for example, with equal weights.

[0334] 2.31. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is an encoding / decoding tool used to address the problem of local illumination variations between the current image and its temporal reference image. LIC is based on a linear model, where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples for the current block. Specifically, LIC can be mathematically modeled using the following equation:

[0335] in In coordinates The prediction signal for the current block at that location; It is composed of motion vectors The reference block it points to; and These are the corresponding scaling factors and offsets applied to the reference block. Figure 36 The LIC procedure is shown. Figure 36 In this context, when applying LIC to a block, the Minimum Mean Square Error (LMSE) method is employed to minimize the number of neighboring samples in the current block (i.e., ...). Figure 36 templates in T ) and its corresponding reference sample in the time-domain reference image (i.e. Figure 36 In T0 or T1 The difference between ) is used to derive the LIC parameters (i.e. and The value of ). Furthermore, to reduce computational complexity, both the template sample and the reference template sample are downsampled (adaptive downsampling) to derive the LIC parameters; that is, only Figure 36 The shaded samples in the data were used for derivation. and .

[0336] To improve encoding and decoding performance, such as Figure 37 As shown, downsampling is not performed on the shorter side. Figure 36 Local lighting compensation is shown. Figure 37 This shows that no downsampling was performed on the short side.

[0337] 2.32. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the motion vector refinement (MV) in the Merge mode, a decoder-side motion vector refinement based on bilateral matching (BM) is applied in the VVC. In the bidirectional prediction operation, a refined MV is searched around the initial MV in reference image lists L0 and L1. The BM method computes the distortion between two candidate blocks in reference image lists L0 and L1. Figure 21 As shown, the SAD between the two blocks is calculated based on each MV candidate (e.g., MV0' and MV1') around the initial MV. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal.

[0338] Figure 38 The motion vector refinement on the decoding side is shown. In VVC, the application of DMVR is limited and is only applied to CUs encoded and decoded in the following modes and features: - CU-level Merge pattern with bidirectional prediction MV.

[0339] - Relative to the current image, one reference image is from the past and another reference image is from the future.

[0340] - The distances (i.e., the difference in point of view) from the two reference images to the current image are the same.

[0341] - Both reference images are short-term reference images.

[0342] - The CU has more than 64 luminance samples.

[0343] - Both the CU height and CU width are greater than or equal to 8 luminance samples.

[0344] - The BCW weight index indicates equal weights.

[0345] - Disable WP for the current block.

[0346] - Do not use CIIP mode for the current block.

[0347] The refined motion vector (MV) derived through the DMVR process is used to generate inter-frame prediction samples and also for temporal motion vector prediction for future image encoding and decoding. The original MV is used in the deblocking process and also for spatial motion vector prediction for future CU encoding and decoding.

[0348] Additional features of DMVR are mentioned in the following sub-entries.

[0349] 2.32.1. Search Scheme In DVMR, the search point revolves around the initial MV, and the MV offset follows the MV difference mirror rule. In other words, any point examined by DMVR, represented by the candidate MV pair (MV0, MV1), obeys the following two equations:

[0350] in This represents the refinement between the initial MV and the refined MV in one of the reference images. The refinement search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage.

[0351] A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of the uncertainty of DMVR refinement, a bias towards the original MV is proposed during the DMVR process. The SAD value between the reference blocks of the initial MV candidate reference is reduced by 1 / 4.

[0352] The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation instead of an additional search with SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration with the minimum SAD at the center.

[0353] In subpixel offset estimation based on parametric error surfaces, the cost at the center location and the costs at the four nearest neighbor locations are used to fit a two-dimensional parabolic error surface equation of the following form:

[0354] in( This corresponds to the score position with the minimum cost, and C corresponds to the minimum cost. The above equation is solved by using the costs of the five search points. Calculated as:

[0355] and The value is automatically constrained between -8 and 8 because all values ​​are positive, and the minimum value is... This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated score ( Added to integer distance refinement MV to obtain subpixel accurate refinement increment MV.

[0356] 2.32.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 of a luminance sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search point surrounds the initial fractional pixel MV with an integer sample offset, so samples at those fractional positions need to be interpolated during the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect is that by using a bilinear filter and utilizing a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not needed by the interpolation process based on the original MV but are needed by the interpolation process based on the refined MV are filled from these available samples.

[0357] 2.32.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luminance samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luminance samples. The maximum cell size for the DMVR search process is limited to 16x16.

[0358] 2.33. Multi-pass decoder-side motion vector refinement In this paper, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector predictions.

[0359] 2.33.1. First pass – Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying the BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.

[0360] BM performs a local search to derive the integer sample precision intDeltaMV and the half-pixel sample precision halfDeltaMv. The local search applies a 3×3 square search pattern to iterate through the search range [-sHor, sHor] in the horizontal direction and the search range [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0361] The cost of bilateral matching is calculated as: bilCost = mvDistanceCost + sadCost. When the block size is cbW... When cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of distortion between reference blocks. The local search of intDeltaMV or halfDeltaMV is terminated when bilCost at the center point of the 3×3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.

[0362] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: · MV0_pass1 = MV0 + deltaMV; · MV1_pass1 = MV1 – deltaMV.

[0363] 2.33.2. Second pass – Sub-block-based bilateral matching MV refinement In the second pass, the refined MV is derived by applying the BM to 16×16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass for the reference image lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.

[0364] For each sub-block, BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and a search range of [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0365] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference subblocks, as follows: bilCost = satdCost costFactor. Search area (2) sHor + 1) (2) sVer + 1) is divided into Figure 22 A maximum of 5 diamond-shaped search regions are shown. Each search region is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order, starting from the top left corner and ending at the bottom right corner. When the minimum bilCost within the current search region is less than a threshold (which is equal to sbW), a loss occurs. When sbH), the integer pixel full search is terminated; otherwise, the integer pixel full search continues to the next search area until all search points have been checked.

[0366] Figure 39 The rhomboid region in the search area is shown. BM performs a local search to derive the half-sample precision halfDeltaMv. The search pattern and cost function are the same as defined in 2.9.1.

[0367] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). The refined MV at the second pass is then derived as: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2); · MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2).

[0368] 2.33.3. Third pass – Sub-block based bidirectional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy, without clipping from the refined MV of the parent-child blocks in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.

[0369] The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) at the third pass are derived as follows: · MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv; · MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv.

[0370] 2.34. Sample-based BDOF In sample-based BDOF, instead of block-based derivation motion refinement (Vx, Vy), it is performed for each sample.

[0371] The encoding / decoding block is divided into 8×8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, a sliding 5×5 window is used for each sample in the sub-block, and the existing BDOF procedure is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional predicted sample values ​​for the center sample of the window.

[0372] 2.35. Extended Merge Forecast In VVC, the Merge candidate list is constructed by including the following five types of candidates in sequence: (1) Airspace MVP from the airspace adjacent to the CU.

[0373] (2) Temporal MVP from the same CU.

[0374] (3) History-based MVP from FIFO table.

[0375] (4) Pair average MVP.

[0376] (5) Zero MV.

[0377] The size of the Merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU encoding / decoding in the Merge mode, the index of the best Merge candidate is encoded using rounded unary binarization (TU). The first bit of the Merge index is encoded / decoded using the context, and bypass encoding / decoding is used for the other bits.

[0378] This section provides the derivation process for each category of merge candidates. Similar to HEVC, VVC also supports parallel derivation of the merge candidate list for all CUs within a given size region.

[0379] 2.35.1. Derivation of Airspace Candidates The derivation of spatial merge candidates in VVC is the same as in HEVC, except that the positions of the first two merge candidates are swapped. At most four merge candidates are selected from those located at the positions shown. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or when intra-frame encoding / decoding is required. After adding the candidate at position A1, a redundancy check is performed on the remaining candidates. This check ensures that candidates with the same motion information are excluded from the list, thus improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the redundancy check. Instead, only... Figure 24 The system uses arrow links to select pairs, and only adds candidates to the list if the corresponding candidates used for redundancy checks do not have the same motion information.

[0380] Figure 40 The locations of the spatial merge candidates are shown.

[0381] Figure 41 The candidate pairs considered for redundancy checks of spatial Merge candidates are shown.

[0382] 2.35.2. Derivation of Time-Domain Candidates In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for the derivation of the co-located CU is explicitly transmitted via signal transmission in the strip header. Figure 25 As shown by the dashed lines, the scaled motion vectors for the temporal merge candidates are obtained by scaling the motion vectors from the co-located CUs using the POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located reference image and the co-located image. The reference image index for the temporal merge candidates is set to 0.

[0383] Figure 42 An illustration of motion vector scaling for temporal Merge candidates is shown.

[0384] like Figure 26 As shown, the position for the temporal candidate is selected between candidate C0 and C1. If the CU at position C0 is unavailable, intra-frame encoded / decoded, or outside the current line of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate. Figure 43 Candidate positions C0 and C1 are shown for the time-domain Merge candidate.

[0385] 2.35.3. Historical Merge Candidate Derivation Historically based MVP (HMVP) merge candidates are added to the merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a CU that is not inter-frame encoded / decoded in a sub-block is encountered, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0386] The HMVP table size S is set to 6, indicating that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When a new candidate is inserted into the table, a constrained First-In-First-Out (FIFO) rule is used, where a redundancy check is first applied to find if a duplicate HMVP already exists in the table. If found, the duplicate HMVP is removed from the table, and all subsequent HMVP candidates are moved forward.

[0387] HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked sequentially and inserted into the candidate list after the TMVP candidates. Redundancy checks are applied between HMVP candidates and spatial or temporal Merge candidates.

[0388] To reduce the number of redundant check operations, the following simplifications are introduced: The number of HMPV candidates used for the Merge list generation is set to (N<= 4) ? M: (8 - N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table.

[0389] Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the process of building the Merge candidate list from HMVP is terminated.

[0390] 2.35.4. Derivation of Pairwise Average Merge Candidates Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing Merge candidate list. These predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid.

[0391] If the Merge list is not full after adding pairwise average Merge candidates, a zero MVP will be inserted at the end until the maximum number of Merge candidates is reached.

[0392] 2.35.5. Merge estimation area The Merge Estimation Region (MER) allows for the independent derivation of Merge candidate lists for CUs within the same Merge Estimation Region (MER). Candidate blocks located within the same MER as the current CU are not included in the generation of the current CU's Merge candidate list. Furthermore, the update process for the historical motion vector prediction candidate list is only updated if (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left brightness sample position of the current CU in the image, and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and is transmitted via signaling as log2_parallel_merge_level_minus2 in the sequence parameter set.

[0393] 2.36. New Merge Candidates 2.36.1. Derivation of non-adjacent Merge candidates In VVC, such as Figure 27 The five spatial neighbor blocks and one temporal nearest neighbor shown were used to derive the Merge candidate.

[0394] We propose using the same pattern as in VVC to derive additional Merge candidates from positions not adjacent to the current block. To achieve this, for each search round i, a virtual block is generated based on the current block, as follows: First, for the current block, the relative position of the virtual block is calculated using the following formula:

[0395] Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.

[0396] Secondly, the width and height of the virtual block are calculated using the following formula:

[0397] Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.

[0398] gridX and gridY are currently set to currWidth and currHeight, respectively.

[0399] Figure 44 The VVC space neighboring blocks of the current block are shown. Figure 28 This illustrates the relationship between the virtual block and the current block. Figure 45 A diagram of the virtual block in the i-th round of search is shown.

[0400] After the virtual block is generated, block A i B i C i D i and E i These can be considered VVC spatial neighbor blocks, and their positions are obtained using the same pattern as the pattern in the VVC. Clearly, if the search round i is 0, the virtual block is the current block. In this case, block A... i B i C i D i and E i It is a spatial neighbor block used in VVC Merge mode.

[0401] When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbor blocks are utilized.

[0402] Non-adjacent spatial merge candidates are inserted into the merge list after the temporal merge candidates in the order B1->A1->C1->D1->E1.

[0403] 2.36.2. STMVP We propose using three spatial merge candidates and one temporal merge candidate to derive the average candidate as the STMVP candidate.

[0404] STMVP is inserted before the Merge candidate in the upper left airspace.

[0405] STMVP candidates were deduplicated along with all previous Merge candidates in the Merge list.

[0406] For airspace candidates, the top three candidates in the current Merge candidate list are used.

[0407] For time-domain candidates, the same position as the co-position of VTM / HEVC is used.

[0408] For airspace candidates, the first, second, and third candidates inserted into the current Merge candidate list before STMVP are denoted as F, S, and T.

[0409] A time-domain candidate with the same position as the VTM / HEVC co-position used in TMVP is denoted as Col.

[0410] The motion vector (denoted as mvLX) of the STMVP candidate in the prediction direction X is derived as follows: 1) If all four Merge candidates have valid reference indices and are all equal to 0 in the prediction direction X (X = 0 or 1),

[0411] 2) If the reference indices of three of the four Merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1),

[0412] 3) If the reference indices of two of the four Merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1),

[0413] Note: STMVP mode is turned off if time-domain candidates are not available.

[0414] 2.36.3. Merge list size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the Merge list is 8.

[0415] 2.37. Geometric Partitioning (GPM) In VVC, geometric segmentation modes are supported for inter-frame prediction. Geometric segmentation modes are transmitted via signaling using a CU-level flag as a merge mode. Other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. For each possible CU size... ,in Excluding 8x64 and 64x8, the geometric segmentation mode supports a total of 64 segmentation types.

[0416] When using this mode, the CU is divided into two parts by a straight line of geometric positioning ( Figure 29 The position of the dividing line is mathematically derived from the angle and offset parameters of a specific segment. Each part of the geometric segment in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each segment, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with conventional bidirectional prediction, only two motion-compensated predictions are needed per CU. The unidirectional prediction motion for each segment is derived using the process described in 2.34.1. Figure 46 An example of GPM partitioning grouped at the same angle is shown.

[0417] If a geometric segmentation mode is used for the current CU, the geometric segmentation index (angle and offset) and two merge indices (one for each segment) are further indicated via signal transmission. The number of maximum GPM candidate sizes is explicitly transmitted in the SPS, and the syntax binarization used for the GPM merge index is specified. After predicting each part of the geometric segment, a blending process with adaptive weights is used to adjust the sample values ​​along the geometric segment edges, as in 2.34.2. This is the prediction signal for the entire CU, and the transformation and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric segmentation mode is stored, as in 2.34.3.

[0418] 2.37.1. Construction of One-Way Prediction Candidate List The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 2.32. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for the geometric segmentation pattern. These motion vectors in... Figure 30 The value is marked with "x". If the corresponding LX motion vector for the nth extended Merge candidate does not exist, the L(1 - X) motion vector of the same candidate is used as the unidirectional predicted motion vector for the geometric segmentation pattern. Figure 47The unidirectional prediction MV selection for geometric segmentation mode is shown.

[0419] 2.37.2. Blending along geometrically segmented edges After predicting each part of the geometric segment using its own motion, a blend is applied to the two predicted signals to derive samples around the segmentation edges. The blending weights for each location of the CU are derived based on the distance between the individual location and the segmentation edge.

[0420] Location The distance to the segmentation edge is derived as follows:

[0421]

[0422] in It is an index used for the angle and offset of geometric segmentation, which depends on the geometric segmentation index transmitted via signal. The sign depends on the angle index. .

[0423] The weights of each part of the geometric segment are derived as follows:

[0424] partIdx depends on the angle index Weight An example in Figure 31 It is shown in the middle. Figure 48 The bending weights using the geometric segmentation pattern are shown. An example of generation.

[0425] 2.37.3. Motion field storage for geometric segmentation patterns Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded by the geometric segmentation pattern.

[0426] The type of motion vector stored for each individual location in the sports field is determined as follows:

[0427] Where motionIdx equals It is recalculated from equation (2-18). partIdx depends on the angle index. .

[0428] If sType equals 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field; otherwise, if sType equals 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.

[0429] Otherwise, if Mv1 and Mv2 come from the same list, only the unidirectional predicted motion Mv2 is stored.

[0430] 2.38. Multiple Hypothesis Prediction In multiple hypothesis prediction (MHP), in addition to inter-frame AMVP mode, regular Merge mode, affine Merge mode, and MMVD mode, up to two additional predictions are transmitted via signal. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.

[0431]

[0432] Weighting factors α As specified in Table 9 below.

[0433] Table 9 – Weighting factors for MHP

[0434] For inter-frame AMVP mode, MHP is only applied when unequal weights in BCW are selected in bidirectional prediction mode.

[0435] Additional assumptions can be either Merge or AMVP mode. In Merge mode, motion information is indicated by the Merge index, and the Merge candidate list is the same as in the geometric segmentation mode. In AMVP mode, the reference index, MVP index, and MVD are transmitted via signals.

[0436] 2.39. Non-adjacent airspace candidates Non-adjacent airspace merge candidates are inserted after the TMVP in the regular merge candidate list. The style of airspace merge candidates is as follows: Figure 32 The distance between non-adjacent spatial domain candidates and the current codec block is shown in the diagram. The distance between the current codec block and the non-adjacent spatial domain candidate is based on the width and height of the current codec block.

[0437] Figure 49 The spatial neighbor blocks used to derive spatial merge candidates are shown.

[0438] 2.40. Template Matching (TM) Template matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference image (i.e., of the same size as the template). For example... Figure 33 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU. This paper employs template matching and makes two modifications: determining the search step size based on AMVR mode, and allowing TM to be cascaded with the bilateral matching process in Merge mode.

[0439] Figure 50 The template matching is shown to be performed over the search area around the initial MV. In AMVP mode, the MVP candidate is determined based on the template matching error, picking the one that achieves the minimum difference between the current block template and the reference block template, and then TM performs MV refinement only on that specific MVP candidate. TM refines the MVP candidate by using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined by using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixel and quarter-pixel MVD precision sequentially according to the AMVR mode specified in Table 10. This search process ensures that the MVP candidate retains the same MV precision as indicated by the AMVR mode after the TM process.

[0440] Table 10 – AMVR Search Styles and Merge Mode Using AMVR

[0441] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 10, TM can proceed up to 1 / 8 pixel MVD precision, or skip those precisions beyond half-pixel MVD precision, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check.

[0442] 2.41. Overlapping Block Motion Compensation (OBMC) Overlapping Block Motion Compensation (OBMC) was previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, it is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. Furthermore, it is applied to both the luma and chroma components. In JEM, MC blocks correspond to codec blocks. When a CU is encoded and decoded in sub-CU modes (including sub-CU merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To handle CU boundaries uniformly, OBMC is performed at the sub-block level on all MC block boundaries, where the sub-block size is set to equal to 4×4, such as... Figure 34 As shown.

[0443] When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of the four connected neighboring sub-blocks (if available and different from the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block.

[0444] The predicted block based on the motion vectors of neighboring sub-blocks is represented as P N ,in N Instructions for neighboring Above , Down square , Left side and right side The index of the sub-block, and the predicted block based on the motion vector of the current sub-block is represented as: P C .when P N When based on the motion information of neighboring sub-blocks that contain the same motion information as the current sub-block, OBMC does not... P N Executed. Otherwise, P N Each sample point was added P C The same points in, that is, P N Four rows / columns were added P C Weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} were used. P N And the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} were used P CThe exception is small MC blocks (i.e., when the height or width of the codec block is equal to 4 or the CU is encoded / decoded in sub-CU mode). For small MC blocks, only P N Two rows / columns were added P C In this case, the weighting factors {1 / 4, 1 / 8} are used. P N And the weighting factors {3 / 4, 7 / 8} were used P C For motion vectors generated based on vertical (horizontal) neighboring sub-blocks P N , P N Samples in the same row (column) are added with the same weight factor. P C .

[0445] Figure 51 The diagram illustrates the application of OBMC to a sub-block. In JEM, for CUs with a size of 256 lumen samples or less, a CU level flag is signaled to indicate whether OBMC is applied to the current CU. For CUs with a size greater than 256 lumen samples or not encoded / decoded in AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to the CU, its effects are considered during the motion estimation phase. The predicted signal formed by OBMC using motion information from the top and left neighboring blocks is used to compensate for the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.

[0446] 2.42. Multiple Transform Selection (MTS) for Kernel Transformation In addition to DCT-II, which is used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding and decoding of both intra-frame and inter-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. Newly introduced transform matrices are DST-VII and DCT-VIII. Table 11 shows the basis functions of the selected DST / DCT.

[0447] Table 11 – Transform basis functions for DCT-II / VIII and DSTVII for N-point inputs

[0448] To maintain the orthogonality of the transformation matrices, the transformation matrices are quantized more precisely than those in HEVC. To keep the intermediate values ​​of the transformation coefficients within the 16-bit range, all coefficients have 10 bits after both the horizontal and vertical transformations.

[0449] To control the MTS scheme, separate enable flags are specified at the SPS level for intra-frame and inter-frame operations. When MTS is enabled at SPS, CU-level flags are signaled to indicate whether MTS is applied. Here, MTS is applied only to luminance. MTS signaling is skipped when one of the following conditions is met.

[0450] – The position of the last significant coefficient for luminance TB is less than 1 (i.e., DC only).

[0451] – The last significant coefficient of luminance TB is located within the MTS zeroing region.

[0452] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are transmitted via signaling to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 12. A unified transform selection for ISP and implicit MTS is used by removing intra-frame mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra-frame block and both intra-frame explicit MTS and inter-frame explicit MTS are enabled, only DST7 is used for both the horizontal and vertical transform kernels. When transform matrix precision is involved, an 8-bit master transform kernel is used. Therefore, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, and 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, and 8-point, 16-point, and 32-point DST-7 and DCT-8) use an 8-bit master transform kernel.

[0453] Table 12 – Transformation and signaling mapping table

[0454] To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are zeroed out for DST-7 and DCT-8 blocks with a size (width or height, or both) equal to 32. Only the coefficients in the 16x16 low-frequency region are retained.

[0455] In HEVC, for example, block residuals can be encoded and decoded in transform skip mode. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not transmitted via signaling when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Implicit MTS can still be enabled when MTS is enabled for inter-frame codec blocks.

[0456] 2.43. Subblock Transformation (SBT) In VTM, a sub-block transform is introduced for inter-frame prediction CUs. In this transform mode, only a sub-part of the residual block is encoded and decoded for the CU. When the inter-frame predicted CU has cu_cbf equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is encoded or decoded. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is encoded and decoded with a presumed adaptive transform, and the remaining portion of the residual block is zeroed out.

[0457] When SBTs are used in inter-frame encoding / decoding CUs, SBT type and SBT position information are transmitted via signals in the bitstream. For example... Figure 35A and Figure 35B As shown, there are two SBT types and two SBT locations. For SBT-V (or SBT-H), the TU width (or height) can be equal to half the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. A 2:2 partition is like a binary tree (BT) partition, while a 1:3 / 3:1 partition is like an asymmetric binary tree (ABT) partition. In an ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 (in luminance samples), a 1:3 / 3:1 partition along that dimension is prohibited. A CU can have a maximum of 8 SBT modes.

[0458] Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (DCT-2 is always used for chroma TB). Two positions in SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are... Figure 35A and 35B The subblock transformation is specified. For example, the horizontal and vertical transformations for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transformations in both dimensions are set to DCT-2. Therefore, the subblock transformation jointly specifies the TU slicing, cbf, and horizontal and vertical kernel transformation types of the residual block.

[0459] SBT is not applied to CUs encoded and decoded in an intra-frame / inter-frame joint mode. Figure 52 The location, type, and transformation type of the SBT are shown.

[0460] 2.44. Adaptive Merge Candidate Reordering Based on Template Matching To improve encoding and decoding efficiency, after constructing the merge candidate list, the order of each merge candidate is adjusted based on the template matching cost. The merge candidates are arranged in the list according to their ascending template matching cost. This is done in subgroups.

[0461] Template matching cost is measured by the sum of absolute differences (SAD) between the current CU's neighboring samples and their corresponding reference samples. If the merge candidate includes bidirectional predicted motion information, then... Figure 36 As shown, the corresponding reference sample is the average of the corresponding reference sample in reference list 0 and the corresponding reference sample in reference list 1. If the Merge candidate contains motion information at the sub-CU level, then as follows... Figure 37 As shown, the corresponding reference sample point is composed of the neighboring sample points of the corresponding reference sub-block.

[0462] like Figure 38 As shown, the sorting process is performed in subgroups. The first three merge candidates are sorted together. The last three merge candidates are sorted together. The template size (width of the left template or height of the top template) is 1. The subgroup size is 3. Figure 53 The neighboring samples used to calculate SAD are shown. Figure 54 The neighboring samples used to calculate SAD for sub-CU level motion information are shown. Figure 55 The sorting process is shown.

[0463] 2.45. Adaptive Merge Candidate List We can assume there are 8 merge candidates. We will take the first 5 merge candidates as the first subgroup and the last 3 merge candidates as the second subgroup (i.e., the last subgroup).

[0464] For the encoder, after constructing the Merge candidate list, as follows Figure 39 As shown, some Merge candidates are adaptively reordered in ascending order of Merge candidate cost.

[0465] More specifically, the template matching cost of the Merge candidates in all subgroups except the last one is calculated; then the Merge candidates except the last one are reordered in their own subgroups; finally, the final list of Merge candidates is obtained. Figure 56 The reordering process in the encoder is shown.

[0466] For the decoder, after constructing the Merge candidate list, such as Figure 40 As shown, some / no merge candidates are adaptively reordered in ascending order at the merge candidate cost. Figure 40 In this context, the subgroup containing the selected (transmitted via signal) Merge candidate is referred to as the selected subgroup. Figure 57 The reordering process in the decoder is shown.

[0467] More specifically, if the selected Merge candidate is in the last subgroup, the Merge candidate list construction process is terminated after deriving the selected Merge candidate, no reordering is performed, and the Merge candidate list is not changed; otherwise, the execution process is as follows: After deriving all Merge candidates in the selected subgroup, the Merge candidate list construction process is terminated; the template matching cost of the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained.

[0468] For both the encoder and decoder, the template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template.

[0469] When deriving reference samples for the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.

[0470] The reference samples (RT) for the template used for bidirectional prediction are obtained by using the reference samples of the template in reference list 0 as follows ( ) and reference samples of the template in reference list 1 ( It is derived by weighted averaging.

[0471]

[0472] The weights (8-w) of the reference templates in reference list 0 and the weights (w) of the reference templates in reference list 1 are determined by the BCW indices of the Merge candidates. The BCW indices equal to {0,1,2,3,4} correspond to w equal to {-2,3,4,5,10}, respectively.

[0473] If the Local Illumination Compensation (LIC) flag of the Merge candidate is true, the reference sample points of the template are derived using the LIC method.

[0474] Template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT.

[0475] The template size is 1. This means that the width of the left template and / or the height of the top template is 1.

[0476] If the encoding / decoding mode is MMVD, the Merge candidates used to derive the base Merge candidates are not reordered.

[0477] If the encoding / decoding mode is GPM, the Merge candidates used to derive the one-way prediction candidate list are not reordered.

[0478] 2.46. Geometric Prediction Patterns with Motion Vector Difference In Geometric Prediction Mode with Motion Vector Difference (GMVD), each geometric segment in GPM can determine whether GMVD is used. If GMVD is selected for a geometric region, the MV of that region is calculated as the sum of the MV of the merged candidates and the MVD. All other processing remains the same as in GPM.

[0479] Using GMVD, MVD is transmitted via signals in pairs of direction and distance. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD in GMVD is shifted left by 2 bits, just like in MMVD.

[0480] 2.47. Affine MMVD In affine MMVD, affine Merge Candidates (referred to as Basic Affine Merge Candidates) are selected, and the MV of the control points is further refined by the MVD information transmitted through the signals.

[0481] The MVD information for the MV of all control points is the same in one prediction direction.

[0482] When the starting MV is a bidirectional prediction MV and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, while the other reference POC is less than the current image's POC), the MV offset of the list 0 MV component added to the starting MV has the opposite value to the MV offset of the list 1 MV; otherwise, when the starting MV is a bidirectional prediction MV and both lists point to the same side of the current image (i.e., both reference POCs are greater than the current image's POC, or both are less than the current image's POC), the MV offset of the list 0 MV component added to the starting MV has the same value as the MV offset of the list 1 MV.

[0483] 2.48. Adaptive Decoder-Side Motion Vector Refinement (ADMVR) In ECM-2.0, if the selected merge candidate satisfies the DMVR condition, a multi-pass decoder-side motion vector refinement (DMVR) method is applied in the regular merge mode. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF).

[0484] The adaptive decoder-side motion vector refinement method consists of two new Merge modes, which are introduced to refine the motion vectors only in one direction (L0 or L1) of the bidirectional predictions of Merge candidates that satisfy the DMVR condition. A multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vectors; however, in the first pass (i.e., the PU level) of the DMVR, either MVD0 or MVD1 is set to zero.

[0485] Similar to the regular Merge pattern, the proposed Merge pattern derives its Merge candidates from spatially adjacent encoded / decoded blocks, TMVPs, non-adjacent blocks, HMVPs, and paired candidates. The difference lies in that only those satisfying the DMVR conditions are added to the candidate list. Both proposed Merge patterns use the same Merge candidate list (i.e., the ADMVR Merge list), and the Merge index is encoded / decoded in the same way as in the regular Merge pattern.

[0486] 2.49. Convolutional Cross-Component Model (CCCM) for Intra-Frame Prediction We propose applying a convolutional cross-component model (CCCM) in a spirit similar to the current CCLM model to predict chroma samples from reconstructed luminance samples. As with CCLM, when using chroma downsampling, the reconstructed luminance samples are downsampled to match a lower-resolution chroma grid.

[0487] Additionally, similar to CCLM, there are options for single-model or multi-model variants using CCCM. The multi-model variant uses two models: one derived for samples above the average luminance reference value, and the other for the remaining samples (following the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.

[0488] 2.49.1. Convolution Filter The proposed convolutional 7-tap filter consists of a 5-tap spatial component with a sign shape, a nonlinear term, and a bias term. The input of the 5-tap spatial component of the filter consists of the center (C) luminance sample that is co-located with the chrominance sample to be predicted, and its upper / north (N), lower / south (S), left / west (W), and right / east (E) neighbors, as shown below. Figure 58 The spatial portion of the convolution filter is shown.

[0489] The nonlinear term P is expressed as the square of the center luminance sample C and scaled to the range of sample values ​​for the content:

[0490] That is, for 10 bits of content, it is calculated as:

[0491] The bias term B represents the scalar offset between the input and output (similar to the offset term in CCLM) and is set to an intermediate chroma value (512 for 10-bit content).

[0492] The output of the filter is calculated as the filter coefficients c. i The convolution with the input values ​​is then limited to the range of valid chromaticity samples:

[0493] 2.49.2. Calculation of Filter Coefficients Filter coefficients c i It is calculated by minimizing the MSE between the predicted chromaticity samples and the reconstructed chromaticity samples in the reference region. Figure 59 The reference region is shown, consisting of six rows of chroma samples above and to the left of the PU. The reference region extends to the right by one PU width and below the PU boundary by one PU height. The region is adjusted to include only available samples. The expansion of the region shown in blue is necessary to support the "edge samples" of the plus-shaped spatial filter and is filled when in unavailable areas. Figure 59 The reference region (and its filling) used to derive the filter coefficients is shown.

[0494] MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chromaticity output. The autocorrelation matrix is ​​decomposed using LDL, and the final filter coefficients are computed using back-substitution. This process roughly follows the computation of ALF filter coefficients in ECM; however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer operations.

[0495] 2.49.3. Bitstream Signaling The use of this mode is signaled via PU-level flags encoded and decoded by CABAC. A new CABAC context is included to support this. When signaling is involved, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is signaled only when the intra-frame prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-mode CCCM).

[0496] 2.49.4. Encoder Operation The encoder performs two new RD checks in the chromaticity prediction mode loop: one for checking the single-model CCCM mode and one for checking the multi-model CCCM mode.

[0497] 2.50. Color Blending Version 2.50.1. Chroma blending with DM, Default, and MMLM modes. In ECM 6.0, there is a chroma blending method that allows the DM mode and the four default modes to be blended with the MMLM_LT mode, as shown below:

[0498] in These are predicted values ​​obtained by applying a non-LM model. These are predicted values ​​obtained by applying the MMLM_LT mode, and This is the final predicted value for the current chroma block. Preset weight combination { } is equal to {1, 3}, {3, 1} or {2, 2}.

[0499] 2.50.2. Chromaticity Blending with LDL Decomposition A novel chromaticity fusion method is proposed as follows:

[0500] in This represents the final predicted chromaticity sample points. This represents the reconstructed brightness sample points. This represents the predicted chroma samples obtained by applying a non-LM mode. For 10-bit content, It was set to equal 512. Model parameters. , and It is derived from the LDL decomposition method used in CCCM, based on the same adjacent sample points (two rows).

[0501] As shown in Table 13, there are three methods for color blending.

[0502] Table 13 – The proposed method's chromaticity blending mode

[0503] 3. Problem 1. In current codec designs, sub-segmentation techniques such as ISP can be applied only to the luma component during intra-frame prediction. However, this approach may limit compression efficiency because it does not adequately consider other components for deeper segmentation in order to achieve compression, especially when blocks or frames have different signal and texture distributions.

[0504] 2. In current codec designs, subsegmentation techniques such as ISP have not yet been extended to other fundamental codec tools, such as cross-component tools in the chroma component. However, combining subsegmentation with more codec tools could contribute a better trade-off between reconstruction quality and bit rate cost.

[0505] 4. Detailed Solution The detailed solutions below should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way.

[0506] In this disclosure, the term "block" may mean a codec block (CB), or a codec unit (CU), or a prediction block (PB), or a prediction unit (PU), or a transform block (TB), or a transform unit (TU), or a codec tree block (CTB), or a codec tree unit (CTU), or a rectangular region of samples / pixels.

[0507] In the following discussion, SatShift(x, n) is defined as:

[0508] Shift(x, n) is defined as .

[0509] In one example, offset0 and / or offset1 are set to (1 < 0.05).<n)> >1 or (1<<(n-1)). In another example, offset0 and / or offset1 are set to 0.

[0510] In another example, offset0 = offset1 = ((1<<n)> >1)-1 or ((1<<(n-1)))-1.

[0511] Clip3(min, max, x) is defined as:

[0512] In the following discussion, “picture,” “image,” and “frame” can have the same meaning.

[0513] In the following discussion, “splitting,” “breaking apart,” and “dividing” can have the same meaning.

[0514] Figures 60-69 Several example segmentation patterns according to some embodiments of this disclosure are shown. Figure 60 In the segmentation pattern shown, a block can be horizontally divided into sub-rectangles. Figure 61 In the segmentation pattern shown, a block can be vertically divided into sub-rectangles. Figure 62 In the segmentation pattern shown, a block can be divided into multiple sub-rectangles. Figure 63 In the segmentation pattern shown, a block can be divided into multiple sub-squares. Figure 64 In the shown segmentation pattern, a block can be horizontally divided into sub-squares. Figure 65 In the shown segmentation pattern, a block can be vertically divided into sub-squares. Figure 66 In the segmentation pattern shown, a block can be divided into horizontally arranged sub-trapezoidal shapes. Figure 67 In the shown segmentation pattern, a block can be divided into vertically arranged sub-trapezoidal shapes. Figure 68 In the segmentation pattern shown, the block can be divided into two sub-triangles. Figure 69 In the segmentation pattern shown, the block can be divided into three sub-triangles. Figure 70 In the segmentation pattern shown, a block can be divided into sub-trapezoidal and sub-triangle shapes. Figure 71 In the segmentation pattern shown, a block can be divided into multiple sub-trapezoidal and sub-triangle shapes.

[0515] sub-split 1. For color components other than luminance, instead of predicting the block as a whole, multiple sub-segments of the divided sub-blocks can be used for prediction using the proposed sub-segmentation method. The sub-segmentation function is denoted as F(), such as... , where p n It is the nth sub-segment.

[0516] a. In one example, more than one sub-segmentation pattern can be used.

[0517] i. In one example, whether and / or which sub-segmentation pattern will be used can be transmitted via signaling in the bitstream, or deduced using encoding / decoding information.

[0518] ii. In one example, the number of sub-segmentation patterns can depend on encoding / decoding information, such as: 1) Block dimension and / or block size and / or block depth.

[0519] 2) Strip / image type and / or segmentation tree type (single tree, or double tree, or local double tree).

[0520] 3) Temporal layer identifier.

[0521] 4) QP.

[0522] 5) The encoding / decoding mode of the current block.

[0523] 6) Encoding / decoding modes for neighboring blocks.

[0524] b. In one example, a subsegmentation method can refer to a set of predictions of geometric subsegments within a block. .

[0525] i. In one example,

[0526] ( 1 ) in B represents the prediction of the i-th geometric sub-segment, and B represents the prediction of the block.

[0527] 1) In one example, the total number of sub-segments n is a fixed value.

[0528] 2) In one example, the total number of sub-segments n is a value that varies depending on the context of the codec.

[0529] ii. In one example,

[0530] ( 2 ) in B represents the prediction of the i-th geometric sub-segment, B represents the prediction of the block, and n is the number of right shifts.

[0531] c. In one example, a sub-segmentation method can refer to segmenting into a set of predicted values. Nonlinear methods.

[0532] i. In one example, a non-linear segmentation method could employ a convolutional neural network.

[0533] 1) In one example,

[0534] ( 3 ) in It is the prediction of the i-th sub-segment. These are the weights of the convolutional layer. It is the bias of the convolutional layer.

[0535]

[0536] ( 4 ) in It is the segmentation method for i-seed segmentation.

[0537]

[0538] ( 5 ) It is a modified linear unit (ReLU), which is the activation function.

[0539] d. The above-disclosed It can be a prediction of the sub-segment of the current component.

[0540] i. The above-disclosed It can be a prediction of sub-segments with different components.

[0541] ii. The above-disclosed It can be a prediction of a sub-segment after filtering or scaling.

[0542] e. In one example, the total number of predicted values ​​n in (1) is equal to 2.

[0543] i. In one example, these two predicted values ​​( , )yes:

[0544] 1) In one example, the splitting method is a fixed predefined method.

[0545] 2) In one example, the segmentation method is a segmentation method that varies depending on the context of the codec.

[0546] f. In one example, a set of one or more partitioning methods can be transmitted via signals in a bitstream.

[0547] i. In one example, one or more sets of segmentation patterns may be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0548] ii. In one example, the first syntax element can be signaled to indicate whether the sub-segmentation method is applied.

[0549] iii. In one example, the second syntax element can be signaled to indicate which set of sub-segmentation patterns is applied or not applied.

[0550] 1) In one example, the second syntax element can only be signaled if the first syntax indicator sub-segment is applied.

[0551] iv. In one example, a single SE can be transmitted via a signal indicating whether subsegmentation is used and which set of patterns is used.

[0552] v. (Multiple) syntax elements (SE) can be encoded or decoded using fixed-length encoding / decoding, EG encoding / decoding, rounding (unary) encoding / decoding, etc.

[0553] vi. (Multiple) SEs can be encoded or decoded using at least one context in arithmetic encoding / decoding.

[0554] vii. (Multiple) SEs can be bypassed for encoding and decoding.

[0555] viii. Only when the sub-segmentation method is permitted can (multiple) SEs be transmitted via signaling.

[0556] ix. (Multiple) SEs can be transmitted via signals in SEI or VUI messages.

[0557] x. In one example, one or more sub-segmentation pattern sets can be transmitted via signaling in a predictable manner.

[0558] 1) In one example, the sub-segmentation pattern can be transmitted via signaling in a predictable manner.

[0559] xi. In one example, the set of sub-segmentation patterns for the first video unit may not be transmitted via signaling; instead, one or more sets of sub-segmentation patterns for the second video unit may be used via signaling, wherein the second video unit is encoded and decoded before the first video unit.

[0560] g. In one example, one or more sets of derived sub-segmentation patterns can be used.

[0561] i. In one example, the derived pattern can be obtained using reconstructed luminance samples and / or neighboring reconstructed luminance / chrominance samples and / or predicted chrominance samples.

[0562] h. In one example, the set of sub-segmentation patterns may depend on the color format and / or color components.

[0563] i. In one example, the set of sub-segmentation patterns can be the same for two chroma components.

[0564] 1) Alternatively, the sub-segmentation pattern sets can be different for the two chroma components.

[0565] ii. In one example, the derivation of the sub-segmentation mode can differ for different color formats.

[0566] 1) In one example, the luminance ISP mode can be used in 4:2:0 / 4:2:2 color formats, but they cannot be used in 4:4:4 color formats.

[0567] i. The multiple sub-segmentation patterns in the proposed method can differ in terms of encoding / decoding information, context, and background.

[0568] i. In one example, the sub-segmentation pattern can depend on the sub-picture, slice, strip, or block size of CTU / CU / PU / TU / CTB / CB / PB / TB.

[0569] 1) In one example, the sub-segmentation mode is p when the block size is 256×256, and the sub-segmentation mode is q when the block size is 128×128.

[0570] ii. In one example, the sub-segmentation pattern may depend on the video content of the video unit.

[0571] 1) In one example, the sub-segmentation mode is p when the video content is a natural sequence, and q when the video content is screen content (SCC).

[0572] j. In one example, whether and / or how the subsegmentation method and its weights are applied can be determined by signaling using at least one syntax element (SE).

[0573] i. In one example, syntax elements can be transmitted via signals in SPS / PPS / APS / image headers / strip headers, etc.

[0574] ii. In one example, the first SE can be signaled to indicate whether the sub-segmentation method is applied.

[0575] iii. In one example, a set of SEs can be transmitted via signals to indicate how the sub-segmentation pattern is predefined.

[0576] 1) In one example, each element in a set of SEs indicates a different pattern.

[0577] k. In one example, whether and / or how one or more sub-segmentation pattern sets from the above methods can be transmitted via signaling in a bitstream, or deduced using encoding / decoding information.

[0578] i. In one example, one or more set indices indicating which set of patterns to use can be transmitted via signaling in the bitstream.

[0579] 1) In one example, a single set index can be used to indicate the set of patterns used for the component.

[0580] 2) In one example, the indication of the pattern set for a component can be transmitted separately via signaling.

[0581] ii. In one example, set indices can be derived using encoding / decoding information.

[0582] 1) In one example, a template-based matching method can be used.

[0583] 2) In one example, the set of pattern indices for a component can be derived together.

[0584] 3) In one example, the set of pattern indices for a component can be derived separately.

[0585] 2. Instead of ISP dividing luma blocks into four or two sub-segments horizontally or vertically based on block size during intra-frame prediction, blocks of luma or other components (such as Cb or Cr) can be segmented into sub-segments based on the codec background, codec information, and codec context.

[0586] a. In one example, a block can be segmented into sub-segments based on the codec background, codec information, and codec context.

[0587] i. In one example, the encoding / decoding information may refer to reconstructed luminance samples and / or neighboring reconstructed luminance samples and / or neighboring reconstructed chrominance samples and / or predicted chrominance samples.

[0588] ii. In one example, the block can be divided into sub-rectangles.

[0589] 1) In one example, such as Figure 60 As shown, the block can be horizontally divided into sub-rectangles.

[0590] 2) In one example, such as Figure 61 As shown, the block can be vertically divided into sub-rectangles.

[0591] 3) In one example, such as Figure 62 As shown, the block can be divided into multiple sub-rectangles.

[0592] iii. In one example, the block can be divided into sub-squares.

[0593] 1) In one example, such as Figure 63 As shown, the block can be divided into multiple sub-squares.

[0594] 2) In one example, such as Figure 64 As shown, the block can be horizontally divided into sub-squares.

[0595] 3) In one example, such as Figure 65 As shown, the block can be vertically divided into sub-squares.

[0596] iv. In one example, the block can be divided into sub-trapezoidal shapes.

[0597] 1) In one example, such as Figure 66 As shown, the block can be divided into horizontally arranged sub-trapezoidal shapes.

[0598] 2) In one example, such as Figure 67 As shown, the block can be divided into vertically arranged sub-trapezoidal shapes.

[0599] v. In one example, a block can be divided into sub-triangles.

[0600] 1) In one example, such as Figure 68 As shown, the block can be divided into two sub-triangles.

[0601] 2) In one example, the block can be divided into three sub-triangles.

[0602] 3) In one example, such as Figure 69 As shown, the block can be divided into more than three sub-triangles.

[0603] vi. In one example, the block can be divided into several sub-trapezoidal and sub-triangle shapes.

[0604] 1) In one example, such as Figure 70 As shown, the block can be divided into sub-triangles and sub-trapezoidal shapes.

[0605] 2) In one example, such as Figure 71 As shown, the block can be divided into several sub-triangles and sub-trapezoidal shapes.

[0606] b. In one example, the sub-segmentation pattern can be transmitted via signaling in the bitstream.

[0607] i. In one example, the sub-segmentation pattern can be transmitted via signaling at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0608] ii. In one example, the first syntax element can be signaled to indicate whether the sub-segmentation method is applied.

[0609] iii. In one example, the second syntax element can be signaled to indicate which level of the sub-segment is applied or not.

[0610] 1) In one example, the second syntax element can only be signaled if the first syntax indicator sub-segment is applied.

[0611] iv. In one example, a single SE can be transmitted via a signal indicating whether a subsegment is used and at which level of the subsegment is used.

[0612] v. (Multiple) SEs can be encoded or decoded using fixed-length encoding / decoding, EG encoding / decoding, rounding (unary) encoding / decoding, etc.

[0613] vi. (Multiple) SEs can be encoded or decoded using at least one context in arithmetic encoding / decoding.

[0614] vii. (Multiple) SEs can be bypassed for encoding and decoding.

[0615] viii. Only when chromatic blending with median values ​​is permitted can (multiple) SEs be transmitted via signal.

[0616] ix. (Multiple) SEs can be transmitted via signals in SEI or VUI messages.

[0617] c. In one example, the same sub-segmentation pattern (such as a prediction pattern or a transformation pattern) can be used for multiple different blocks.

[0618] i. In one example, the mode can be transmitted / derived via signaling in the same way as the overall codec block.

[0619] ii. In one example, the mode may be transmitted / derived via signaling in a manner different from that of the overall codec block.

[0620] 1) In one example, the allowed patterns for a sub-segment block can be a subset of all patterns.

[0621] a) For example, the allowed modes for sub-segmented chroma blocks can be cross-component prediction modes, such as DM or DIMD or CCLM or CCCM.

[0622] b) For example, sub-segmented chroma blocks must be encoded and decoded using specific modes such as CCLM, without signaling.

[0623] d. Multiple sub-segmentation patterns can be applied to the proposed sub-segmentation method.

[0624] i. In one example, whether and / or how to apply the sub-segmentation method and its multiple patterns can be transmitted via a signal using at least one syntax element (SE).

[0625] 1) In one example, syntax elements can be transmitted via signals in SPS / PPS / APS / image header / strip header / CTU / CU, etc.

[0626] 2) In one example, the first syntax element can be signaled to indicate whether a sub-segmentation with multiple patterns is applied.

[0627] 3) In one example, the second syntax element can be signaled to indicate which set of sub-segmentation patterns is applied or not applied.

[0628] a) In one example, the second syntax element can only be transmitted via signaling if the first syntax indicates that a sub-segmentation method with multiple patterns is applied.

[0629] 4) In one example, a single SE can be transmitted via a signal indicating whether subsegmentation is used and which set of patterns is used.

[0630] 5) Multiple SEs can be encoded or decoded using fixed-length encoding / decoding, EG encoding / decoding, rounding (unary) encoding / decoding, etc.

[0631] 6) (Multiple) SEs can be encoded or decoded using at least one context in arithmetic encoding and decoding.

[0632] 7) (Multiple) SEs can be bypassed for encoding and decoding.

[0633] 8) Multiple SEs may be transmitted via signal only if subsegments with multiple modes are permitted to be used.

[0634] 9) (Multiple) SEs can be transmitted via signals in SEI or VUI messages.

[0635] e. In one example, a block of Cr can be segmented into sub-segments based on the codec background, codec information, and codec context.

[0636] f. In one example, blocks in an RGB component can be segmented into sub-segments based on the codec background, codec information, and codec context.

[0637] g. In one example, some component blocks can be segmented into sub-segments during intra-frame prediction based on the codec background, codec information, and codec context.

[0638] h. In one example, some component blocks can be segmented into sub-segments during inter-frame prediction based on the codec background, codec information, and codec context.

[0639] 3. In one example, subsegments can be encoded / decoded sequentially.

[0640] a. For example, the first sub-segment is encoded / decoded, and then the second sub-segment can be encoded / decoded, which may depend on the reconstructed samples of the first sub-segment.

[0641] b. Alternatively, subsegments can be encoded / decoded in parallel.

[0642] i. For example, the first sub-segment is encoded / decoded, and the second sub-segment can be encoded / decoded independently, without depending on the first sub-segment.

[0643] 4. In one example, the encoding and decoding order of sub-segments can be fixed.

[0644] a. The encoding / decoding order can be from left to right.

[0645] b. The encoding / decoding order can be from top to bottom.

[0646] c. The encoding / decoding order can be from right to left.

[0647] d. The encoding / decoding order can be from bottom to top.

[0648] e. The encoding / decoding order can be transmitted via signals within the bitstream.

[0649] f. The encoding / decoding order can be deduced without signaling.

[0650] 5. A sub-segmentation method is proposed that can be combined with different prediction tools.

[0651] a. In one example, prediction patterns from linear or nonlinear models can be used in conjunction with sub-segmentation methods.

[0652] b. In one example, the prediction pattern from the angle prediction pattern can be used in conjunction with the sub-segmentation method.

[0653] i. In one example, in the intra-chroma prediction pipeline, the angular prediction mode of the sub-segment can be derived from the co-position luma block or the neighboring chroma block.

[0654] 1) In one example, such as Figure 72 As shown, the prediction mode of sub-segmentation can be derived from the co-position brightness block based on the position of the sub-segmentation.

[0655] ii. In one example, if the chromaticity components are predicted using a chromaticity fusion method, the angular prediction pattern of the sub-segment can be derived from the co-position luminance block or the neighboring chromaticity block.

[0656] 1) In one example, such as Figure 72As shown, the prediction mode of sub-segmentation can be derived from the co-position brightness block based on the position of the sub-segmentation.

[0657] c. In one example, the IBC or template matching method can be used in conjunction with the sub-segmentation method.

[0658] i. In one example, in a chroma intra-prediction pipeline with IBC or TMP mode, the block vector or motion vector of a sub-segment can be derived from a co-position luma block or a neighboring chroma block.

[0659] 1) In one example, such as Figure 73 As shown, the block vector or motion vector of a sub-segment can be derived from the co-position brightness block based on the position of the sub-segment.

[0660] d. In one example, non-LM prediction patterns can be used in conjunction with sub-segmentation methods.

[0661] e. In one example, the CCLM prediction pattern can be used in conjunction with a sub-segmentation method.

[0662] i. In one example, in a chroma intra-frame prediction pipeline with CCLM mode, the luma or chroma reference points for sub-segments can be generated from co-located luma blocks or neighboring chroma blocks.

[0663] 1) In one example, the luminance or chrominance reference points of a subsegment can be derived from the co-position luminance block based on the location of the subsegment.

[0664] f. In one example, the CCLM-slope prediction pattern can be used in conjunction with a sub-segmentation method.

[0665] i. In one example, in a chroma intra-frame prediction pipeline with CCLM slope flags, the luma or chroma reference points for a sub-segment can be generated from a co-located luma block or a neighboring chroma block.

[0666] 1) In one example, the luminance or chrominance reference points of a subsegment can be derived from the co-position luminance block based on the location of the subsegment.

[0667] g. In one example, the GLM prediction pattern can be used in conjunction with a sub-segmentation method.

[0668] i. In one example, in a chroma intra-frame prediction pipeline with GLM mode, the luma or chroma reference points for sub-segments can be generated from co-located luma blocks or neighboring chroma blocks.

[0669] 1) In one example, the luminance or chrominance reference points of a subsegment can be derived from the co-position luminance block based on the location of the subsegment.

[0670] h. In one example, the CCCM prediction pattern can be used in conjunction with the sub-segmentation method.

[0671] i. In one example, in a chroma intra-frame prediction pipeline with the CCCM flag, the luma or chroma reference points for a sub-segment can be generated from a co-located luma block or a neighboring chroma block.

[0672] 1) In one example, the luminance or chrominance reference points of a subsegment can be derived from the co-position luminance block based on the location of the subsegment.

[0673] i. In one example, TMP / DIMD / TIMD / MRL / MIP prediction patterns can be used in conjunction with sub-segmentation methods.

[0674] 6. When using prediction tools to decompress and compress video units, the above methods can also be applied to sub-segmentation methods for the luminance component in decoders and encoders.

[0675] General aspects 7. Whether and / or how the methods disclosed above can be applied to signal transmission at the block level / sequence level / picture group level / picture level / strip level / piece group level, such as in the codec structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0676] 8. Whether and / or how the methods disclosed above are applied may depend on the encoded / decoded information, such as block size, color format, single / double tree segmentation, color components, and stripe / image type.

[0677] 9. The methods presented in this document can be used in other codec tools that require subsegmentation.

[0678] 5. Examples Further details of embodiments of this disclosure related to block partitioning will now be described. The embodiments of this disclosure should be considered as examples for explaining general concepts and should not be interpreted in a narrow sense. Furthermore, these embodiments may be applied individually or in any combination.

[0679] As used herein, the term "block" can refer to color components, sub-pictures, pictures, stripes, slices, codec tree units (CTUs), CTU rows, CTU groups, codec units (CUs), prediction units (PUs), transform units (TUs), codec tree blocks (CTBs), codec blocks (CBs), prediction blocks (PBs), transform blocks (TBs), sub-blocks of video blocks, sub-regions within video blocks, and video processing units comprising multiple samples / pixels, etc. Blocks can be rectangular or non-rectangular.

[0680] Figure 74 A flowchart of a method 7400 for video processing according to some embodiments of the present disclosure is shown. Method 7400 can be implemented during the conversion between a current video block and a video bitstream. Figure 74 As shown, method 7400 begins at 7402, where the prediction for the current video block is determined based on multiple sub-segments of the current video block and at least one intra-frame prediction scheme. The current video block is a chroma block, and the multiple sub-segments are obtained by segmenting the current video block.

[0681] For example, multiple sub-segments can be obtained by dividing the current video block based on a segmentation pattern. In one example, each of the multiple sub-segments is rectangular in shape. That is, the block can be divided into multiple rectangular sub-segments, such as... Figure 62 As shown. Alternatively, each of the multiple sub-segments can be square in shape. That is, a block can be divided into multiple square sub-segments, such as... Figure 63 As shown. Reference Figures 60-61 and Figures 64-71 This document describes some additional example segmentation patterns. It should be understood that the possible implementations of the segmentation patterns described herein are merely illustrative and should not be construed as limiting this disclosure in any way.

[0682] At least one intra-frame prediction scheme includes an angle prediction mode, an intra-block copy (IBC) mode, a template matching mode, and / or a cross-component prediction mode. For example, referencing Figure 6 Several possible angle prediction modes are described. Template matching modes can be the intra-template matching (TMP) mode described in Section 2.19 above. Cross-component prediction modes can include cross-component linear model (CCLM) mode, CCLM-slope mode, gradient linear model (GLM) mode, convolution cross-component model (CCCM) mode, etc.

[0683] At 7404, the conversion is performed based on a prediction for the current video block. In some embodiments, the conversion may include encoding the current video block into a bitstream. Alternatively or additionally, the conversion may include decoding the current video block from the bitstream. It should be understood that the above descriptions and / or examples are for illustrative purposes only. The scope of this disclosure is not limited in this respect.

[0684] In light of the above, the sub-segmentation scheme is permitted to be used in combination with angle prediction mode, intra-block copy (IBC) mode, template matching mode, and / or cross-component prediction mode for chroma blocks. Compared to conventional solutions, the proposed method can advantageously improve encoding / decoding efficiency and quality, especially when blocks have different signal and texture distributions.

[0685] In some embodiments, at least one intra-frame prediction scheme includes determining an angular prediction mode for a first sub-segment of a plurality of sub-segments. Furthermore, the angular prediction mode for the first sub-segment is determined based on either a co-occurrence luma sub-segment or a neighboring chroma sub-segment of the first sub-segment. In one example, the angular prediction mode for the first sub-segment may be determined as the angular prediction mode for the co-occurrence luma sub-segment of the first sub-segment. Alternatively, the angular prediction mode for the first sub-segment may be determined as the angular prediction mode for the neighboring chroma sub-segment of the first sub-segment. For example, as... Figure 72 As shown, the angle prediction mode for the first sub-segment can be derived from the co-position brightness sub-segment of the first sub-segment based on the position of the first sub-segment.

[0686] In some embodiments, the prediction for the current video block is determined based on a chroma fusion mode. In this case, the final prediction for the first sub-segment can be determined by fusing multiple predictions for the first sub-segment, and the multiple predictions include the prediction for the first sub-segment determined using the angle prediction mode described above.

[0687] In some additional embodiments, at least one intra-frame prediction scheme includes an IBC mode for determining predictions for a second sub-segment among a plurality of sub-segments. Additionally, the block vector (BV) for the IBC mode is determined based on either a co-occurrence luma sub-segment or a neighboring chroma sub-segment of the second sub-segment. In one example, the block vector for the IBC mode may be determined as a block vector for a co-occurrence luma sub-segment of the second sub-segment. Alternatively, the block vector for the IBC mode may be determined as a block vector for a neighboring chroma sub-segment of the second sub-segment. For example, as... Figure 73 As shown, the block vector for IBC mode is derived from the co-positional luminance sub-segment of the second sub-segment based on the position of the second sub-segment.

[0688] In some embodiments, at least one intra-frame prediction scheme includes determining a template matching mode for predicting a third sub-segment among a plurality of sub-segments. Additionally, the block vector for the template matching mode is determined based on either a co-occurrence luma sub-segment or a neighboring chroma sub-segment of the third sub-segment. In one example, the block vector for the template matching mode may be determined as a block vector for the co-occurrence luma sub-segment of the third sub-segment. Alternatively, the block vector for the template matching mode may be determined as a block vector for the neighboring chroma sub-segment of the third sub-segment. For example, as... Figure 73 As shown, the block vector used for template matching mode is derived from the co-position brightness sub-segment of the third sub-segment based on the position of the third sub-segment.

[0689] In some embodiments, at least one intra-frame prediction scheme includes a cross-component prediction mode for determining a prediction for a fourth sub-segment among a plurality of sub-segments. Additionally, a luminance reference sample for the cross-component prediction mode is determined based on a co-position luminance sub-segment of the fourth sub-segment. For example, the luminance reference sample may be a reconstructed luminance sample, etc. For example, the luminance reference sample for the cross-component prediction mode is derived from the co-position luminance sub-segment of the fourth sub-segment based on the position of the fourth sub-segment.

[0690] Additionally or alternatively, the chromaticity reference points for the cross-component prediction mode are determined based on the neighboring chromaticity sub-segments of the fourth sub-segment. For example, the chromaticity reference points may be reconstructed chromaticity points, etc. In one example, the chromaticity reference points for the cross-component prediction mode are determined as chromaticity reference points for the neighboring chromaticity sub-segments of the fourth sub-segment. For example, the chromaticity reference points for the cross-component prediction mode are derived from the neighboring chromaticity sub-segments of the fourth sub-segment based on the position of the fourth sub-segment.

[0691] In view of the above, the solutions according to some embodiments of this disclosure can advantageously improve encoding and decoding efficiency and encoding and decoding quality.

[0692] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. In this method, a prediction for a current video block is determined based on multiple sub-segments of the current video block and at least one intra-prediction scheme. The current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: angular prediction mode, intra-block copy (IBC) mode, template matching mode, or cross-component prediction mode. Furthermore, the bitstream is generated based on the prediction for the current video block.

[0693] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. In this method, a prediction for a current video block is determined based on multiple sub-segments of the current video block and at least one intra-prediction scheme. The current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: angular prediction mode, intra-block copy (IBC) mode, template matching mode, or cross-component prediction mode. Furthermore, the bitstream is generated based on the prediction for the current video block and stored in a non-transitory computer-readable recording medium.

[0694] Implementations of this disclosure may be described in accordance with the following entries, and its features may be combined in any reasonable manner.

[0695] Item 1. A method for video processing, comprising: a conversion between a current video block and a bitstream of the video; determining a prediction for the current video block based on a plurality of sub-segments of the current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the plurality of sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme comprises at least one of: an angle prediction mode, an intra-block copy (IBC) mode, a template matching mode, or a cross-component prediction mode; and performing the conversion based on the prediction for the current video block.

[0696] Item 2. The method according to Item 1, wherein the at least one intra-frame prediction scheme includes determining an angular prediction mode for a first sub-segment of the plurality of sub-segments, and the angular prediction mode for the first sub-segment is determined based on a co-position luminance sub-segment of the first sub-segment or a neighboring chrominance sub-segment of the first sub-segment.

[0697] Item 3. The method according to Item 2, wherein the angle prediction mode for the first sub-segment is determined to be either the angle prediction mode for the co-position luminance sub-segment of the first sub-segment or the angle prediction mode for the neighboring chromaticity sub-segment of the first sub-segment.

[0698] Item 4. The method according to any one of items 2-3, wherein the angle prediction mode for the first sub-segment is derived from the co-position brightness sub-segment of the first sub-segment based on the position of the first sub-segment.

[0699] Item 5. The method according to any one of items 2-4, wherein the prediction for the current video block is determined based on a chroma fusion mode.

[0700] Item 6. The method according to Item 5, wherein the final prediction for the first sub-segment is determined by fusing multiple predictions for the first sub-segment, and the multiple predictions include the prediction for the first sub-segment determined using the angle prediction mode.

[0701] Item 7. The method according to any one of items 1-6, wherein the at least one intra-frame prediction scheme includes determining an IBC mode for a second sub-segment of the plurality of sub-segments, and the block vector for the IBC mode is determined based on a co-position luma sub-segment of the second sub-segment or a neighboring chroma sub-segment of the second sub-segment.

[0702] Item 8. The method according to Item 7, wherein the block vector for the IBC mode is determined as a block vector for the co-position luminance sub-segment of the second sub-segment or a block vector for the neighboring chrominance sub-segment of the second sub-segment.

[0703] Item 9. The method according to any one of items 7-8, wherein the block vector for the IBC mode is derived from the co-positional luminance sub-segment of the second sub-segment based on the position of the second sub-segment.

[0704] Item 10. The method according to any one of items 1-9, wherein the at least one intra-frame prediction scheme includes determining a template matching mode for determining a prediction for a third sub-segment among the plurality of sub-segments, and a block vector for the template matching mode is determined based on a co-position luma sub-segment of the third sub-segment or a neighboring chroma sub-segment of the third sub-segment.

[0705] Item 11. The method according to Item 10, wherein the block vector for the template matching mode is determined as a block vector for the co-position luminance sub-segment of the third sub-segment or a block vector for the neighboring chrominance sub-segment of the third sub-segment.

[0706] Item 12. The method according to any one of items 10-11, wherein the block vector for the template matching mode is derived from the co-position luminance sub-segment of the third sub-segment based on the position of the third sub-segment.

[0707] Item 13. The method according to any one of Items 1-12, wherein the cross-component prediction mode includes at least one of the following: cross-component linear model (CCLM) mode, CCLM-slope mode, gradient linear model (GLM) mode, or convolutional cross-component model (CCCM) mode.

[0708] Item 14. The method according to any one of items 1-13, wherein the at least one intra-frame prediction scheme includes a cross-component prediction mode for determining a prediction for a fourth sub-segment among the plurality of sub-segments, and a luminance reference sample for the cross-component prediction mode is determined based on a co-position luminance sub-segment of the fourth sub-segment.

[0709] Item 15. The method according to Item 14, wherein the luminance reference sample for the cross-component prediction mode is derived from the co-position luminance sub-segment of the fourth sub-segment based on the position of the fourth sub-segment.

[0710] Item 16. The method according to any one of items 1-15, wherein the at least one intra-frame prediction scheme includes a cross-component prediction mode for determining a prediction for a fourth sub-segment among the plurality of sub-segments, and chroma reference points for the cross-component prediction mode are determined based on neighboring chroma sub-segments of the fourth sub-segment.

[0711] Item 17. The method according to Item 16, wherein the chromaticity reference point for the cross-component prediction mode is determined as the chromaticity reference point for the neighboring chromaticity sub-segment of the fourth sub-segment.

[0712] Item 18. The method according to any one of items 16-17, wherein the chromaticity reference points for the cross-component prediction mode are derived from the neighboring chromaticity sub-segments of the fourth sub-segment based on the position of the fourth sub-segment.

[0713] Item 19. The method according to any one of items 1-18, wherein the shape of each of the plurality of sub-segments is a rectangle or a square.

[0714] Item 20. The method according to any one of items 1-19, wherein the conversion includes encoding the current video block into the bitstream.

[0715] Item 21. The method according to any one of items 1-19, wherein the conversion includes decoding the current video block from the bitstream.

[0716] Item 22. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-21.

[0717] Item 23. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of items 1-21.

[0718] Item 24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: determining a prediction for the current video block based on a plurality of sub-segments of a current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the plurality of sub-segments being obtained by segmenting the current video block, and the at least one intra-prediction scheme comprising at least one of: an angle prediction mode, an intra-block copy (IBC) mode, a template matching mode, or a cross-component prediction mode; and generating the bitstream based on the prediction for the current video block.

[0719] Item 25. A method for storing a bitstream of video, comprising: determining a prediction for the current video block based on a plurality of subsegments of a current video block and at least one intra-frame prediction scheme, wherein the current video block is a chroma block, the plurality of subsegments being obtained by segmenting the current video block, and the at least one intra-frame prediction scheme comprising at least one of: an angle prediction mode, an intra-block copy (IBC) mode, a template matching mode, or a cross-component prediction mode; generating the bitstream based on the prediction for the current video block; and storing the bitstream in a non-transitory computer-readable recording medium.

[0720] Example device Figure 75 A block diagram of a computing device 7500 in which various embodiments of the present disclosure may be implemented is shown. The computing device 7500 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0721] It should be understood that, Figure 75 The computing device 7500 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0722] like Figure 75 As shown, computing device 7500 includes general-purpose computing device 7500. Computing device 7500 may include at least one or more processors or processing units 7510, memory 7520, storage unit 7530, one or more communication units 7540, one or more input devices 7550, and one or more output devices 7560.

[0723] In some embodiments, the computing device 7500 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 7500 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0724] Processing unit 7510 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 7520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to provide the parallel processing capability of computing device 7500. Processing unit 7510 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0725] Computing device 7500 typically includes various computer storage media. Such media can be any media accessible by computing device 7500, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 7520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 7530 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 7500.

[0726] The computing device 7500 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 75 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0727] The communication unit 7540 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 7500 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 7500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0728] Input device 7550 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 7560 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 7540, computing device 7500 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 7500 can also communicate with one or more devices that enable a user to interact with computing device 7500, or, if needed, with any device (e.g., network card, modem, etc.) that enables computing device 7500 to communicate with one or more other computing devices. This communication can be performed via an input / output (I / O) interface (not shown).

[0729] In some embodiments, some or all of the components of computing device 7500 may be arranged in a cloud computing architecture, rather than integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0730] In embodiments of this disclosure, computing device 7500 can be used to implement video encoding / decoding. Memory 7520 may include one or more video codec modules 7525 having one or more program instructions. These modules are accessible and executable by processing unit 7510 to perform the functions of the various embodiments described herein.

[0731] In an example embodiment of performing video encoding, input device 7550 may receive video data as input 7570 to be encoded. The video data may be processed, for example, by video codec module 7525 to generate an encoded bitstream. The encoded bitstream may be provided as output 7580 via output device 7560.

[0732] In an example embodiment of performing video decoding, input device 7550 may receive an encoded bitstream as input 7570. The encoded bitstream may be processed, for example, by a video codec module 7525 to generate decoded video data. The decoded video data may be provided as output 7580 via output device 7560.

[0733] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For the conversion between a current video block and the bitstream of the video, a prediction for the current video block is determined based on multiple sub-segments of the current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: angle prediction mode, intra-block copy (IBC) mode, template matching mode, or cross-component prediction mode. as well as The transformation is performed based on the prediction for the current video block.

2. The method of claim 1, wherein the at least one intra-frame prediction scheme includes determining an angle prediction mode for a first sub-segment of the plurality of sub-segments, and the angle prediction mode for the first sub-segment is determined based on a co-position luminance sub-segment of the first sub-segment or a neighboring chrominance sub-segment of the first sub-segment.

3. The method of claim 2, wherein the angle prediction mode for the first sub-segment is determined to be either the angle prediction mode for the co-position luminance sub-segment of the first sub-segment or the angle prediction mode for the neighboring chromaticity sub-segment of the first sub-segment.

4. The method according to any one of claims 2-3, wherein the angle prediction mode for the first sub-segment is derived from the co-position brightness sub-segment of the first sub-segment based on the position of the first sub-segment.

5. The method according to any one of claims 2-4, wherein the prediction for the current video block is determined based on a chroma fusion mode.

6. The method of claim 5, wherein the final prediction for the first sub-segment is determined by fusing multiple predictions for the first sub-segment, and the multiple predictions include the prediction for the first sub-segment determined using the angle prediction mode.

7. The method according to any one of claims 1-6, wherein the at least one intra-frame prediction scheme includes determining an IBC mode for a second sub-segment of the plurality of sub-segments, and the block vector for the IBC mode is determined based on a co-position luma sub-segment of the second sub-segment or a neighboring chroma sub-segment of the second sub-segment.

8. The method of claim 7, wherein the block vector for the IBC mode is determined as a block vector for the co-position luminance sub-segment of the second sub-segment or a block vector for the neighboring chrominance sub-segment of the second sub-segment.

9. The method according to any one of claims 7-8, wherein the block vector for the IBC mode is derived from the co-positional luminance sub-segment of the second sub-segment based on the position of the second sub-segment.

10. The method according to any one of claims 1-9, wherein the at least one intra-frame prediction scheme includes determining a template matching mode for predicting a third sub-segment among the plurality of sub-segments, and a block vector for the template matching mode is determined based on a co-position luma sub-segment of the third sub-segment or a neighboring chroma sub-segment of the third sub-segment.

11. The method of claim 10, wherein the block vector for the template matching mode is determined to be a block vector for the co-position luminance sub-segment of the third sub-segment or a block vector for the neighboring chrominance sub-segment of the third sub-segment.

12. The method according to any one of claims 10-11, wherein the block vector for the template matching mode is derived from the co-position luminance sub-segment of the third sub-segment based on the position of the third sub-segment.

13. The method according to any one of claims 1-12, wherein the cross-component prediction mode comprises at least one of the following: Cross-component linear model (CCLM) mode, CCLM - Slope Mode Gradient Linear Model (GLM) pattern, or Convolutional Cross-Component Model (CCCM) mode.

14. The method according to any one of claims 1-13, wherein the at least one intra-frame prediction scheme includes a cross-component prediction mode for determining a prediction for a fourth sub-segment among the plurality of sub-segments, and a luminance reference sample for the cross-component prediction mode is determined based on a co-position luminance sub-segment of the fourth sub-segment.

15. The method of claim 14, wherein the luminance reference sample for the cross-component prediction mode is derived from the co-position luminance sub-segment of the fourth sub-segment based on the position of the fourth sub-segment.

16. The method according to any one of claims 1-15, wherein the at least one intra-frame prediction scheme includes a cross-component prediction mode for determining a prediction for a fourth sub-segment among the plurality of sub-segments, and chroma reference points for the cross-component prediction mode are determined based on neighboring chroma sub-segments of the fourth sub-segment.

17. The method of claim 16, wherein the chromaticity reference point for the cross-component prediction mode is determined as the chromaticity reference point for the neighboring chromaticity sub-segment of the fourth sub-segment.

18. The method of any one of claims 16-17, wherein the chromaticity reference points for the cross-component prediction mode are derived from the neighboring chromaticity sub-segments of the fourth sub-segment based on the position of the fourth sub-segment.

19. The method according to any one of claims 1-18, wherein the shape of each of the plurality of sub-segments is a rectangle or a square.

20. The method according to any one of claims 1-19, wherein the conversion comprises encoding the current video block into the bitstream.

21. The method according to any one of claims 1-19, wherein the conversion comprises decoding the current video block from the bitstream.

22. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-21.

23. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-21.

24. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: A prediction for the current video block is determined based on multiple sub-segments of the current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: angle prediction mode, intra-block copy (IBC) mode, template matching mode, or cross-component prediction mode. as well as The bitstream is generated based on the prediction for the current video block.

25. A method for storing a bitstream of video, comprising: A prediction for the current video block is determined based on multiple sub-segments of the current video block and at least one intra-prediction scheme, wherein the current video block is a chroma block, the multiple sub-segments are obtained by segmenting the current video block, and the at least one intra-prediction scheme includes at least one of the following: angle prediction mode, intra-block copy (IBC) mode, template matching mode, or cross-component prediction mode. The bitstream is generated based on the prediction for the current video block; as well as The bitstream is stored in a non-transitory computer-readable recording medium.