Methods, apparatus and media for video processing

By using Spatial Geometric Partitioning Mode (SGPM) candidate and Merge mode encoding and decoding, the problem of improving encoding and decoding efficiency in existing technologies is solved, and more efficient video encoding and decoding performance is achieved.

CN122139362APending Publication Date: 2026-06-02DOUYIN VISION CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2024-11-06
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially in spatial geometry prediction modes, where it is difficult to further improve encoding and decoding performance.

Method used

The Spatial Geometric Partitioning (SGPM) candidate is adopted to construct the prediction or reconstruction of video units through historical or neighbor information. The Merge mode is used to encode and decode video units to improve encoding and decoding performance.

Benefits of technology

The application of spatial geometric segmentation mode improves the encoding and decoding efficiency of video and enhances the performance of video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122139362A_ABST
    Figure CN122139362A_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure provide a solution for video processing. A method for video processing is proposed. The method includes: a conversion between video units and a bitstream of video; obtaining one or more Spatial Geometric Partitioning Pattern (SGPM) candidates for video units based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of video units based on the one or more SGPM candidates; and performing the conversion based on the prediction or construction of video units.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to spatial geometry prediction modes with Merge patterns. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. This method includes: a conversion between video units and a bitstream of the video; obtaining one or more Spatial Geometric Partitioning Pattern (SGPM) candidates for the video units based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; and performing the conversion based on the prediction or construction of the video units. This approach helps improve the encoding and decoding performance of SGPM.

[0005] In the second aspect, another method for video processing is proposed. This method includes: a conversion between video units and the video bitstream; constructing a candidate list for video units based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in SGPM Merge mode; obtaining a prediction or reconstruction of the video units based on the candidate list; and performing the conversion based on the prediction or reconstruction of the video units. This approach helps improve the encoding and decoding performance of SGPM.

[0006] In a third aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform a method according to either the first or second aspect of this disclosure.

[0007] In a fourth aspect, a non-transitory computer-readable storage medium is provided. This non-transitory computer-readable storage medium stores instructions that cause a processor to perform a method according to the first or second aspect of this disclosure.

[0008] In a fifth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: obtaining one or more spatial geometric segmentation pattern (SGPM) candidates for video units of the video based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; and generating a bitstream based on the prediction or construction of the video units.

[0009] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores video bitstreams generated by a method performed by an apparatus for video processing. The method includes: constructing a candidate list of video units for the video based on one or more spatial geometric partitioning mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in SGPM Merge mode; obtaining predictions or reconstructions of the video units based on the candidate list; and generating a bitstream based on the predictions or reconstructions of the video units.

[0010] In a seventh aspect, a method for storing a bitstream of video is proposed. The method includes: obtaining one or more Spatial Geometric Partition Pattern (SGPM) candidates for video units of a video based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; generating a bitstream based on the prediction or construction of the video units; and storing the bitstream in a non-transitory computer-readable recording medium.

[0011] In the eighth aspect, a method for storing a bitstream of video is proposed. The method includes: constructing a candidate list of video units for a video based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in SGPM Merge mode; obtaining a prediction or reconstruction of the video units based on the candidate list; generating a bitstream based on the prediction or reconstruction of the video units; and storing the bitstream in a non-transitory computer-readable recording medium.

[0012] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0013] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0014] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 An example of an encoder block diagram is shown; Figure 5 67 intra-frame prediction modes are shown; Figure 6 Reference samples for wide-angle intra-frame prediction are shown; Figure 7 This illustrates the problem of discontinuity when the orientation exceeds 45°; Figure 8 The spatial GPM candidates are shown; Figure 9 The GPM template is shown; Figure 10 GPM mixing is shown; Figure 11 The locations of the spatial merge candidates are shown; Figure 12 The candidate pairs considered for redundancy checks of spatial merge candidates are shown. Figure 13 This is a schematic diagram of motion vector scaling for time-domain Merge candidates; Figure 14 Candidate positions C0 and C1 for the temporal Merge candidate are shown; Figure 15 The VVC spatial neighboring blocks of the current block are shown; Figure 16 This is a schematic diagram of the virtual block in the i-th round of search; Figure 17 An example of GPM partitioning grouped at the same angle is shown; Figure 18 The unidirectional prediction MV selection for geometric segmentation patterns is shown; Figure 19 The bending weights using the geometric segmentation pattern are shown. An example of generation; Figure 20 The spatial neighbor block used to derive spatial merge candidates is shown; Figure 21 This illustrates template matching performed on the search area surrounding the initial MV; Figure 22The location, type, and transformation type of the SBT are shown; Figure 23 The neighboring sample points used to calculate SAD are shown; Figure 24 The neighboring samples used to calculate SAD for sub-CU level motion information are shown; Figure 25 The sorting process is shown; Figure 26 The reordering process in the encoder is shown; Figure 27 The reordering process in the decoder is shown; Figure 28 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 29 A flowchart of a method for video processing according to embodiments of the present disclosure is shown; and Figure 30 A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0015] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation

[0016] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0017] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0018] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0019] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0020] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0021] Example Environment Figure 1 This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0022] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0023] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0024] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0025] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0026] Figure 2 This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0027] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0028] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0029] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0030] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.

[0031] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0032] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0033] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0034] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0035] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating a reference image in list 0 or list 1 that includes the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0036] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images in lists 0 and 1 that include multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0037] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0038] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0039] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0040] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0041] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0042] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0043] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0044] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0045] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0046] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0047] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0048] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0049] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0050] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0051] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0052] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0053] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0054] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0055] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some respects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0056] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0057] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0058] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.

[0059] 1. Brief Overview This disclosure relates to video codec techniques. Specifically, it relates to Spatial Geometric Prediction Mode (SGPM) and SGPM Merge Mode, as well as other codec tools in image / video codecs. This disclosure can be applied to existing video codec standards such as HEVC or VVC. It may also be applicable to future video codec standards or video codecs.

[0060] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.

[0061] 2.1. Encoder / decoder streams of typical video codecs Figure 4 An example of a VVC encoder block diagram is shown, comprising three loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Compensation (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding compensation and applying finite impulse response (FIR) filters, respectively, and by utilizing the side information from the encoding and decoding through signal transmission compensation and filter coefficients. ALF is located in the last processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.

[0062] 2.2. Intra-mode encoding and decoding with 67 intra-prediction modes Figure 5 Sixty-seven intra-frame prediction modes are shown. This is to capture arbitrary edge directions presented in natural video, such as... Figure 5 As shown, the number of directional intra-prediction modes has been expanded from 33 used in HEVC to 65, while the planar and DC modes remain unchanged. These denser directional intra-prediction modes are applicable to all block sizes and both luma and chroma intra-prediction.

[0063] In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division is needed to generate intra-prediction values ​​using DC mode. In VVC, blocks can have rectangular shapes, which typically requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks.

[0064] 2.2.1. Wide-angle intra-frame prediction Although 67 modes are defined in VVC, the precise prediction direction for a given intra-prediction mode index depends on the block shape. Regular angular intra-prediction directions are defined clockwise from 45 degrees to -135 degrees. In VVC, for non-square blocks, several regular angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are transmitted via signaling using the original mode index, which is then remapped to the wide-angle mode index after resolution. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding / decoding method remains unchanged.

[0065] To support these predicted directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as follows: Figure 6 As shown, the Figure 6 Reference samples for wide-angle intra-frame prediction are shown.

[0066] The number of modes replaced in the wide-angle directional mode depends on the block aspect ratio. The replaced intra-prediction modes are shown in Table 1.

[0067]

[0068] Figure 7 This illustrates the problem of discontinuities when the orientation exceeds 45°. For example... Figure 7 As shown, in the case of wide-angle intra-frame prediction, two vertically adjacent predicted samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap. The negative impact of pα. If the wide-angle mode represents a non-fractional offset. There are 8 wide-angle modes that satisfy this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted through these modes, the samples in the reference buffer are directly copied without applying any interpolation. This modification reduces the number of samples that need to be smoothed. Furthermore, it aligns the design of non-fractional modes in regular prediction modes with that of the wide-angle mode.

[0069] In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, with the number of entries expanded from 35 to 67 to align with the expansion of intra-prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra-prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​in the mapping table entries to more accurately convert the prediction angles for chroma blocks.

[0070] 2.3. Inter-frame prediction For each inter-frame prediction CU, motion parameters consist of a motion vector, a reference picture index, and a reference picture list usage index, along with additional information required for inter-frame prediction sample generation using new encoding / decoding features of the VVC. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded / decoded motion vector increments, or reference picture indices. A Merge mode is specified, where motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates and additional scheduling introduced in the VVC. The Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where the motion vector for each reference picture list, the corresponding reference picture index, the reference picture list usage flag, and other necessary information are explicitly transmitted via signaling for each CU.

[0071] 2.4. Intra-Block Copying (IBC) Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well known to significantly improve the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer precise. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-encoded CUs are considered a third prediction mode, distinct from intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.

[0072] On the encoder side, hash-based motion estimation for IBC is performed. The encoder performs RD check on blocks with a width or height no greater than 16 luminance samples. For non-Merge mode, block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a local search based on block matching is performed.

[0073] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4x4 sub-blocks. For the larger current block, a hash key match with a reference block is determined when all hash keys of all 4x4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the lowest cost is selected.

[0074] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.

[0075] At the CU level, IBC mode is transmitted via signaling using a flag, and it can be transmitted via signaling as either IBC AMVP mode or IBC skip / Merge mode as follows: IBC Skip / Merge Mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and paired candidates.

[0076] IBC AMVP mode: Block vector difference is encoded and decoded in the same way as motion vector difference. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the upper neighbor (if IBC encoded and decoded). When either neighbor is unavailable, the default block vector is used as the prediction value. A flag is transmitted via signaling to indicate the block vector prediction value index.

[0077] 2.5. Spatial Geometric Partitioning Model (SGPM) SGPM is an intra-frame mode similar to GPM in its inter-frame coding / decoding tools, where two prediction components are generated from the intra-frame prediction process. In this mode, a candidate list is constructed, where each entry contains a segmentation partition and two intra-frame prediction modes, such as... Figure 8 As shown, 26 segmentation modes and 3 intra-frame prediction modes were used to form a combination. The length of the candidate list was set to 16. The selected candidate indices were transmitted via signaling.

[0078] Use templates ( Figure 9 The list is reordered, where the SAD between the template's prediction and reconstruction is used for sorting. The template size is fixed at 1.

[0079] For each segmentation pattern, the same intra-to-inter-frame GPM list derivation is used to derive the IPM list for each segment. The IPM list size is set to 3. In the list, the TIMD derivation pattern is replaced by two derivation patterns with horizontal and vertical directions.

[0080] The SGPM pattern is applied with limited block sizes: 4 <= width <= 64, 4 <= height <= 64, width < height * 8, height < width * 8, width * height >= 32.

[0081] The PPS flag is encoded and decoded to indicate whether non-mixing of two intra-frame predictions is allowed. When the PPS flag is set to false, the following adaptive mixing is also applied to the spatial GPM, where... Figure 10 The mixing depth τ shown is derived as follows: If min(width, height) == 4, then 1 / 2τ is chosen. Otherwise, if min(width, height) == 8, then τ is selected. Otherwise, if min(width, height) == 16, then 2τ is chosen. Otherwise, if min(width, height) == 32, then 4τ is selected. Otherwise, 8τ is selected.

[0082] Otherwise (with the PPS flag set to true), 1 / 4τ is always used for blocks encoded via spatial GPM to ensure that blending is not used when SGPM blocks have perfectly horizontal or vertical split angles, and a much narrower blending width is used when SGPM blocks have other split angles. It should be noted that this flag is set to true in the current Common Test Conditions (CTC) for screen content video.

[0083] 2.6. Multi-Model Learning (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, using a threshold as the average of the luminance reconstruction neighboring samples, the reconstructed neighboring samples are classified into two categories. The linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Slope adjustment is applied to both the cross-component linear model (CCLM) and multi-model LM predictions. This adjustment is a linear function that maps luminance values ​​to chrominance values, tilted relative to a center point determined by the average luminance values ​​of the reference samples.

[0084] 2.7. Extended Merge Forecast In VVC, the Merge candidate list is constructed by including the following five types of candidates in sequence: (1) Airspace MVP from the airspace adjacent to the CU.

[0085] (2) Temporal MVP from the same CU.

[0086] (3) History-based MVP from FIFO table.

[0087] (4) Pair average MVP.

[0088] (5) Zero MV.

[0089] The size of the merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU encoded in Merge mode, the index of the best merge candidate is encoded using rounded unary binarization (TU). The first bit of the merge index is encoded using the context, and bypass encoding is used for the remaining bits.

[0090] This section provides the derivation process for each category of Merge candidates. Similar to HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a region of a specific size.

[0091] 2.7.1. Derivation of Airspace Candidates The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Figure 11 At most four merged candidates are selected from the candidates at the indicated positions. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only if one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or if it is intra-frame encoded / decoded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only... Figure 12 The system uses arrow links to select pairs, and only adds candidates to the list if the corresponding candidates used for redundancy checks do not have the same motion information.

[0092] 2.7.2. Derivation of Time-Domain Candidates In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for the derivation of the co-located CU is explicitly transmitted via signal transmission in the strip header. Figure 13 As shown by the dashed lines, the scaled motion vectors of the temporal merge candidate are obtained by scaling the motion vectors of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located reference image and the co-located image. The reference image index of the temporal merge candidate is set to 0.

[0093] like Figure 14 As shown, the position for the temporal candidate is selected between candidate C0 and C1. If the CU at position C0 is unavailable, intra-frame encoded or decoded, or outside the current line of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate.

[0094] 2.7.3. Derivation of Merge Candidates Based on History Historically based MVP (HMVP) merge candidates are added to the merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-frame encoding / decoding CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0095] The HMVP table size S is set to 6, indicating that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a first-in, first-out (FIFO) rule of constraints is utilized, where a redundancy check is first applied to find if a duplicate HMVP exists in the table. If found, the duplicate HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest few HMVP candidates in the table are checked sequentially and inserted into the candidate list, following the TMVP candidates. Redundancy checks are applied to HMVP candidates for both spatial and temporal merge candidates.

[0096] To reduce the number of redundant check operations, the following simplifications are introduced: Is the number of HMPV candidates used in the Merge list generation set to (N<= 4)? M: (8 (N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table.

[0097] Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the process of building the Merge candidate list from the HMVP is terminated.

[0098] 2.7.4. Derivation of Pairwise Average Merge Candidates Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing Merge candidate list. These predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid.

[0099] If the Merge list is not full after adding pairwise average Merge candidates, a zero MVP will be inserted at the end until the maximum number of Merge candidates is reached.

[0100] 2.7.5. Merge Estimation Region The Merge Estimation Region (MER) allows for the independent derivation of Merge candidate lists for CUs within the same Merge Estimation Region (MER). Candidate blocks located within the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. Furthermore, the update process for the historical motion vector prediction candidate list is only updated when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left brightness sample position of the current CU in the image, and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and is transmitted via signaling as log2_parallel_merge_level_minus2 in the sequence parameter set.

[0101] 2.8. New Merge Candidates 2.8.1. Derivation of Non-Adjacent Merge Candidates In VVC, Figure 15 The five spatial neighbor blocks and one temporal nearest neighbor shown were used to derive the Merge candidate.

[0102] We propose using the same pattern as in VVC to derive additional Merge candidates from positions not adjacent to the current block. To achieve this, for each search round i, the virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated using the following formula: Offsetx =-i×gridX, Offsety = -i×gridY Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.

[0103] Secondly, the width and height of the virtual block are calculated using the following formula: newWidth = i×2×gridX+ currWidthnewHeight = i×2×gridY +currHeight.

[0104] Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.

[0105] gridX and gridY are currently set to currWidth and currHeight, respectively.

[0106] Figure 16 This is a schematic diagram of the virtual blocks in the i-th round of the search. After the virtual blocks are generated, blocks Ai, Bi, Ci, Di, and Ei can be considered as VVC spatial neighbor blocks of the virtual blocks, and their positions are obtained using the same pattern as the pattern in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, blocks Ai, Bi, Ci, Di, and Ei are spatial neighbor blocks used in the VVC Merge pattern.

[0107] When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbor blocks are utilized.

[0108] Non-adjacent spatial merge candidates are inserted into the merge list after the temporal merge candidates in the order B1->A1->C1->D1->E1.

[0109] 2.8.2.STMVP We propose using three spatial merge candidates and one temporal merge candidate to derive the average candidate as the STMVP candidate.

[0110] STMVP is inserted before the Merge candidate in the upper left airspace.

[0111] STMVP candidates were deduplicated along with all previous Merge candidates in the Merge list.

[0112] For airspace candidates, the top three candidates in the current Merge candidate list are used.

[0113] For time-domain candidates, use the same position as the VTM / HEVC co-position.

[0114] For airspace candidates, the first, second, and third candidates inserted into the current Merge candidate list before STMVP are denoted as F, S, and T.

[0115] A time-domain candidate with the same position as the VTM / HEVC co-position used in TMVP is denoted as Col.

[0116] The motion vector (denoted as mvLX) of the STMVP candidate in the prediction direction X is derived as follows: 1) If all four Merge candidates have valid reference indices and are all equal to 0 in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_S + mvLX_T + mvLX_Col)>>2 2) If the reference indices of three of the four Merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F × 3 + mvLX_S × 3 + mvLX_Col × 2)>>3 or mvLX = (mvLX_F × 3 + mvLX_T × 3 + mvLX_Col × 2)>>3 or mvLX = (mvLX_S × 3 + mvLX_T × 3 + mvLX_Col × 2)>>3 3) If the reference indices of two of the four Merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_Col)>>1 or mvLX = (mvLX_S + mvLX_Col)>>1 or mvLX = (mvLX_T + mvLX_Col)>>1 Note: STMVP mode is turned off if time-domain candidates are not available.

[0117] 2.8.3. Merge list size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8.

[0118] 2.9. Geometric Partitioning (GPM) In VVC, geometric segmentation modes are supported for inter-frame prediction. Geometric segmentation modes are transmitted via signaling using a CU-level flag as a merge mode. Other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. For each possible CU size... (in Excluding 8x64 and 64x8, the geometric segmentation mode supports a total of 64 segments.

[0119] Figure 17 An example of GPM partitioning grouped at the same angle is shown. When using this mode, the CU is divided into two parts by geometrically positioned straight lines ( Figure 17 The position of the dividing line is mathematically derived from the angle and offset parameters of a specific segment. Each part of the geometric segment in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each segment, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with conventional bidirectional prediction, only two motion-compensated predictions are needed per CU. The unidirectional prediction motion for each segment is derived using the process described in 2.9.1.

[0120] If a geometric segmentation mode is used for the current CU, the geometric segmentation mode (angle and offset) and two merge indices (one for each segment) are further indicated via signal transmission. The number of maximum GPM candidate dimensions is explicitly transmitted in SPS, and the syntax binarization used for the GPM merge index is specified. After predicting each part of the geometric segmentation, a mixing process with adaptive weights is used, and the sample values ​​along the geometric segmentation edges are adjusted, as shown in 2.20.2. This is the prediction signal for the entire CU, and the transformation and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric segmentation mode is stored.

[0121] 2.9.1. Construction of One-Way Prediction Candidate List The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 2.7. Let n denote the index of the unidirectional predicted motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional predicted motion vector for the geometric segmentation pattern. These motion vectors in... Figure 18 The value is marked with "x". If the corresponding LX motion vector of the nth extended Merge candidate does not exist, then the L(1) of the same candidate... The X motion vector is used as a unidirectional predictive motion vector for the geometric segmentation pattern.

[0122] 2.9.2. Blending along geometrically segmented edges After predicting each part of the geometric segment using its own motion, a blend is applied to the two predicted signals to derive samples around the segmentation edges. The blending weights at each location of the CU are derived based on the distance between the individual location and the segmentation edge.

[0123] Location The distance to the segmentation edge is derived as follows: (2-1) (2-2) (2-3) (2-4) in It is an index for the angle and offset of the geometric segmentation, which depends on the geometric segmentation index emitted by the signal. The sign depends on the angle index. .

[0124] The weights of each part of the geometric segment are derived as follows: (2-5) (2-6) (2-7) partIdx depends on the angle index Weight An example in Figure 19 It is shown in the middle. Figure 19 The mixed weights using geometric segmentation patterns are shown. An example of generation.

[0125] 2.9.3. Motion field storage for geometric segmentation patterns Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded by the geometric segmentation pattern.

[0126] The type of motion vector stored for each individual location in the sports field is determined as follows: (2-8) Where motionIdx equals It is recalculated from equation (2-18). partIdx depends on the angle index. .

[0127] If sType equals 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field; otherwise, if sType equals 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.

[0128] Otherwise, if Mv1 and Mv2 come from the same list, only the unidirectional predicted motion Mv2 is stored.

[0129] 2.10. Non-adjacent airspace candidates Figure 20 The diagram shows the spatial neighbor blocks used to derive spatial merge candidates. Non-adjacent spatial merge candidates are inserted after the TMVP in the regular merge candidate list. The style of the spatial merge candidates is shown in... Figure 20 The distance between non-adjacent spatial domain candidates and the current codec block is shown in the diagram. The distance between the current codec block and the non-adjacent spatial domain candidate is based on the width and height of the current codec block.

[0130] 2.11. Template Matching (TM) Template matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference image (i.e., of the same size as the template). For example... Figure 21 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU. The template matching previously proposed in JVET-J0021 is adopted with two modifications: the search step size is determined based on the AMVR mode, and in the Merge mode, the TM can be cascaded with the bilateral matching process.

[0131] In AMVP mode, MVP candidates are determined based on template matching error, selecting the one that minimizes the difference between the current block template and the reference block template. TM then performs MV refinement only on that specific MVP candidate. TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. AMVP candidates can be further refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially according to the AMVR mode specified in Table 2. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the AMVR mode after the TM process.

[0132] Table 2 – Search patterns for AMVR and search patterns using AMVR's Merge mode

[0133] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 2, TM can proceed up to 1 / 8 pixel MVD precision, or skip those precisions beyond half-pixel MVD precision, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check.

[0134] 2.12. Multiple Transform Selection (MTS) for Core Transformation In addition to DCT-II, which is already used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding and decoding of both inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 3 shows the basis functions of the selected DST / DCT.

[0135]

[0136] To maintain the orthogonality of the transformation matrices, the transformation matrices are quantized more precisely than those in HEVC. To keep the intermediate values ​​of the transformation coefficients within the 16-bit range, all coefficients must have 10 bits after both the horizontal and vertical transformations.

[0137] To control the MTS scheme, separate enable flags are specified at the SPS level for intra-frame and inter-frame operations. When MTS is enabled at SPS, CU-level flags are signaled to indicate whether MTS is applied. Here, MTS is applied only to luminance. MTS signaling is skipped when one of the following conditions is met.

[0138] The position of the last effective coefficient of luminance TB is less than 1 (i.e., DC only).

[0139] The last effective coefficient of luminance TB is located in the MTS zero region.

[0140] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are transmitted via signaling to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 4. A unified transform selection for ISP and implicit MTS is used by removing intra-frame mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra-frame block and both intra-frame explicit MTS and inter-frame explicit MTS are enabled, only DST7 is used for both the horizontal and vertical transform cores. An 8-bit master transform core is used when transform matrix precision is involved. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8, use an 8-bit master transform core.

[0141]

[0142] To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are zeroed for DST-7 and DCT-8 blocks with dimensions (width or height, or both) equal to 32. Only the coefficients in the 16x16 low-frequency region are retained.

[0143] In HEVC, for example, block residuals can be encoded and decoded using transform skip mode. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Implicit MTS can also be enabled when MTS is enabled for inter-frame encoding / decoding blocks.

[0144] 2.13. Subblock Transformation (SBT) In VTM, a sub-block transform is introduced for inter-frame prediction CUs. In this transform mode, only a sub-part of the residual block is encoded / decoded for the CU. When the inter-frame prediction CU has cu_cbf equal to 1, the signal cu_sbt_flag can be used to indicate whether the entire residual block or a sub-part of the residual block is encoded / decoded. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is encoded / decoded using a presumed adaptive transform, and another portion of the residual block is zeroed out.

[0145] When SBTs are used in inter-frame encoding / decoding CUs, SBT type and SBT position information are transmitted via signals in the bitstream. For example... Figure 22 As shown, there are two SBT types and two SBT locations. For SBT-V (or SBT-H), the TU width (or height) can be equal to half the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. A 2:2 partition is like a binary tree (BT) partition, while a 1:3 / 3:1 partition is like an asymmetric binary tree (ABT) partition. In an ABT partition, only small regions include non-zero residuals. If one dimension of the CU is 8 in the luma samples, a 1:3 / 3:1 partition along that dimension is not allowed. A CU can have a maximum of 8 SBT modes.

[0146] Position-dependent transform core selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TB always uses DCT-2). Two positions in SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal and vertical transforms at each SBT position are... Figure 22 The transformations are specified in the code. For example, the horizontal and vertical transformations at SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transformations in both dimensions are set to DCT-2. Therefore, the sub-block transformations jointly specify the TU slice, cbf, and the horizontal and vertical core transformation types of the residual block. Figure 22 The location, type, and transform type of the SBT are shown. The SBT is not applied to CUs encoded and decoded using intra-frame / inter-frame joint mode.

[0147] 2.14. Adaptive Merge Candidate Reordering Based on Template Matching To improve encoding and decoding efficiency, after constructing the merge candidate list, the order of each merge candidate is adjusted based on the template matching cost. The merge candidates are arranged in the list according to their ascending template matching cost. This is done in subgroups.

[0148] Template matching cost is measured by the sum of absolute differences (SAD) between the current CU's neighboring samples and their corresponding reference samples. If the merge candidate includes bidirectional predicted motion information, then... Figure 23 As shown, the corresponding reference sample is the average of the corresponding reference sample in reference list 0 and the corresponding reference sample in reference list 1. If the Merge candidate includes motion information at the sub-CU level, then as follows... Figure 24 As shown, the corresponding reference sample point is composed of the neighboring sample points of the corresponding reference sub-block.

[0149] like Figure 25 As shown, the sorting process is performed in subgroups. The first three merge candidates are sorted together. The last three merge candidates are sorted together.

[0150] The template size (width of the left template or height of the top template) is 1. The subgroup size is 3.

[0151] We can assume there are 8 merge candidates. We will take the first 5 merge candidates as the first subgroup and the last 3 merge candidates as the second subgroup (i.e., the last subgroup).

[0152] For the encoder, after constructing the Merge candidate list, as follows Figure 26 As shown, some Merge candidates are adaptively reordered in ascending order of Merge candidate cost.

[0153] More specifically, the template matching cost is calculated for the merge candidates in all subgroups except the last one; then the merge candidates in their own subgroups except the last one are reordered; finally, the final merge candidate list is obtained.

[0154] For the decoder, after constructing the Merge candidate list, such as Figure 27 As shown, some / no merge candidates are adaptively reordered in ascending order at the merge candidate cost. Figure 27 In this context, the subgroup containing the selected (transmitted via signal) Merge candidate is referred to as the selected subgroup.

[0155] More specifically, if the selected Merge candidate is in the last subgroup, the Merge candidate list construction process is terminated after deriving the selected Merge candidate, no reordering is performed, and the Merge candidate list remains unchanged; otherwise, the process is as follows: After deriving all Merge candidates in the selected subgroup, the Merge candidate list construction process is terminated; the template matching cost for the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained.

[0156] For both the encoder and decoder: the template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template.

[0157] When deriving reference samples for the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.

[0158] The reference samples (RT) for the template used for bidirectional prediction are obtained by using the reference samples of the template in reference list 0 as follows ( ) and reference samples of the template in reference list 1 ( It is derived by weighted averaging.

[0159] (2-9) The weights (8-w) of the reference templates in reference list 0 and the weights (w) of the reference templates in reference list 1 are determined by the BCW indices of the Merge candidates. The BCW indices equal to {0,1,2,3,4} correspond to w equal to {-2,3,4,5,10}, respectively.

[0160] If the Local Illumination Compensation (LIC) flag of the Merge candidate is true, the reference sample points of the template are derived using the LIC method.

[0161] Template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT.

[0162] The template size is 1. This means that the width of the left template and / or the height of the top template is 1.

[0163] If the encoding / decoding mode is MMVD, the Merge candidates used to derive the basic Merge candidates are not reordered. If the encoding / decoding mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered.

[0164] 2.15. Geometric segmentation mode with motion vector difference In the Geometry Segmentation Mode with Motion Vector Difference (GMVD), each geometric segment in the GPM can determine whether GMVD is used. If GMVD is selected for a geometric region, the region's MV is calculated as the sum of the MV of the merged candidates and the MVD. All other processing remains the same as in the GPM.

[0165] Using GMVD, MVD is transmitted as a pair of direction and distance via signals. Nine candidate distances are involved (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD in GMVD is shifted left by 2, just like in MMVD.

[0166] 2.16. Geometric segmentation pattern with affine prediction (GPM-affine) GPM is further extended to enable affine motion compensation (AMC). Therefore, GPM segmentation can be predicted via AMC inter-frame prediction, non-AMC inter-frame prediction, or intra-frame prediction. Furthermore, a GPM segmentation predicted via AMC can be combined with another GPM segmentation predicted via AMC, non-AMC, or intra-frame prediction.

[0167] When AMC is applied, similar to the construction of the one-way predictive merge candidate list for GPM in VVC, the one-way predictive affine merge candidate list is constructed from the sub-block-based merge candidate list after discarding sub-TMVP candidates. AMC is performed for GPM segmentation using the control point motion vectors (CPMVs) of the merge candidates in the one-way predictive affine merge candidate list.

[0168] For each GPM segment, a gpm_affine_flag is signaled to indicate whether AMC is applied to the GPM segment. Depending on whether AMC or non-AMC is applied, different arithmetic context models are used to signal merge candidate indices for the GPM segment.

[0169] In the current implementation, AMC is not allowed for GPM-MMVD and GPM-TM.

[0170] 2.17. Regression-based GPM mixture An additional implicit GPM mode is proposed, in which two integer mixing matrices (W0 and W1) are derived from the template (top row, left column). The mixing matrix is ​​modeled as an affine linear function of the sample location (x, y) in the current CU: W 0 (x,y) = ax + by + c and W 1 (x,y) = 1 - W 0 (x,y) (2-10) The parameters (a, b, c) are derived from the reference template using the same solver (MSE minimization) as used for CCCM, GLM, or GL-CCCM. A list of candidates is constructed from the regular GPM candidates and reordered using the template cost.

[0171] The GPM implicit mode is signaled via a CU-level flag (gpm_implicit_flag). If gpm_implicit_flag is true, merge-idx is encoded / decoded to signal a pair of GPM candidates to be used. If gpm_implicit_flag is false, regular GPM syntax elements are signaled.

[0172] 3. Problem In the current design of SGPM, a candidate list is constructed, where each entry contains a segmentation and two intra-prediction modes. The construction of the candidate list involves information from the current block from several aspects. First, template matching costs are used to reorder the candidate list. Second, the three intra-prediction modes used to construct the candidate list are derived using neighbor samples. However, historical and / or neighbor information specific to SGPM has not been investigated, as it could contribute to improving the encoding and decoding performance of SGPM.

[0173] 4. Detailed Solution The detailed solutions below should be considered as examples for explaining general concepts. These solutions should not be interpreted in a narrow sense. Furthermore, these solutions can be combined in any way.

[0174] In this disclosure, the SGPM candidate includes one segmentation mode and two intra-frame prediction modes.

[0175] Spatial Geometric Prediction Model with Merge Mode (SGPM Merge Model) 1. One or more new SGPM candidates are proposed that can be used to obtain prediction / reconstruction of video cells, wherein one or more SGPM candidates are different from the SGPM candidates used in video cells using SGPM encoding and decoding as disclosed in ECM-10.1.

[0176] a. In one example, a new SGPM candidate can be derived from spatially adjacent (adjacent and / or non-adjacent) video cells.

[0177] b. In one example, a new SGPM candidate can be derived from a temporal video unit.

[0178] c. In one example, new SGPM candidates can be derived from the history table / list used to store SGPM candidates.

[0179] d. In one example, a new SGPM candidate can be constructed.

[0180] i. In one example, the encoding and decoding information of neighboring video units can be used to construct new SGPM candidates.

[0181] 1) In one example, the intra-frame prediction mode of neighboring video units can be used.

[0182] e. In one example, one or more new SGPM candidates can be used to construct the candidate list for the current video unit.

[0183] f. In one example, at least one SGPM segmentation can be predicted using new SGPM candidates.

[0184] 2. In one example, one or more SGPM candidates can be used to construct one or more new candidate lists, and the prediction / reconstruction of video units is obtained using the candidate lists. The codec mode is represented as the SGPM Merge mode.

[0185] a. In one example, an SGPM candidate can refer to a triplet of a segmentation pattern and two intra-prediction patterns.

[0186] i. Alternative sites: SGPM candidates can refer to one or more partitioning patterns.

[0187] ii. Alternatively, SGPM candidates may refer to one or more intra-frame prediction modes.

[0188] iii. Alternative sites: SGPM candidates may refer to one or more prediction samples.

[0189] b. In one example, one or more SGPM candidates may come from video units encoded and decoded using SGPM and / or SGPM Merge mode or other SGPM modes.

[0190] c. In one example, reordering can be used to create a new candidate list.

[0191] i. In one example, reordering could depend on template matching or template matching cost.

[0192] ii. In one example, the reordering can be the same as that used in SGPM.

[0193] d. In one example, which candidate list is used can be signaled, deduced, or predefined.

[0194] e. In one example, which SGPM candidate from the candidate list is used in the SGPM Merge mode to obtain prediction / reconstruction samples can be transmitted via signaling.

[0195] i. In one example, one or more syntax elements can be used to indicate SGPM candidates.

[0196] f. In one example, which SGPM candidate from the SGPM candidate list is used in the SGPM Merge mode to obtain prediction / reconstruction samples can be predefined or derived.

[0197] g. In one example, different SGPM splits can use different SGPM Merge candidate lists.

[0198] 3. In one example, one of the SGPM candidates or elements in the SGPM candidates for a video unit that was encoded / decoded before the current video unit was encoded / decoded can be used for the current video unit.

[0199] a. In one example, video units that were encoded / decoded before the current video unit can be in different strips / slices / subpictures / pictures / CTUs / CTU rows.

[0200] b. In one example, video units that were encoded / decoded before the current video unit can be in the same strip / slice / sub-picture / picture / CTU / CTU row as the current video unit.

[0201] c. In one example, the video unit that was encoded / decoded before the current video unit can be a neighboring spatial domain (e.g., adjacent and / or non-adjacent) video unit.

[0202] d. In one example, reused SGPM candidates can be stored in a list / table (e.g., a historical SGPM candidate table).

[0203] i. In one example, the list / table can be updated during the encoding / decoding process.

[0204] ii. In one example, the maximum size of the list / table can be predefined, transmitted via signal, or derived.

[0205] iii. In one example, the list / table can be reinitialized at the beginning of the strip / slice / sub-image / picture / CTU / CTU.

[0206] 1) In one example, the list / table can be reinitialized to an empty list / table.

[0207] 2) In one example, the list / table can be reinitialized using one or more predefined / derived / signaled SGPM candidates.

[0208] iv. In one example, how and / or whether to use / update the list / table can depend on the encoding / decoding information.

[0209] 1) In one example, encoding / decoding information can refer to block dimension / size / location.

[0210] e. In one example, when the current video unit is a chroma video unit, the reused SGPM candidate can come from the luma video unit and / or the chroma video unit.

[0211] 4. Whether and / or how to apply the SGPM Merge mode can depend on the encoding / decoding information, which can refer to: a. Whether specific encoding / decoding methods, such as SGPM, are allowed. b. Block dimensions and / or block size c. Block depth d. Strip / image type and / or segmentation tree type (single tree, dual tree, or local dual tree) e. Temporal layer identifier f. Block position g. CTU / strip / film / image / image resolution h. Color format i. Color components i. In one example, SGPM and / or SGPM Merge modes can be applied to all color components.

[0212] ii. In one example, when the SGPM and / or SGPM Merge mode is applied to the chromaticity component, it may be different from the SGPM and / or SGPM Merge mode applied to the luminance component.

[0213] iii. In one example, whether and / or how the SGPM and / or SGPM Merge pattern is applied to the first component may depend on whether and / or how the SGPM and / or SGPM Merge pattern is applied to the second component.

[0214] 1) In one example, the first component may refer to the chromaticity component (e.g., Cb and / or Cr), and the second component may refer to the luminance component (e.g., Y).

[0215] 2) In one example, the SGPM and / or SGPM Merge patterns can be applied to the first component in the same way as the second component.

[0216] a) Alternatively, the way the SGPM and / or SGPM Merge pattern is applied to the first component may differ from that of the second component.

[0217] iv. In one example, the SGPM Merge mode can be applied to the luminance component but not to the chrominance component.

[0218] 1) In one example, the luminance component can refer to Y in the YCbCr color space or G in the RGB color space.

[0219] 2) In one example, the chromaticity components can refer to Cb and / or Cr in the YCbCr color space or R and / or B in the RGB color space.

[0220] 5. The indication of SGPM Merge mode can be conditionally transmitted via signaling, wherein the conditions may include: a. Block dimensions and / or block size b. Block depth c. Strip / image type and / or segmentation tree type (single tree, dual tree, or local dual tree) d. Temporal layer identifier e. Block location f. CTU / strip / film / image / image resolution g. Color format h. Color components 6. Whether the current block is encoded or decoded in SGPM Merge mode can be transmitted via signals using one or more syntax elements (SE).

[0221] a. In one example, syntax elements may be binarized using fixed-length encoding, rounding unary encoding, unary encoding, or EG encoding, or encoded as flags.

[0222] b. In one example, syntax elements can be either bypassed or context-encoded.

[0223] i. The context may depend on encoded / decoded information, such as block dimensions, and / or block size, and / or stripe / picture type, and / or information about neighboring blocks (adjacent or non-adjacent), and / or information about other encoding / decoding tools used for the current block, and / or information about the temporal layer.

[0224] c. In one example, one or more syntax elements may be transmitted via signaling at the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0225] d. In one example, syntax elements can be encoded and decoded in a predictive manner.

[0226] e. In one example, syntax elements can be conditionally encoded or decoded.

[0227] i. For example, a second SE may be signaled to indicate whether SGPM Merge mode is used only when the first SE indicates that SGPM Merge mode is applicable.

[0228] 1) The first SE can be located at the sequence header / image header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0229] 2) The second SE can be block-specific.

[0230] ii. For example, a third SE may be signaled to indicate how to perform the SGPM Merge mode only when the second SE indicates the use of the SGPM Merge mode.

[0231] 1) In one example, the third SE can be used to indicate which SGPM candidate is used.

[0232] 7. It is proposed that the mixing matrix used in SGPM and / or IBC-GPM can be derived using the same method as in regression-based GPM.

[0233] a. In one example, whether and / or how to use the mixing matrix can be predefined or derived.

[0234] b. In one example, whether and / or how a hybrid matrix can be transmitted via signaling.

[0235] i. In a single example, one or more syntax elements may be used.

[0236] c. In one example, whether and / or how to use a blending matrix may differ for different video content (e.g., content captured by a camera or screen content).

[0237] General aspects 9. In the above examples, a video unit can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region including more than one sample or pixel.

[0238] 10. Whether and / or how the methods disclosed above can be applied to be transmitted via signal at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0239] 11. Whether and / or how to apply the above methods may depend on the following information: a. Messages transmitted via signals in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Piece Group Header / Codec Tree Unit (CTU) / Codec Unit (CU) / CTU Line / CTU Group / TU / PU Block / Video Codec Unit; b. Location of CU / PU / TU / block / video codec unit; c. The block dimensions of the current block and / or its neighboring blocks; d. The block shape of the current block and / or its neighboring blocks; e. The encoding / decoding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode; f. Indication of color format (such as 4:2:0, 4:4:4); g. Encoding / decoding tree structure; h. Strip / panel type and / or image type; i. Color components (e.g., can be applied only to the chromaticity component or the luminance component); j. Temporal layer ID; k. Standard grade / level / tier.

[0240] Figure 28 A flowchart of a method 2800 for video processing according to an embodiment of the present disclosure is shown. Method 2800 is implemented during the conversion between video units of a video and a bitstream of a video.

[0241] At box 2810, the conversion between video units and video bitstreams for a video unit is performed, based on at least one of the following: historical information or proximity information, to obtain one or more Spatial Geometric Partitioning Mode (SGPM) candidates for the video unit. In some embodiments, a video unit includes at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, region containing more than one sample point or pixel.

[0242] At box 2820, a prediction or reconstruction of a video cell is obtained based on one or more SGPM candidates. For example, one or more new SGPM candidates may be used to obtain the prediction / reconstruction of a video cell, wherein one or more SGPM candidates are different from the SGPM candidates used in the video cell encoded and decoded in SGPM.

[0243] At box 2830, a transformation is performed based on the prediction or construction of video units. In some embodiments, the transformation includes encoding video units into a bitstream. In some embodiments, the transformation includes decoding video units from the bitstream. In this way, the efficiency and performance of SGPM can be improved by taking into account historical and proximity information.

[0244] In some embodiments, one or more SGPM candidates are derived from at least one of spatially adjacent neighboring video units or spatially non-adjacent neighboring video units. In some other embodiments, one or more SGPM candidates are derived from temporal video units. In some still embodiments, one or more SGPM candidates are derived from a history table or list used to store SGPM candidates.

[0245] In some embodiments, one or more SGPM candidates are constructed. For example, one or more SGPM candidates are constructed using codec information of at least one neighboring video unit. In some embodiments, the intra-prediction mode of at least one neighboring video unit is used.

[0246] In some embodiments, one or more SGPM candidates are used to construct a candidate list for video units. In some other embodiments, at least one SGPM segmentation utilizes one of the one or more SGPM candidates for prediction.

[0247] In some embodiments, one of the SGPM candidates or elements in the SGPM candidates for a video unit encoded or decoded before the current video unit is encoded or decoded, or a video unit decoded before the current video unit is decoded, is used for the current video unit. In some embodiments, the video unit encoded or decoded before the current video unit is in one of a different strip, a different slice, a different subpicture, a different picture, a different codec tree unit (CTU), or a different CTU row. In some other embodiments, the video unit encoded or decoded before the current video unit is in one of the same strip, the same slice, the same subpicture, the same picture, the same CTU, or the same CTU row as the current video unit.

[0248] In some embodiments, the video unit that was encoded or decoded before the current video unit is a spatially adjacent video unit. For example, a spatially adjacent video unit may be adjacent to or not adjacent to the current video unit.

[0249] In some embodiments, one or more reused SGPM candidates are stored in a list or table. For example, the table is a historical SGPM candidate table. In some embodiments, the list or table is updated during the encoding / decoding or decoding process. In some other embodiments, the maximum size of the list or table is predefined, transmitted via signaling, or derived.

[0250] In some embodiments, the list or table is reinitialized at the beginning of one of the following: strip, slice, sub-picture, picture, CTU, CTU row. In some embodiments, the list is reinitialized to an empty list. Alternatively, the table is reinitialized to an empty table. In some embodiments, the list or table is reinitialized using one or more predefined, deduced, or signaled SGPM candidates.

[0251] In some embodiments, the manner and / or whether to use or update the list or table depends on the codec information. For example, the codec information includes at least one of the following: block dimension, block size, or block location. In some embodiments, when the current video unit is a chroma video unit, the reused SGPM candidates come from at least one of the luma video unit or the chroma video unit.

[0252] In some embodiments, the blending matrix used in at least one of SGPM or Intra-Block Copy Geometric Partitioning Mode (IBC-GPM) is derived using the same method as in regression-based GPM. In some embodiments, whether to use a blending matrix and / or the manner in which the blending matrix is ​​used is predefined or derived. In some other embodiments, whether to use a blending matrix and / or the manner in which the blending matrix is ​​used is determined by signal transmission.

[0253] In some embodiments, one or more syntax elements are used to indicate whether and / or how a blending matrix is ​​used. In some embodiments, the use of a blending matrix for a first type of video content and / or the manner in which a blending matrix is ​​used for the first type of video content differs from the use of a blending matrix for a second type of video content. For example, the first type of video content is content captured by a camera, and the second type of video content is screen content.

[0254] In some embodiments, an indication of whether one or more SGPM candidates are obtained for a video unit and / or how one or more SGPM candidates are obtained for a video unit is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, an indication of whether one or more SGPM candidates are obtained for a video unit and / or how one or more SGPM candidates are obtained for a video unit is indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.

[0255] In some embodiments, whether and / or how one or more SGPM candidates for a video unit are obtained depends on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, strip header, slice group header, codec tree unit (CTU), codec unit (CU), CTU row, CTU group, TU, PU block, video codec unit, position of CU block, position of PU block, position of TU block, position of video codec unit, block dimension of current block, block dimension of neighboring blocks of current block, block shape of current block, block shape of neighboring blocks of current block, codec mode of block, indication of color format, codec tree structure, strip group type, slice group type, strip picture type, slice picture type, color components, ID of temporal layer, standard grade, standard level, or standard layer.

[0256] In some embodiments, the block encoding / decoding mode is at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, or non-IBC sub-block mode. In some embodiments, the color format is indicated as 4:2:0 or 4:4:4. Alternatively, color components are applied to one of the following: chroma component or luminance component.

[0257] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: obtaining one or more spatial geometric segmentation pattern (SGPM) candidates for video units of the video based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; and generating a bitstream based on the prediction or construction of the video units.

[0258] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: obtaining one or more Spatial Geometric Partition Pattern (SGPM) candidates for video units of the video based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; generating a bitstream based on the prediction or construction of the video units; and storing the bitstream in a non-transitory computer-readable recording medium.

[0259] Figure 29 A flowchart of a method 2900 for video processing according to an embodiment of the present disclosure is shown. Method 2900 is implemented during the conversion between video units of a video and a bitstream of a video.

[0260] At box 2910, for the conversion between video units and the video bitstream, a candidate list for video units is constructed based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units. Video units are encoded and decoded in SGPM Merge mode. In some embodiments, a video unit includes at least one of the following: color components, sub-pictures, stripes, slices, codec tree units (CTUs), CTU rows, CTU groups, codec units (CUs), prediction units (PUs), transform units (TUs), codec tree blocks (CTBs), codec blocks (CBs), prediction blocks (PBs), transform blocks (TBs), blocks, sub-blocks of blocks, sub-regions within blocks, and regions containing more than one sample or pixel.

[0261] At box 2920, the prediction or reconstruction of a video cell is obtained based on a candidate list. In one example, one or more SGPM candidates can be used to construct one or more new candidate lists, and the prediction / reconstruction of the video cell is obtained using the candidate lists. For example, SGPM candidates of neighboring video cells and / or previously encoded / decoded video cells can be used to generate new candidate lists. For example, a combination of orientation and intra-frame mode is used to construct the SGPM candidate list, and the index of the combination is transmitted via signal transmission. The combination of orientation and intra-frame mode can be obtained directly from the codec block.

[0262] At box 2930, a transformation is performed based on the prediction or reconstruction of video units. In some embodiments, the transformation includes encoding video units into a bitstream. In some embodiments, the transformation includes decoding video units from the bitstream. In this way, by taking into account historical and proximity information, the efficiency and performance of SGPM can be improved.

[0263] In some embodiments, one or more SGPM candidates comprise a triplet of a segmentation mode and two intra-prediction modes. In some other embodiments, one or more SGPM candidates comprise one or more segmentation modes. Alternatively, one or more SGPM candidates comprise one or more intra-prediction modes. In some further embodiments, one or more SGPM candidates comprise one or more prediction samples. In some embodiments, one or more SGPM candidates originate from video units encoded using at least one of the following: SGPM, SGPM Merge mode, or other SGPM modes.

[0264] In some embodiments, reordering is applied to the candidate list. For example, reordering may depend on template matching or template matching cost. In some embodiments, the reordering is the same as that used in SGPM.

[0265] In some embodiments, which candidate list is used is transmitted, derived, or predefined. In some other embodiments, which SGPM candidate in the candidate list is used in the SGPM Merge mode to obtain prediction or reconstruction samples is transmitted via signaling. For example, one or more syntax elements are used to indicate one or more SGPM candidates. In some embodiments, which SGPM candidate in the candidate list is used in the SGPM Merge mode to obtain prediction or reconstruction samples is predefined or derived.

[0266] In some embodiments, the first SGPM segmentation uses a first SGPM Merge candidate list, and the second SGPM segmentation uses a second SGPM Merge candidate list. In this case, the first SGPM segmentation may be different from the second SGPM segmentation.

[0267] In some embodiments, whether a video unit is encoded or decoded in SGPM Merge mode is transmitted via signaling using one or more syntax elements (SE). In some embodiments, one or more syntax elements are binarized using one of the following: fixed-length encoding / decoding, rounding unary encoding / decoding, unary encoding / decoding, or EG encoding / decoding, or one or more syntax elements are encoded as flags.

[0268] In some embodiments, one or more syntax elements are either bypassed or context-encoded. In some embodiments, the context depends on the encoded information. For example, the encoded information includes at least one of the following: block dimension, block size, stripe type, picture type, information about neighboring blocks, information about other codecs used for the video unit, or information about the temporal layer. In some embodiments, one or more syntax elements are signaled at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice header.

[0269] In some embodiments, one or more syntax elements are encoded / decoded predictively. In some other embodiments, one or more syntax elements are encoded / decoded conditionally. For example, if a first SE indicates that SGPM Merge mode is applicable, a second SE is signaled to indicate whether SGPM Merge mode is used. In some embodiments, the first SE is at one of the following: sequence header, picture header, SPS, VPS, DPS, DCI, PPS, APS, strip header, or slice header. In some embodiments, the second SE is for blocks.

[0270] In some embodiments, if the second SE indicates the use of the SGPM Merge mode, the third SE is signaled to indicate how the SGPM Merge mode should be performed. In some embodiments, the third SE is used to indicate which SGPM candidate is used.

[0271] In some embodiments, whether and / or how the SGPM Merge mode is applied depends on the encoding / decoding information. In some embodiments, the encoding / decoding information includes at least one of the following: whether the encoding / decoding method is permitted, block dimension, block size, block depth, stripe type, picture type, segmentation tree type, temporal layer identifier, block location, CTU resolution, stripe resolution, slice resolution, subpicture resolution, picture resolution, color format, or color components.

[0272] In some embodiments, the encoding / decoding method is SGPM. Alternatively or additionally, the segmentation tree type includes one of the following: single tree, dual tree, or partial dual tree.

[0273] In some embodiments, at least one of the SGPM or SGPM Merge modes is applied to all color components. In some other embodiments, if at least one of the SGPM or SGPM Merge modes is applied to the chroma component, it differs from at least one of the SGPM or SGPM Merge modes applied to the luminance component.

[0274] In some embodiments, whether at least one of the SGPM or SGPM Merge modes is applied to the first component depends on whether at least one of the SGPM or SGPM Merge modes is applied to the second component. Alternatively or additionally, the manner in which at least one of the SGPM or SGPM Merge modes is applied to the first component depends on the manner in which at least one of the SGPM or SGPM Merge modes is applied to the second component. In some embodiments, the first component is a chromaticity component, and the second component is a luminance component.

[0275] In some embodiments, at least one of the SGPM or SGPM Merge patterns is applied to the first component in the same way as the second component. Alternatively, at least one of the SGPM or SGPM Merge patterns is applied to the first component in a different way than the second component.

[0276] In some embodiments, the SGPM Merge mode is applied to the luminance component but not to the chrominance component. In some embodiments, the luminance component is Y in the YCbCr color space or G in the RGB color space. In some other embodiments, the chrominance component is Cb and / or Cr in the YCbCr color space or R and / or B in the RGB color space.

[0277] In some embodiments, the indication of the SGPM Merge mode is transmitted via signaling based on conditions. For example, the conditions include at least one of the following: block dimension, block size, block depth, strip type, picture type, segmentation tree type, temporal layer identifier, block location, CTU resolution, strip resolution, slice resolution, sub-picture resolution, picture resolution, color format, or color components.

[0278] In some embodiments, the indication of whether to construct a candidate list for video units and / or how to construct the candidate list for video units is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level. In some embodiments, the indication of whether to construct a candidate list for video units and / or how to construct the candidate list for video units is indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.

[0279] In some embodiments, whether and / or how a candidate list for video units is constructed depends on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, stripe header, slice group header, codec tree unit (CTU), codec unit (CU), CTU row, CTU group, TU, PU block, video codec unit, position of CU block, position of PU block, position of TU block, position of video codec unit, block dimension of the current block, block dimension of the current block's neighboring blocks, block shape of the current block, block shape of the current block's neighboring blocks, codec mode of the block, color format indication, codec tree structure, stripe group type, slice group type, stripe picture type, slice picture type, color components, temporal layer ID, standard grade, standard level, or standard layer. In some embodiments, the codec mode of the block is at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, or non-IBC sub-block mode. In some embodiments, the color format indication is 4:2:0 or 4:4:4. In some other embodiments, the color component is applied to either the chromaticity component or the luminance component.

[0280] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by means of a video processing apparatus. The method includes: constructing a candidate list of video units for the video based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in an SGPM Merge mode; obtaining a prediction or reconstruction of the video units based on the candidate list; and generating a bitstream based on the prediction or reconstruction of the video units.

[0281] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: constructing a candidate list of video units for the video based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in an SGPM Merge mode; obtaining a prediction or reconstruction of the video units based on the candidate list; generating a bitstream based on the prediction or reconstruction of the video units; and storing the bitstream in a non-transitory computer-readable recording medium.

[0282] The implementation of this disclosure can be described by the following entries, the features of which can be combined in any reasonable manner.

[0283] Item 1. A video processing method comprising: a conversion between video units and a bitstream of video; obtaining one or more spatial geometric partitioning pattern (SGPM) candidates for the video units based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; and performing the conversion based on the prediction or construction of the video units.

[0284] Item 2. According to the method of Item 1, one or more SGPM candidates are derived from at least one of spatially adjacent neighboring video units or spatially non-adjacent neighboring video units.

[0285] Item 3. According to the method of Item 1, one or more SGPM candidates are derived from temporal video units.

[0286] Item 4. According to the method of Item 1, one or more SGPM candidates are derived from a history table or list used to store SGPM candidates.

[0287] Item 5. One or more SGPM candidates are constructed according to the method in Item 1.

[0288] Item 6. According to the method of Item 5, one or more SGPM candidates are constructed using the encoding and decoding information of at least one neighboring video unit.

[0289] Item 7. According to the method of Item 6, an intra-frame prediction mode of at least one neighboring video unit is used.

[0290] Item 8. In accordance with the method of Item 1, one or more SGPM candidates are used to construct a candidate list of video units.

[0291] Item 9. According to the method of Item 1, at least one SGPM segmentation is predicted using one of the one or more SGPM candidates.

[0292] Item 10. According to the method of Item 1, wherein one of the SGPM candidates or elements in the SGPM candidates for a video unit that was encoded or decoded before the current video unit was encoded or decoded, or a video unit that was decoded before the current video unit was decoded, is used for the current video unit.

[0293] Item 11. According to the method of Item 10, wherein the video units that were encoded or decoded before the current video unit are in one of the following: different strips, different slices, different subpictures, different pictures, different codec tree units (CTUs), or different CTU rows.

[0294] Item 12. According to the method of Item 10, the video unit that was encoded or decoded before the current video unit is in one of the same strip, the same slice, the same sub-picture, the same picture, the same CTU, or the same CTU row as the current video unit.

[0295] Item 13. According to the method of Item 10, the video unit that was encoded or decoded before the current video unit is a spatially adjacent video unit.

[0296] Item 14. According to the method of Item 13, the spatially adjacent video unit is either adjacent to or not adjacent to the current video unit.

[0297] Item 15. According to the method of Item 10, one or more of the reused SGPM candidates are stored in a list or table.

[0298] Item 16. According to the method of Item 15, where the table is a historical SGPM candidate table.

[0299] Item 17. According to the method of Item 15, the list or table is updated during the encoding / decoding or decoding process.

[0300] Item 18. According to the method of Item 15, the maximum size of the list or table is predefined, transmitted by signal, or derived.

[0301] Item 19. According to the method of Item 15, the list or table is reinitialized at the beginning of one of the following: strip, slice, sub-picture, picture, CTU, CTU row.

[0302] Item 20. According to the method of Item 19, where the list is reinitialized to an empty list, or the table is reinitialized to an empty table.

[0303] Item 21. According to the method of Item 19, the list or table is reinitialized using one or more SGPM candidates that are predefined, derived, or transmitted by signaling.

[0304] Item 22. The method according to Item 15, wherein the manner in which a list or table is used or updated and / or whether a list or table is used or updated depends on the encoding / decoding information.

[0305] Item 23. According to the method of Item 22, wherein the encoding / decoding information includes at least one of the following: block dimension, block size, or block location.

[0306] Item 24. The method according to Item 10, wherein when the current video unit is a chroma video unit, the reused SGPM candidate comes from at least one of the luma video unit or the chroma video unit.

[0307] Item 25. According to the method of Item 1, the mixing matrix used in at least one of SGPM or Intra-Block Replication Geometric Segmentation Mode (IBC-GPM) is derived using the same method as in regression-based GPM.

[0308] Item 26. According to the method of Item 25, whether to use a blending matrix and / or the manner of using a blending matrix is ​​predefined or derived.

[0309] Item 27. The method according to Item 25, wherein a hybrid matrix is ​​used and / or the manner in which a hybrid matrix is ​​used is transmitted via signal transmission.

[0310] Item 28. According to the method of Item 27, one or more syntax elements are used to indicate whether a mixed matrix is ​​used and / or the manner in which a mixed matrix is ​​used.

[0311] Item 29. The method of Item 25, wherein the use of a blending matrix for the first type of video content and / or the manner in which a blending matrix is ​​used for the first type of video content differs from that used for the second type of video content.

[0312] Item 30. According to the method of Item 29, the first type of video content is content captured by a camera, and the second type of video unit is screen content.

[0313] Item 31. The method according to any of Items 1-30, wherein the indication of whether and / or how to obtain one or more SGPM candidates for a video unit is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0314] Item 32. The method according to any of Items 1-30, wherein the indication of whether and / or how to obtain one or more SGPM candidates for a video unit is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

[0315] Item 33. The method according to any one of Items 1-30, wherein whether and / or how one or more SGPM candidates for a video unit are obtained depends on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, strip header, slice group header, codec tree unit (CTU), codec unit (CU), CTU row, CTU group, TU, PU block, video codec unit, position of CU block, position of PU block, position of TU block, position of video codec unit, block dimension of the current block, block dimension of the current block's neighboring blocks, block shape of the current block, block shape of the current block's neighboring blocks, codec mode of the block, indication of color format, codec tree structure, strip group type, slice group type, strip picture type, slice picture type, color components, ID of temporal layer, standard grade, standard level, or standard layer.

[0316] Item 34. According to the method of Item 33, the encoding / decoding mode of the block is at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

[0317] Item 35. The method according to Item 33, wherein the color format is indicated as 4:2:0 or 4:4:4, or wherein the color components are applied to one of the following: chromaticity component or luminance component.

[0318] Item 36. A video processing method comprising: conversion between video units of a video and a bitstream of a video; constructing a candidate list for video units based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in SGPM mode; obtaining a prediction or reconstruction of the video units based on the candidate list; and performing the conversion based on the prediction or reconstruction of the video units.

[0319] Item 37. According to the method of Item 36, one or more SGPM candidates include a triplet of a segmentation mode and two intra-prediction modes.

[0320] Item 38. According to the method of Item 36, one or more SGPM candidates include one or more partitioning patterns.

[0321] Item 39. According to the method of Item 36, one or more SGPM candidates include one or more intra-frame prediction modes.

[0322] Item 40. According to the method of Item 36, one or more SGPM candidates include one or more prediction samples.

[0323] Item 41. According to the method of Item 36, one or more SGPM candidates are derived from video units encoded and decoded in at least one of the following ways: SGPM, SGPMMerge mode, or other SGPM modes.

[0324] Item 42. According to the method of Item 36, where reordering is applied to the candidate list.

[0325] Item 43. According to the method of Item 42, where the reordering depends on template matching or template matching cost.

[0326] Item 44. The method of Item 42, wherein the reordering is the same as that used in SGPM.

[0327] Item 45. According to the method of Item 36, which candidate list is used is either transmitted by signaling, derived, or predefined.

[0328] Item 46. According to the method of Item 36, which SGPM candidate in the candidate list is used in the SGPMMerge mode to obtain the predicted or reconstructed samples transmitted by the signal.

[0329] Item 47. According to the method of Item 46, one or more syntax elements are used to indicate one or more SGPM candidates.

[0330] Item 48. According to the method of Item 36, which SGPM candidate in the candidate list is used in the SGPMMerge mode to obtain the predicted or reconstructed samples is predefined or derived.

[0331] Item 49. The method according to Item 36, wherein the first SGPM segmentation uses a first SGPM Merge candidate list, the second SGPM segmentation uses a second SGPM Merge candidate list, and the first SGPM segmentation is different from the second SGPM segmentation.

[0332] Item 50. According to the method of Item 36, whether or not the video unit is encoded or decoded in SGPMMerge mode is transmitted via a signal using one or more syntax elements (SE).

[0333] Item 51. According to the method of Item 50, one or more syntax elements are binarized using one of the following: fixed-length encoding / decoding, rounding unary encoding / decoding, unary encoding / decoding, or EG encoding / decoding, or one or more syntax elements are encoded as flags.

[0334] Item 52. According to the method of Item 50, one or more syntax elements are either bypassed or context-encoded.

[0335] Item 53. According to the method of Item 52, where the context depends on the encoded / decoded information.

[0336] Item 54. According to the method of Item 53, the encoded / decoded information includes at least one of the following: block dimension, block size, stripe type, picture type, information of neighboring blocks, information of other encoding / decoding tools used for the video unit, or information of the temporal layer.

[0337] Item 55. According to the method of Item 50, one or more syntax elements are transmitted via signaling at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

[0338] Item 56. According to the method of Item 50, one or more syntax elements are encoded and decoded in a predictive manner.

[0339] Item 57. According to the method of Item 50, one or more syntax elements are encoded or decoded conditionally.

[0340] Item 58. According to the method of Item 57, wherein if the first SE indicates that the SGPMMerge mode is applicable, the second SE is signaled to indicate whether the SGPMMerge mode is used.

[0341] Item 59. According to the method of Item 58, wherein the first SE is at one of the following: sequence header, image header, SPS, VPS, DPS, DCI, PPS, APS, strip header, or slice header.

[0342] Item 60. According to the method of Item 58, where the second SE is for the block.

[0343] Item 61. According to the method of Item 57, wherein if the second SE indicates the use of SGPMMerge mode, the third SE is signaled to indicate how to perform SGPMMerge mode.

[0344] Item 62. According to the method of Item 61, the third SE is used to indicate which SGPM candidate is used.

[0345] Item 63. According to the method of Item 36, the application of SGPMMerge mode and / or the manner in which SGPMMerge mode is applied depends on the encoding / decoding information.

[0346] Item 64. According to the method of Item 63, the encoding / decoding information includes at least one of the following: whether the encoding / decoding method is permitted, block dimension, block size, block depth, stripe type, picture type, segmentation tree type, temporal layer identifier, block location, CTU resolution, stripe resolution, slice resolution, subpicture resolution, picture resolution, color format, or color components.

[0347] Item 65. The method according to Item 64, wherein the encoding / decoding method is SGPM, and / or wherein the segmentation tree type includes one of the following: single tree, dual tree, or local dual tree.

[0348] Item 66. According to the method of Item 64, at least one of the SGPM or SGPMMerge modes is applied to all color components.

[0349] Item 67. The method according to Item 64, wherein if at least one of the SGPM or SGPMMerge modes is applied to the chromaticity component, it is different from at least one of the SGPM or SGPMMerge modes applied to the luminance component.

[0350] Item 68. The method according to Item 64, wherein whether at least one of the SGPM or SGPMMerge patterns is applied to the first component depends on whether at least one of the SGPM or SGPMMerge patterns is applied to the second component, and / or wherein the manner in which at least one of the SGPM or SGPMMerge patterns is applied to the first component depends on the manner in which at least one of the SGPM or SGPMMerge patterns is applied to the second component.

[0351] Item 69. According to the method of Item 68, wherein the first component is a chromaticity component and the second component is a luminance component.

[0352] Item 70. The method according to Item 68, wherein at least one of the SGPM or SGPMMerge patterns is applied to the first component in the same manner as for the second component, or wherein at least one of the SGPM or SGPMMerge patterns is applied to the first component in a different manner than for the second component.

[0353] Item 71. The method according to Item 70, wherein the SGPMMerge mode is applied to the luminance component but not to the chrominance component.

[0354] Item 72. According to the method of Item 71, wherein the luminance component is Y in the YCbCr color space or G in the RGB color space.

[0355] Item 73. The method according to Item 71, wherein the chromaticity components are Cb and / or Cr in the YCbCr color space, or R and / or B in the RGB color space.

[0356] Item 74. The method according to Item 36, wherein the indication of the SGPMMerge mode is transmitted via signal based on conditions.

[0357] Item 75. According to the method of Item 74, wherein the conditions include at least one of the following: block dimension, block size, block depth, strip type, picture type, segmentation tree type, temporal layer identifier, block location, CTU resolution, strip resolution, slice resolution, subpicture resolution, picture resolution, color format, or color components.

[0358] Item 76. The method according to any one of Items 36-75, wherein the indication of whether and / or how to construct a candidate list for video units is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0359] Item 77. The method according to any one of Items 36-75, wherein the indication of whether and / or how to construct a candidate list for video units is indicated in one of the following: sequence header, picture header, column parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

[0360] Item 78. The method of any one of Items 36-75, wherein whether and / or how to construct a candidate list for video units depends on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, strip header, slice group header, codec tree unit (CTU), codec unit (CU), CTU row, CTU group, TU, PU block, video codec unit, position of CU block, position of PU block, position of TU block, position of video codec unit, block dimension of the current block, block dimension of the current block's neighboring blocks, block shape of the current block, block shape of the current block's neighboring blocks, codec mode of the block, indication of color format, codec tree structure, strip group type, slice group type, strip picture type, slice picture type, color components, ID of temporal layer, standard grade, standard level, or standard layer.

[0361] Item 79. According to the method of Item 78, the encoding / decoding mode of the block is at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

[0362] Item 80. According to the method of Item 78, the color format is indicated as 4:2:0 or 4:4:4.

[0363] Item 81. The method according to Item 78, wherein the color component is applied to one of the following: chromaticity component or luminance component.

[0364] Item 82. The method according to any one of Items 1-81, wherein the video unit comprises at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, region containing more than one sample point or pixel.

[0365] Item 83. The method according to any one of items 1-82, wherein the conversion includes encoding video units into a bitstream.

[0366] Item 84. The method according to any one of items 1-82, wherein the conversion includes decoding video units from a bitstream.

[0367] Item 85. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1-84.

[0368] Item 86. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform a method according to any one of items 1-84.

[0369] Item 87. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: obtaining one or more spatial geometric segmentation pattern (SGPM) candidates for video units of the video based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; and generating a bitstream based on the prediction or construction of the video units.

[0370] Item 88. A method for storing a bitstream of video, comprising: obtaining one or more spatial geometric segmentation pattern (SGPM) candidates for video units of the video based on at least one of the following: historical information or proximity information; obtaining a prediction or construction of the video units based on the one or more SGPM candidates; generating a bitstream based on the prediction or construction of the video units; and storing the bitstream in a non-transitory computer-readable recording medium.

[0371] Item 89. A non-transitory computer-readable recording medium for storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: constructing a candidate list of video units for the video based on one or more spatial geometric partitioning mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in SGPMMerge mode; obtaining a prediction or reconstruction of the video units based on the candidate list; and generating a bitstream based on the prediction or reconstruction of the video units.

[0372] Item 90. A method for storing a bitstream of video, comprising: constructing a candidate list of video units for a video based on one or more spatial geometric partitioning mode (SGPM) candidates for neighboring video units, wherein the video units are encoded and decoded in SGPM mode; obtaining a prediction or reconstruction of the video units based on the candidate list; generating a bitstream based on the prediction or reconstruction of the video units; and storing the bitstream in a non-transitory computer-readable recording medium.

[0373] Example device Figure 30 A block diagram of a computing device 3000 in which various embodiments of the present disclosure may be implemented is shown. The computing device 3000 may be implemented as a source device 110 (or video encoder 114) or a destination device 120 (or video decoder 124), or may be included in a source device 110 (or video encoder 114) or a destination device 120 (or video decoder 124).

[0374] It should be understood that, Figure 30 The computing device 3000 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0375] like Figure 30 As shown, computing device 3000 includes general-purpose computing device 3000. Computing device 3000 may include at least one or more processors or processing units 3010, memory 3020, storage unit 3030, one or more communication units 3040, one or more input devices 3050, and one or more output devices 3060.

[0376] In some embodiments, the computing device 3000 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 3000 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0377] Processing unit 3010 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 3020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 3000. Processing unit 3010 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0378] Computing device 3000 typically includes various computer storage media. Such media can be any media accessible by computing device 3000, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 3020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 3030 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 3000.

[0379] The computing device 3000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 30 Not shown, but a disk drive for reading from and / or writing to a removable non-volatile disk, and an optical disc drive for reading from and / or writing to a removable non-volatile optical disc may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0380] The communication unit 3040 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in the computing device 3000 can be implemented by a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 3000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0381] Input device 3050 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 3060 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 3040, computing device 3000 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 3000 can also communicate with one or more devices that enable a user to interact with computing device 3000, or, if necessary, with any device that enables computing device 3000 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be performed via an input / output (I / O) interface (not shown).

[0382] In some embodiments, some or all of the components of computing device 3000 may be deployed in a cloud computing architecture, rather than being integrated into a single device. In a cloud computing architecture, components may be remotely provided and work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring end users to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications via a WAN that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, along with the corresponding data, may be stored on servers at remote locations. Computing resources in a cloud computing environment may be consolidated or distributed across remote data center locations. Cloud computing infrastructure may provide services through shared data centers, although they appear as a single access point to users. Therefore, a cloud computing architecture can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, the components and functionality described herein may be provided by conventional servers or installed directly or otherwise on client devices.

[0383] In embodiments of this disclosure, computing device 3000 may be used to implement video encoding / decoding. Memory 3020 may include one or more video codec modules 3025 having one or more program instructions. These modules are accessible and executable by processing unit 3010 to perform the functions of the various embodiments described herein.

[0384] In an example embodiment of performing video encoding, input device 3050 may receive video data as input 3070 to be encoded. The video data may be processed, for example, by video codec module 3025 to generate an encoded bitstream. The encoded bitstream may be provided as output 3080 via output device 3060.

[0385] In an example embodiment performing video decoding, input device 3050 may receive an encoded bitstream as input 3070. The encoded bitstream may be processed, for example, by video codec module 3025 to generate decoded video data. The decoded video data may be provided as output 3080 via output device 3060.

[0386] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A video processing method, comprising: For the conversion between video units and the bitstream of the video, one or more spatial geometric segmentation pattern (SGPM) candidates for the video unit are obtained based on at least one of the following: historical information or proximity information; Based on the one or more SGPM candidates, the prediction or construction of the video unit is obtained; and The conversion is performed based on the prediction or the construction of the video unit.

2. The method of claim 1, wherein the one or more SGPM candidates are derived from at least one of spatially adjacent neighboring video units or spatially non-adjacent neighboring video units.

3. The method of claim 1, wherein the one or more SGPM candidates are derived from temporal video units.

4. The method of claim 1, wherein the one or more SGPM candidates are derived from a history table or list used to store SGPM candidates.

5. The method of claim 1, wherein the one or more SGPM candidates are constructed.

6. The method of claim 5, wherein the one or more SGPM candidates are constructed using codec information of at least one neighboring video unit.

7. The method of claim 6, wherein the intra-frame prediction mode of the at least one neighboring video unit is used.

8. The method of claim 1, wherein the one or more SGPM candidates are used to construct a candidate list of the video units.

9. The method of claim 1, wherein at least one SGPM segmentation is predicted using one of the one or more SGPM candidates.

10. The method of claim 1, wherein one of the SGPM candidates or elements of the SGPM candidates for a video unit encoded or decoded before the current video unit is encoded or decoded before the current video unit is decoded is used for the current video unit.

11. The method of claim 10, wherein the video unit encoded or decoded prior to the current video unit is in one of a different strip, a different slice, a different sub-picture, a different picture, a different codec tree unit (CTU), or a different CTU row.

12. The method of claim 10, wherein the video unit that was encoded or decoded before the current video unit is in one of the same strip, the same slice, the same sub-picture, the same picture, the same CTU, or the same CTU row as the current video unit.

13. The method of claim 10, wherein the video unit that was encoded or decoded before the current video unit is a spatially adjacent video unit.

14. The method of claim 13, wherein the spatially adjacent video unit is adjacent to or not adjacent to the current video unit.

15. The method of claim 10, wherein one or more reused SGPM candidates are stored in a list or table.

16. The method of claim 15, wherein the table is a historical SGPM candidate table.

17. The method of claim 15, wherein the list or table is updated during the encoding / decoding or decoding process.

18. The method of claim 15, wherein the maximum size of the list or table is predefined, transmitted by a signal, or derived.

19. The method of claim 15, wherein the list or table is reinitialized at the beginning of one of the following: strip, slice, sub-image, picture, CTU, CTU row.

20. The method of claim 19, wherein the list is reinitialized to an empty list, or the table is reinitialized to an empty table.

21. The method of claim 19, wherein the list or table is reinitialized using one or more SGPM candidates that are predefined, derived, or transmitted via signaling.

22. The method of claim 15, wherein the manner in which the list or table is used or updated and / or whether the list or table is used or updated depends on the encoding / decoding information.

23. The method of claim 22, wherein the encoding / decoding information includes at least one of the following: block dimension, block size, or block location.

24. The method of claim 10, wherein when the current video unit is a chroma video unit, the reused SGPM candidate comes from at least one of a luma video unit or a chroma video unit.

25. The method of claim 1, wherein the mixing matrix used in at least one of SGPM or Intra-Block Replication Geometry Segmentation Mode (IBC-GPM) is derived using the same method as in regression-based GPM.

26. The method of claim 25, wherein whether to use the mixing matrix and / or the manner of using the mixing matrix is ​​predefined or derived.

27. The method of claim 25, wherein the signal is transmitted whether or not the mixing matrix is ​​used and / or in a manner that uses the mixing matrix.

28. The method of claim 27, wherein one or more syntax elements are used to indicate whether the blending matrix is ​​used and / or the manner in which the blending matrix is ​​used.

29. The method of claim 25, wherein the use of the blending matrix for the first type of video content and / or the manner in which the blending matrix is ​​used for the first type of video content differs from that for the second type of video content.

30. The method of claim 29, wherein the first type of video content is content captured by a camera, and the second type of video unit is screen content.

31. The method according to any one of claims 1-30, wherein an indication of whether or not the one or more SGPM candidates for the video unit are obtained and / or how the one or more SGPM candidates for the video unit are obtained is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.

32. The method according to any one of claims 1-30, wherein an indication of whether the one or more SGPM candidates for the video unit are obtained and / or how the one or more SGPM candidates for the video unit are obtained is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

33. The method according to any one of claims 1-30, wherein whether or not the one or more SGPM candidates for the video unit are obtained and / or how the one or more SGPM candidates for the video unit are obtained depends on at least one of the following: The message indicated in one of the following: DPS, SPS, VPS, PPS, APS, image header, stripe header, slice group header, codec tree unit (CTU), codec unit (CU), CTU line, CTU group, TU, PU block, video codec unit. The location of the CU block The location of the PU block The location of the TU block, The location of the video encoding / decoding unit. The block dimension of the current block. The block dimension of the current block's neighboring blocks. The shape of the current block. The block shape of the current block's neighboring blocks. Block encoding / decoding modes, Indicators of color format, Encoder tree structure, Strip group type, Film set type, Strip image type, Image type Color components, Time-domain ID, Standard level Standard level, or Standard layer.

34. The method of claim 33, wherein the encoding / decoding mode of the block is at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

35. The method of claim 33, wherein the color format indication is 4:2:0 or 4:4:4, or The color component is applied to either the chromaticity component or the luminance component.

36. A video processing method, comprising: For the conversion between video units and the bitstream of the video, a candidate list for the video unit is constructed based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, wherein the video unit is encoded and decoded in SGPM Merge mode; Based on the candidate list, the prediction or reconstruction of the video unit is obtained; and The conversion is performed based on the prediction or reconstruction of the video unit.

37. The method of claim 36, wherein the SGPM candidate in the one or more SGPM candidates comprises a triplet of a segmentation mode and two intra-prediction modes.

38. The method of claim 36, wherein the SGPM candidates in the one or more SGPM candidates include one or more partitioning patterns.

39. The method of claim 36, wherein the SGPM candidates in the one or more SGPM candidates include one or more intra-frame prediction modes.

40. The method of claim 36, wherein the SGPM candidate in the one or more SGPM candidates comprises one or more prediction samples.

41. The method of claim 36, wherein the one or more SGPM candidates are video units encoded and decoded in at least one of the following: SGPM, SGPM Merge mode, or other SGPM modes.

42. The method of claim 36, wherein reordering is applied to the candidate list.

43. The method of claim 42, wherein the reordering depends on template matching or template matching cost.

44. The method of claim 42, wherein the reordering is the same as that used in SGPM.

45. The method of claim 36, wherein which candidate list is used is transmitted by signaling, derived, or predefined.

46. ​​The method of claim 36, wherein which SGPM candidate in the candidate list is used in SGPM Merge mode to obtain predicted or reconstructed samples is transmitted via signal.

47. The method of claim 46, wherein one or more syntax elements are used to indicate the one or more SGPM candidates.

48. The method of claim 36, wherein which SGPM candidate in the candidate list is used in the SGPM Merge mode to obtain predicted or reconstructed samples is predefined or derived.

49. The method of claim 36, wherein the first SGPM segmentation uses a first SGPM Merge candidate list, the second SGPM segmentation uses a second SGPM Merge candidate list, and the first SGPM segmentation is different from the second SGPM segmentation.

50. The method of claim 36, wherein whether the video unit is encoded or decoded in SGPM Merge mode is transmitted via a signal using one or more syntax elements (SE).

51. The method of claim 50, wherein the one or more syntax elements are binarized using one of the following: fixed-length encoding / decoding, rounding unary encoding / decoding, unary encoding / decoding, or EG encoding / decoding, or the one or more syntax elements are encoded as flags.

52. The method of claim 50, wherein the one or more syntax elements are bypassed or context-encoded.

53. The method of claim 52, wherein the context depends on the encoded / decoded information.

54. The method of claim 53, wherein the encoded / decoded information includes at least one of the following: block dimension, block size, stripe type, picture type, information of neighboring blocks, information of other encoding / decoding tools used for the video unit, or information of the temporal layer.

55. The method of claim 50, wherein the one or more syntax elements are transmitted via signaling at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice header.

56. The method of claim 50, wherein the one or more syntax elements are encoded and decoded in a predictive manner.

57. The method of claim 50, wherein the one or more syntax elements are encoded or decoded conditionally.

58. The method of claim 57, wherein if the first SE indicates that the SGPM Merge mode is applicable, the second SE is signaled to indicate whether the SGPM Merge mode is used.

59. The method of claim 58, wherein the first SE is in one of the following: sequence header, image header, SPS, VPS, DPS, DCI, PPS, APS, strip header, or slice header.

60. The method of claim 58, wherein the second SE is directed to a block.

61. The method of claim 57, wherein if the second SE indicates the use of SGPM Merge mode, the third SE is signaled to indicate how to perform SGPM Merge mode.

62. The method of claim 61, wherein the third SE is used to indicate which of the SGPM candidates is used.

63. The method of claim 36, wherein whether or not the SGPM Merge mode is applied and / or the manner in which the SGPMMerge mode is applied depends on the encoding / decoding information.

64. The method of claim 63, wherein the encoding / decoding information comprises at least one of the following: Are encoding / decoding methods allowed? Block dimension, Block size, Block depth, Strip type, Image type, Segmentation tree type, Temporal layer identifier, Block location, CTU resolution, Strip resolution, Image resolution, Sub-image resolution, Image resolution, Color format, or Color components.

65. The method of claim 64, wherein the encoding / decoding method is SGPM, and / or The segmentation tree type mentioned therein includes one of the following: single tree, double tree, or local double tree.

66. The method of claim 64, wherein at least one of SGPM or SGPM Merge mode is applied to all color components.

67. The method of claim 64, wherein if at least one of the SGPM or SGPM Merge modes is applied to the chromaticity component, it is different from at least one of the SGPM or SGPM Merge modes applied to the luminance component.

68. The method of claim 64, wherein whether at least one of the SGPM or SGPM Merge patterns is applied to the first component depends on whether at least one of the SGPM or SGPM Merge patterns is applied to the second component, and / or The manner in which at least one of the SGPM or SGPM Merge patterns is applied to the first component depends on the manner in which at least one of the SGPM or SGPM Merge patterns is applied to the second component.

69. The method of claim 68, wherein the first component is a chromaticity component and the second component is a luminance component.

70. The method of claim 68, wherein the manner in which at least one of the SGPM or SGPM Merge modes is applied to the first component is the same as that applied to the second component, or The manner in which at least one of the SGPM or SGPM Merge modes is applied to the first component differs from the manner applied to the second component.

71. The method of claim 70, wherein the SGPM Merge mode is applied to the luminance component but not to the chrominance component.

72. The method according to claim 71, wherein the luminance component is Y in the YCbCr color space or G in the RGB color space.

73. The method according to claim 71, wherein the chromaticity components are Cb and / or Cr in the YCbCr color space, or R and / or B in the RGB color space.

74. The method of claim 36, wherein the indication of the SGPM Merge mode is transmitted via signaling based on conditions.

75. The method of claim 74, wherein the condition comprises at least one of the following: Block dimension, Block size, Block depth, Strip type, Image type, Segmentation tree type, Temporal layer identifier, Block location, CTU resolution, Strip resolution, Image resolution, Sub-image resolution, Image resolution, Color format, or Color components.

76. The method according to any one of claims 36-75, wherein the indication of whether to construct the candidate list for the video unit and / or how to construct the candidate list for the video unit is indicated at one of the following: sequence level, Image group level, Image quality, strip level, or Film series level.

77. The method of any one of claims 36-75, wherein the indication of whether to construct the candidate list for the video unit and / or how to construct the candidate list for the video unit is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

78. The method according to any one of claims 36-75, wherein whether and / or how the candidate list for the video unit is constructed depends on at least one of the following: The message indicated in one of the following: DPS, SPS, VPS, PPS, APS, image header, stripe header, slice group header, codec tree unit (CTU), codec unit (CU), CTU line, CTU group, TU, PU block, video codec unit. The location of the CU block The location of the PU block The location of the TU block, The location of the video encoding / decoding unit. The block dimension of the current block. The block dimension of the current block's neighboring blocks. The shape of the current block. The block shape of the current block's neighboring blocks. Block encoding / decoding modes, Indicators of color format, Encoder tree structure, Strip group type, Film set type, Strip image type, Image type Color components, Time-domain ID, Standard level Standard level, or Standard layer.

79. The method of claim 78, wherein the encoding / decoding mode of the block is at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, or non-IBC sub-block mode.

80. The method of claim 78, wherein the color format indication is 4:2:0 or 4:4:

4.

81. The method of claim 78, wherein the color component is applied to one of the following: a chromaticity component or a luminance component.

82. The method according to any one of claims 1-81, wherein the video unit comprises at least one of the following: Color components, Sub-images, strip, piece, Code-decode tree unit (CTU) CTU line, CTU group, Codec Unit (CU) Prediction Unit (PU) Transformer Unit (TU) Code-decode tree block (CTB). Code block (CB) Predicted blocks (PB). Transform block (TB) piece, Sub-blocks of a block Sub-regions within the block A region containing more than one sample point or pixel.

83. The method according to any one of claims 1-82, wherein the conversion comprises encoding the video unit into the bitstream.

84. The method according to any one of claims 1-82, wherein the conversion comprises decoding the video unit from the bitstream.

85. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-84.

86. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1-84.

87. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: One or more spatial geometric segmentation pattern (SGPM) candidates for video units of the video are obtained based on at least one of the following: historical information or proximity information; Based on the one or more SGPM candidates, the prediction or construction of the video unit is obtained; and The bitstream is generated based on the prediction or construction of the video unit.

88. A method for storing a bitstream of video, comprising: One or more spatial geometric segmentation pattern (SGPM) candidates for video units of the video are obtained based on at least one of the following: historical information or proximity information; The prediction or construction of the video unit is obtained based on the one or more SGPM candidates; The bitstream is generated based on the prediction or construction of the video unit; as well as The bitstream is stored in a non-transitory computer-readable recording medium.

89. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: Based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, a candidate list of video units for the video is constructed, wherein the video units are encoded and decoded in SGPM Merge mode. Based on the candidate list, the prediction or reconstruction of the video unit is obtained; and The bitstream is generated based on the prediction or reconstruction of the video unit.

90. A method for storing a bitstream of video, comprising: Based on one or more Spatial Geometric Partitioning Mode (SGPM) candidates for neighboring video units, a candidate list of video units for the video is constructed, wherein the video units are encoded and decoded in SGPM Merge mode. Based on the candidate list, the prediction or reconstruction of the video unit is obtained; The bitstream is generated based on the prediction or reconstruction of the video unit; as well as The bitstream is stored in a non-transitory computer-readable recording medium.