Methods, apparatus and media for video processing

By applying inter-frame prediction and encoding/decoding tools, including intra-frame block copying and intra-frame template matching prediction, the video encoding/decoding process is optimized and the encoding/decoding efficiency is improved.

CN122139359APending Publication Date: 2026-06-02DOUYIN VISION CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2024-10-25
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially in the process of converting video units to bitstreams. The application of tools such as intra-frame block copying and intra-frame template matching prediction has not been fully optimized.

Method used

Inter-frame prediction and encoding/decoding tools, including intra-block copy (IBC) or intra-template matching prediction (intraTMP), are employed and applied to multiple sub-segments or sub-blocks of a video unit, which are then converted through a first target encoding/decoding mode to improve encoding/decoding performance.

Benefits of technology

By applying inter-frame prediction and encoding/decoding tools, the encoding/decoding performance of video is improved, and the encoding/decoding efficiency is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122139359A_ABST
    Figure CN122139359A_ABST
Patent Text Reader

Abstract

Embodiments of this disclosure provide a solution for video processing. A method for video processing is proposed. The method includes: for conversion between video units and video bitstreams of a video, applying a first target encoding / decoding mode to the video units, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; and performing the conversion based on the first target encoding / decoding mode.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to video processing techniques, and more specifically, to geometric segmentation patterns with intra-frame block copying. Background Technology

[0002] Today, digital video capabilities are being applied to all aspects of people's lives. Various video compression technologies have been proposed for video encoding / decoding, such as MPEG-2, MPEG-4, ITU-TH.263, ITU-TH.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-TH.265 High Efficiency Video Codec (HEVC) standard, and Multifunctional Video Codec (VVC) standard. However, the overall expectation is to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention

[0003] Embodiments of this disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method includes: for a conversion between video units and a bitstream of video, applying a first target encoding / decoding mode to the video units, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; and performing the conversion based on the first target encoding / decoding mode. Compared to conventional solutions, the method according to the first aspect of this disclosure can improve encoding / decoding performance by applying the first target encoding / decoding mode.

[0005] In a second aspect, an apparatus for video processing is provided. The apparatus includes a processor and a non-transitory memory having instructions thereon. When executed by the processor, the instructions cause the processor to perform the method according to the first aspect of this disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. This non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of this disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: applying a first target encoding / decoding mode to video units of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; and generating a bitstream based on the first target encoding / decoding mode.

[0008] In a fifth aspect, a method for storing a bitstream of video is proposed. The method includes: applying a first target encoding / decoding mode to video units of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of the following: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; generating a bitstream based on the first target encoding / decoding mode; and storing the bitstream in a non-transitory computer-readable recording medium.

[0009] This summary aims to present, in a simplified form, the selected concepts further described below in the detailed embodiments. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Attached Figure Description

[0010] The above and other objects, features, and advantages of exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.

[0011] Figure 1 A block diagram illustrating an example video codec system according to some embodiments of the present disclosure is shown; Figure 2 A block diagram illustrating a first example video encoder according to some embodiments of the present disclosure is shown; Figure 3 A block diagram illustrating an example video decoder according to some embodiments of the present disclosure is shown; Figure 4 An example of an encoder block diagram is shown; Figure 5 67 intra-frame prediction modes are shown; Figure 6A and Figure 6B Reference samples used for wide-angle intra-frame prediction are shown respectively; Figure 7The discontinuity problem is shown when the orientation exceeds 45°; Figure 8A and Figure 8B The MMVD search points are shown respectively; Figure 9 A schematic diagram for symmetric MVD mode is shown; Figure 10 The extended CU region used in BDOF is shown; Figure 11 The top and left neighbor blocks used in the CIIP weight derivation are shown; Figure 12A and Figure 12B The affine motion models based on control points are shown respectively; Figure 13 The affine MVF for each sub-block is shown; Figure 14 The location of the inherited affine motion prediction value is shown; Figure 15 This demonstrates the inheritance of control point motion vectors; Figure 16 The locations of candidate positions for constructing the affine Merge pattern are shown; Figure 17 A schematic diagram illustrating the use of motion vectors for the proposed combination method is shown; Figure 18 The sub-block MV VSB and pixel Δv(i,j) are shown; Figure 19A and Figure 19B The SbTMVP process in VVC is shown below; Figure 20 Local illumination compensation is shown; Figure 21 This shows no downsampling for the short side; Figure 22 This shows the refinement of motion vectors on the decoding side; Figure 23 The diamond-shaped area in the search region is shown; Figure 24 The locations of the spatial merge candidates are shown; Figure 25 The candidate pairs considered for redundancy checks of spatial merge candidates are shown. Figure 26 A schematic diagram of motion vector scaling for temporal Merge candidates is shown; Figure 27 The candidate positions, C0 and C1, for the time-domain Merge candidate are shown; Figure 28 The VVC spatial neighboring blocks of the current block are shown; Figure 29 A schematic diagram of the virtual block in the i-th round of search is shown; Figure 30 An example of GPM partitioning grouped at the same angle is shown; Figure 31 The unidirectional prediction MV selection for geometric segmentation patterns is shown; Figure 32 This demonstrates the generation of blended weights using geometric segmentation patterns. Examples; Figure 33 The spatial neighboring blocks used to derive spatial merge candidates are shown; Figure 34 This demonstrates template matching execution over the search area surrounding the initial MV; Figure 35 A schematic diagram of a sub-block in an OBMC application is shown; Figure 36 The location, type, and transformation type of the SBT are shown; Figure 37 The neighboring samples used to calculate SAD are shown; Figure 38 The neighboring samples used to calculate SAD for sub-CU level motion information are shown; Figure 39 The sorting process is shown; Figure 40 The reordering process in the encoder is shown; Figure 41 The reordering process in the decoder is shown; Figure 42 The IBC reference area is shown, depending on the current CU location; Figure 43 An example of symmetry is shown in a screen content image; Figure 44A A schematic diagram of BV adjustment for horizontal flipping is shown; Figure 44B A schematic diagram of BV adjustment for vertical flipping is shown; Figure 45 An example of how to derive AR-BVP is shown; Figure 46 The five positions in Bn are shown; Figure 47 The intra-frame template matching search area used is shown; Figure 48 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; Figure 49A block diagram of a computing device in which various embodiments of the present disclosure may be implemented is shown.

[0012] In all accompanying drawings, the same or similar reference numerals usually refer to the same or similar elements. Detailed Implementation

[0013] The principles of this disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described for illustrative purposes only and to help those skilled in the art understand and implement this disclosure, and do not imply any limitation on the scope of this disclosure. In addition to the methods described below, the disclosure described herein can be implemented in various other ways.

[0014] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0015] The terms "an embodiment," "embodiment," "example embodiment," etc., used in this disclosure refer to embodiments that may include specific features, structures, or characteristics, but not every embodiment is required to include that specific feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in conjunction with an example embodiment, it is claimed that, whether explicitly described or not, such a feature, structure, or characteristic affecting its relation to other embodiments is within the knowledge of those skilled in the art.

[0016] It should be understood that although the terms “first” and “second”, etc., may be used herein to describe various elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments. As used herein, the singular forms “a,” “an,” and “the” are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising,” “including,” “having,” “containing,” and / or “comprising” as used herein indicate the presence of the said features, elements, and / or components, but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.

[0018] Example Environment Figure 1This is a block diagram illustrating an example video encoding / decoding system 100 from which the techniques of this disclosure may be utilized. As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0019] Video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, interfaces for receiving video data from video content providers, computer graphics systems for generating video data, and / or combinations thereof.

[0020] Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded images and associated data. An encoded image is an encoded representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. Encoded video data can be directly transmitted to destination device 120 via network 130A through I / O interface 116. Encoded video data may also be stored on storage medium / server 130B for access by destination device 120.

[0021] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may acquire encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or it may be external to the destination device 120, which is configured to interface with an external display device.

[0022] The video encoder 114 and the video decoder 124 can operate according to video compression standards such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other existing and / or future standards.

[0023] Figure 2This is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be... Figure 1 An example of a video encoder 114 in system 100 is shown.

[0024] The video encoder 200 can be configured to implement any or all of the technologies disclosed herein. Figure 2 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0025] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0026] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode, in which at least one reference picture is the picture in which the current video block is located.

[0027] Furthermore, although some components (such as motion estimation unit 204 and motion compensation unit 205) can be integrated, for interpretable purposes, these components are... Figure 2 The examples are shown separately.

[0028] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0029] The mode selection unit 203 can, for example, select one of several coding modes (intra-coding or inter-coding) based on the error result, and provide the resulting intra-coded or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference image. In some examples, the mode selection unit 203 can select an intra-inter-prediction joint prediction (CIIP) mode, in which prediction is based on inter-prediction signals and intra-prediction signals. In the case of inter-prediction, the mode selection unit 203 can also select a resolution for the block based on the motion vector (e.g., sub-pixel precision or integer pixel precision).

[0030] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and decoded samples of images from buffer 213 other than the image associated with the current video block.

[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-strip, P-strip, or B-strip. As used herein, an "I-strip" can refer to a portion of an image composed of macroblocks, all of which are based on macroblocks within the same image. Furthermore, as used herein, in some aspects, "P-strip" and "B-strip" can refer to portions of an image composed of macroblocks independent of macroblocks within the same image.

[0032] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating a reference image in list 0 or list 1 that includes the reference video block, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0033] Alternatively, in other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search reference images in list 0 to find a reference video block for the current video block, and can also search reference images in list 1 to find another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference images in lists 0 and 1 that include multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and multiple motion vectors of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.

[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoder's decoding process. Alternatively, in some embodiments, the motion estimation unit 204 can reference the motion information of another video block to transmit the motion information of the current video block via a signal. For example, the motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0035] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0036] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0037] As discussed above, the video encoder 200 can transmit motion vectors via signals in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0038] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0040] In other examples, such as in skip mode, residual data for the current video block may not exist, and residual generation unit 207 may not perform a subtraction operation.

[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0042] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.

[0044] After the video block is reconstructed by reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.

[0046] Figure 3 This is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be... Figure 1 An example of video decoder 124 in system 100 is shown.

[0047] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 3 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0048] exist Figure 3 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally contrasted with the encoding process described with respect to the video encoder 200.

[0049] Entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy-encoded video data, and motion compensation unit 302 can determine motion information from the entropy-decoded video data, including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which involves deriving several most likely candidates based on data from neighboring PBs and reference pictures. Motion information typically includes horizontal motion vector displacement values ​​and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B-strip, an identifier of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially or temporally neighboring blocks.

[0050] The motion compensation unit 302 can generate motion compensation blocks, possibly by performing interpolation based on an interpolation filter. Identifiers for interpolation filters used with sub-pixel precision can be included in the syntax elements.

[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate the interpolated values ​​of sub-integer pixels for the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.

[0052] Motion compensation unit 302 may use at least some of the syntax information to determine the size of the blocks used to encode the encoded video sequence (multiple frames) and / or (multiple stripes), segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a pattern indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence. As used herein, in some respects, a “strip” can refer to a data structure that can be decoded independently of other stripes of the same image in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A strip can be an entire image or a region of an image.

[0053] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Dequantization unit 304 dequantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0054] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates decoded video for presentation on a display device.

[0055] Some exemplary embodiments of this disclosure will be described in detail below. It should be noted that section headings are used in this document for ease of understanding and not to limit the embodiments disclosed in a section to that section. Furthermore, although some embodiments are described with reference to multi-function video codecs or other specific video codecs, the disclosed techniques are also applicable to other video codec techniques. Furthermore, although some embodiments describe video encoding steps in detail, it should be understood that the decoding steps corresponding to decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates. 1. Summary of the Invention This disclosure relates to video codec technology. Specifically, it relates to Geometric Partition Mode (GPM) and Intra-Block Copy (IBC), how to use IBC in GPM mode, and other codec tools in image / video codecs. It can be applied to existing video codec standards such as HEVC or VVC. It can also be applied to future video codec standards or video codecs.

[0057] 2. Introduction Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed the H.261 and H.263 standards, while ISO / IEC developed MPEG-1 and MPEG-4 Vision. These two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Codec (AVC) standard, and the H.265 / HEVC standard. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established to work on the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.

[0058] 2.1. Encoder / decoder streams of typical video codecs Figure 4 An example of a VVC encoder block diagram is shown, comprising three loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Compensation (SAO), and ALF. Unlike DF, which uses predefined filters, SAO and ALF utilize the original samples of the current image, reducing the mean square error between the original and reconstructed samples by adding compensation and applying finite impulse response (FIR) filters, respectively, and by utilizing the side information from the encoding and decoding through signal transmission compensation and filter coefficients. ALF is located in the last processing stage of each image and can be viewed as a tool attempting to capture and repair artifacts caused by previous stages.

[0059] 2.2. Intra-mode encoding and decoding with 67 intra-prediction modes To capture arbitrary edge directions presented in natural video, such as Figure 5 As shown, the number of directional intra-prediction modes has been expanded from 33 used in HEVC to 65, while the planar and DC modes remain unchanged. These denser directional intra-prediction modes are applicable to all block sizes and both luma and chroma intra-prediction.

[0060] In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division is needed to generate intra-prediction values ​​using DC mode. In VVC, blocks can have rectangular shapes, which typically requires division for each block. To avoid division for DC prediction, only the longer side is used to calculate the average of non-square blocks.

[0061] 2.2.1. Wide-angle intra-frame prediction Although 67 modes are defined in VVC, the precise prediction direction for a given intra-prediction mode index depends on the block shape. Regular angular intra-prediction directions are defined clockwise from 45 degrees to -135 degrees. In VVC, for non-square blocks, several regular angular intra-prediction modes are adaptively replaced with wide-angle intra-prediction modes. The replaced modes are transmitted via signaling using the original mode index, which is then remapped to the wide-angle mode index after resolution. The total number of intra-prediction modes remains unchanged at 67, and the intra-mode encoding / decoding method remains unchanged.

[0062] To support these predicted directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as follows: Figure 6A and Figure 6B As shown, they illustrate reference samples used for wide-angle intra-frame prediction.

[0063] The number of modes replaced in the wide-angle directional mode depends on the aspect ratio of the block. The replaced intra-prediction modes are shown in Table 1.

[0064] Table 1 – Intra-prediction modes replaced by wide-angle mode

[0065] Figure 7 This illustrates the problem of discontinuities when the orientation exceeds 45°. For example... Figure 7 As shown, in the case of wide-angle intra-frame prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the increased gap. p α The negative impact of wide-angle mode. If the wide-angle mode represents a non-fractional offset. There are 8 wide-angle modes that satisfy this condition, namely [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted through these modes, the samples in the reference cache are directly copied without applying any interpolation. This modification reduces the number of samples that need to be smoothed. In addition, it aligns the design of non-fractional modes in regular prediction modes with that of wide-angle modes.

[0066] In VVC, in addition to 4:2:0, 4:2:2 and 4:4:4 chroma formats are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, with the number of entries expanded from 35 to 67 to align with the expansion of intra-prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra-prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values ​​in the mapping table entries to more accurately convert the prediction angles for chroma blocks.

[0067] 2.3. Inter-frame prediction For each inter-frame prediction CU, motion parameters consist of a motion vector, a reference picture index, and a reference picture list usage index, along with additional information required for inter-frame prediction sample generation using new encoding / decoding features of the VVC. Motion parameters can be transmitted via signaling in an explicit or implicit manner. When a CU is encoded / decoded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no encoded / decoded motion vector increments, or reference picture indices. A Merge mode is specified, where motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates and additional scheduling introduced in the VVC. The Merge mode can be applied to any inter-frame prediction CU, not just skip mode. An alternative to the Merge mode is explicit transmission of motion parameters, where the motion vector for each reference picture list, the corresponding reference picture index, the reference picture list usage flag, and other necessary information are explicitly transmitted via signaling for each CU.

[0068] 2.4. Intra-Block Copying (IBC) Intra-Block Copy (IBC) is a tool used in the HEVC extension on SCC. It is well known to significantly improve the encoding and decoding efficiency of screen content material. Since IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed within the current image. The luma block vector of an IBC-encoded CU is integer precise. The chroma block vector is also rounded to integer precision. When combined with AMVR, IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. IBC-encoded CUs are considered a third prediction mode, distinct from intra-frame or inter-frame prediction modes. IBC mode is suitable for CUs with a width and height of 64 luma samples or less.

[0069] On the encoder side, hash-based motion estimation for IBC is performed. The encoder performs RD check on blocks with a width or height no greater than 16 luminance samples. For non-Merge mode, block vector search is first performed using a hash-based search. If the hash search does not return valid candidates, a local search based on block matching is performed.

[0070] In hash-based search, hash key matching (32-bit CRC) between the current block and reference blocks is extended to all allowed block sizes. Hash key calculation for each location in the current image is based on 4x4 sub-blocks. For the larger current block, a hash key match with a reference block is determined when all hash keys of all 4x4 sub-blocks match the hash key at the corresponding reference location. If multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the lowest cost is selected.

[0071] In block matching search, the search scope is set to cover both the previous CTU and the current CTU.

[0072] At the CU level, IBC mode is transmitted via signaling using a flag, and it can be transmitted via signaling as either IBC AMVP mode or IBC skip / Merge mode as follows: IBC Skip / Merge Mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC codec blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates, and paired candidates.

[0073] IBC AMVP mode: Block vector difference is encoded and decoded in the same way as motion vector difference. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the upper neighbor (if IBC encoded and decoded). When either neighbor is unavailable, the default block vector is used as the prediction value. A flag is transmitted via signaling to indicate the block vector prediction value index.

[0074] 2.5. Multi-Model Learning (MMLM) The CCLM included in VVC is extended by adding three multi-model LM (MMLM) modes. In each MMLM mode, using a threshold as the average of the luminance reconstruction neighboring samples, the reconstructed neighboring samples are classified into two categories. The linear model for each category is derived using the least mean square (LMS) method. For the CCLM mode, the LMS method is also used to derive the linear model. Slope adjustment is applied to both the cross-component linear model (CCLM) and multi-model LM predictions. This adjustment is a linear function that maps luminance values ​​to chrominance values, tilted relative to a center point determined by the average luminance values ​​of the reference samples.

[0075] 2.6. Merge Pattern with MVD (MMVD) In addition to the implicitly derived motion information being directly used in the Merge pattern for generating prediction samples of the current CU, a Merge pattern with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is transmitted via signaling immediately after the regular Merge flag is sent to indicate whether the MMVD pattern is used in the CU.

[0076] In MMVD, after a Merge candidate is selected, it is further refined through MVD information transmitted via signals. This further information includes a Merge candidate flag, an index specifying the amplitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the top two candidates in the Merge list is selected as the MV basis. The MMVD candidate flag is transmitted via signals to specify which candidate to use between the first and second Merge candidates.

[0077] The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. Figure 8A and Figure 8B The MMVD search point is shown. (Example) Figure 8A and Figure 8B As shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2.

[0078] Table 2 – Relationship between Distance Index and Predefined Offset

[0079] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent the four directions shown in Table 3. It is important to note that the meaning of the MVD sign can vary depending on the information of the starting MV. When the starting MV is a unidirectional or bidirectional prediction MV, and both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the sign in Table 3 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV, and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), and the POC difference in list 0 is greater than the POC difference in list 1, the sign in Table 3 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Otherwise, if the POC difference in list 1 is greater than the POC difference in list 0, the sign in Table 3 specifies the sign of the MV offset added to the list 1 MV component of the starting MV, and the sign of the list 0 MV has the opposite value.

[0080] MVD is scaled based on the difference in POC in each direction. If the difference in POC is the same in both lists, no scaling is needed. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, then the MVD of list 1 is scaled, as shown below. Figure 26 As shown, the POC difference of L0 is defined as td, and the POC difference of L1 is defined as tb. If the POC difference of L1 is greater than that of L0, the MVD of list 0 is scaled in the same way. If the initial MV is unidirectionally predicted, the MVD is added to the available MV.

[0081] Table 3 – Symbols of MV Offsets Specifyed by Direction Index

[0082] 2.7. Symmetric MVD Encoding and Decoding In VVC, in addition to the normal one-way and two-way predictive MVD signaling modes, the symmetric MVD mode of two-way predictive MVD signaling is also applied. In the symmetric MVD mode, the reference image indices of both List 0 and List 1, as well as the motion information of the MVD of List 1, are not transmitted through signals but are derived.

[0083] The decoding process of the symmetric MVD mode is as follows: 1) At the strip level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to 0.

[0084] Otherwise, if the most recent reference image in list 0 and the most recent reference image in list 1 form a forward and backward reference pair or a backward and forward reference pair, then BiDirPredFlag is set to 1, and both the reference images in list 0 and list 1 are short-term reference images. Otherwise, BiDirPredFlag is set to 0.

[0085] 2) At the CU level, if the CU is bidirectional predictive codec and BiDirPredFlag is equal to 1, the symmetry mode flag indicating whether the symmetry mode is used is explicitly transmitted via signaling.

[0086] When the symmetry mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly transmitted via signaling. The reference indices of list 0 and list 1 are each set to equal a pair of reference images. MVD1 is set to equal to (-MVD0). The final motion vector is shown in the following formula.

[0087]

[0088] Figure 9 A schematic diagram of the symmetric MVD model is shown.

[0089] In the encoder, symmetric MVD motion estimation begins with an initial MV evaluation. A set of initial MV candidates includes MVs obtained from a one-way prediction search, MVs obtained from a two-way prediction search, and MVs from an AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.

[0090] 2.8. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF (formerly known as BIO) was included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires far less computation, especially in terms of the number of multiplications and multiplier size.

[0091] BDOF is used to refine the bidirectional prediction signal of the CU at the 4x4 sub-block level. BDOF is applied to the CU if all of the following conditions are met: The CU is encoded and decoded using a “true” bidirectional prediction mode, meaning that one of the two reference images is displayed before the current image in the order of display, and the other is displayed after the current image in the order of display.

[0092] The distances from the two reference images to the current image (i.e., the difference in point of view) are the same.

[0093] Both reference images are short-term reference images.

[0094] CU is not encoded or decoded using affine mode or SbTMVP Merge mode.

[0095] The CU has more than 64 luminance samples.

[0096] Both the CU height and CU width are greater than or equal to 8 luminance samples.

[0097] The BCW weight index indicates equal weights.

[0098] WP is not enabled for the current CU.

[0099] The current CU does not use CIIP mode.

[0100] BDOF is applied only to the luminance component. As the name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4x4 sub-block, motion refinement is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Motion refinement is then used to adjust the bidirectional prediction sample values ​​in the 4x4 sub-blocks. The following steps are applied during the BDOF process.

[0101] First, the horizontal and vertical gradients of the two predicted signals. It is calculated by directly calculating the difference between two neighboring sample points, that is,

[0102] in It is a list Coordinates of the predicted signal in The sample value at the location, and shift1 is calculated based on the luminance bit depth bitDepth as shift1 = max(6, bitDepth-6).

[0103] Then, the autocorrelation and cross-correlation of the gradients. Calculated as

[0104] in

[0105] in It is a 6x6 window surrounding a 4x4 sub-block, and The values ​​are set to min(1, bitDepth) respectively. 11) and min(4, bitDepth) 8).

[0106] Then the motion is refined. The cross-correlation and autocorrelation terms are derived using the following formulas:

[0107] in It is a floor function, and .

[0108] Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4x4 sub-block:

[0109] Finally, the BDOF samples of CU are calculated by adjusting the bidirectional prediction samples as follows:

[0110] These values ​​were chosen such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.

[0111] To derive the gradient values, the list outside the current CU boundary is used. Some predicted samples in ) It needs to be generated. Figure 10 The extended CU region used in BDOF is shown. For example... Figure 10 As shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region (white area) are generated by directly taking reference samples at nearby integer positions (using floor() operations on coordinates) without interpolation, and a normal 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values ​​are used only for gradient calculation. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled from their nearest neighbors (i.e., repeated).

[0112] When the width and / or height of a CU is greater than 16 luminance samples, it will be divided into sub-blocks with a width and / or height equal to 16 luminance samples, and the sub-block boundaries will be considered as CU boundaries in the BDOF process. The maximum cell size for the BDOF process is limited to 16x16. The BDOF process can be skipped for each sub-block. The BDOF process is not applied to the sub-block when the SAD between the initial L0 and L1 prediction samples is less than a threshold. The threshold is set to equal to (8 * W * ( H >> 1 ), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 and L1 prediction samples calculated in the DVMR process is reused here.

[0113] Bidirectional optical flow (BDOF) is disabled if BCW is enabled for the current block, meaning the BCW weight index indicates unequal weights. Similarly, BDOF is disabled if WP is enabled for the current block, meaning either luma_weight_lx_flag is 1 for either of the two reference images. BDOF is also disabled when the CU is encoded / decoded using symmetric MVD mode or CIIP mode.

[0114] 2.9. Inter-frame and Intra-frame Joint Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU includes at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and height are less than 128 luma samples, an additional flag is transmitted via signaling to indicate whether Inter-Frame Intra-Frame Joint Prediction (CIIP) mode is applied to the current CU. As the name suggests, CIIP prediction combines inter-frame prediction signals with intra-frame prediction signals. The inter-frame prediction signal in CIIP mode... The inter-frame prediction process is derived using the same procedure as the regular Merge mode; and the intra-frame prediction signal... The standard intra-frame prediction process with a planar pattern is derived. Then, the intra-frame prediction signal and the inter-frame prediction signal are combined using a weighted average, where the weight values ​​are based on the encoding / decoding patterns of the top neighbor block and the left neighbor block (in...). Figure 11 Depicted in the middle, Figure 11 The top neighbor block and left neighbor block used in the CIIP weight derivation are shown below: If the top neighbor is available and is intra-coded, set isIntraTop to 1; otherwise, set isIntraTop to 0. If the left neighbor is available and intra-frame encoded, set isIntraLeft to 1; otherwise, set isIntraLeft to 0. If (isIntraLeft + isIntraTop) equals 2, then wt is set to 3; Otherwise, if (isIntraLeft + isIntraTop) equals 1, then wt is set to 2; Otherwise, set wt to 1.

[0115] The CIIP predictions are formed as follows:

[0116] 2.10. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 12A and Figure 12B Affine motion models based on control points are shown respectively, for example, Figure 12A It is a 4-parameter affine model, and Figure 12B This is a 6-parameter affine model. For example... Figure 12A and Figure 12B As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).

[0117] For a 4-parameter affine motion model, the location of the sample point in the block ( x, y The motion vector at point () is derived as:

[0118] For a 6-parameter affine motion model, the location of the sample points in the block ( x, y The motion vector at point () is derived as:

[0119] in( mv 0x , mv 0y ) is the motion vector of the upper left control point, ( mv 1x , mv 1y ) is the motion vector of the upper right control point, and ( mv 2x , mv 2y ) is the motion vector of the lower left control point.

[0120] To simplify motion compensation prediction, block-based affine transformation prediction is applied. To derive the motion vector for each 4x4 luma sub-block, the motion vector of the center sample point of each sub-block is calculated according to the above equation (e.g., ...). Figure 13 As shown, Figure 13 The affine MVF for each sub-block is shown and rounded to 1 / 16 fractional precision. A motion-compensated interpolation filter is then applied to generate a prediction for each sub-block with a derived motion vector. The sub-block size for the chroma component is also set to 4x4. The MV of the 4x4 chroma sub-block is calculated as the average of the MVs of the four corresponding 4x4 luma sub-blocks.

[0121] Similar to translational motion inter-frame prediction, there are two affine motion inter-frame prediction modes: affine Merge mode and affine AMVP mode.

[0122] 2.10.1. Affine Merge Prediction The AF_MERGE mode can be applied to CUs with a width and height greater than or equal to 8. In this mode, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPVM candidates, and the one to be used for the current CU is indicated by a signal transmission index. The following three types of CPVM candidates are used to form the affine merge candidate list: Inherited affine Merge candidates inferred from the CPMV of neighboring CUs.

[0123] Affine Merge candidate CPMVP constructed using translational MV derivation of neighboring CUs.

[0124] Zero MV.

[0125] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of the neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Figure 14 The location of the inherited affine motion prediction values ​​is shown. Candidate blocks are as follows: Figure 14 As shown. For the left-side prediction, the scan order is A0->A1, while for the top-side prediction, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between candidates from two inherited sides. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. As shown, if the neighboring lower-left block A is encoded and decoded in affine mode, the motion vectors of the upper-left, upper-right, and lower-left corners of the CU including block A are obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are based on... Computed. When block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated according to... Calculated.

[0126] Figure 15 The inheritance of control point motion vectors is shown.

[0127] The constructed affine candidate means that the candidate is constructed by combining the neighboring translational motion information of each control point. Figure 16 The locations of candidate positions for the affine Merge pattern used in construction are shown. The motion information of the control points is derived from... Figure 16 The spatial and temporal nearest neighbors shown are derived. CPMV k (k=1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. If available, the TMVP is used as CPMV4.

[0128] After obtaining the motion signatures (MVs) of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used for sequential construction: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, { CPMV1, CPMV2}, { CPMV1, CPMV3}.

[0129] Combining three CPMVs constructs a 6-parameter affine merge candidate, and combining two CPMVs constructs a 4-parameter affine merge candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.

[0130] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list.

[0131] 2.10.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with a width and height both greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated by sequentially using the following four types of CPVM candidates: Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.

[0132] A constructible AMVP candidate CPMVP derived using the translation MV of neighboring CUs.

[0133] Translation MV from the neighboring CU.

[0134] Zero MV.

[0135] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AVMP candidates, only affine CUs with the same reference picture as those in the current block are considered. No deduplication is applied when inserting inherited affine motion predictions into the candidate list.

[0136] The constructed AMVP candidate is from Figure 16 The specified spatial nearest neighbor derivation is shown. The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture index of neighboring blocks is checked. The block that is first inter-frame encoded / decoded in the checking order and has the same reference picture as the current CU is used. Only one block is considered if the current CU is encoded / decoded in a 4-parameter affine mode, and... mv 0 and mv If all three CPMVs are available, they are added as candidates in the affine AMVP list. When the current CU is encoded / decoded using a 6-parameter affine mode and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.

[0137] If the affine AMVP list still has fewer than 2 candidates after checking the inherited and constructed AMVP candidates, then mv 0、 mv 1 and mv2 will be added sequentially as translation MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to populate the affine AMVP list.

[0138] 2.10.3. Affine Motion Information Storage In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is used only to generate inherited CPMVPs in the affine Merge mode and affine AMVP mode for the most recently encoded / decoded CU. Subblock MVs derived from the CPMV are used for motion compensation, MV derivation of the Merge / AMVP list for translation MVs, and deblocking.

[0139] To avoid image row caching for additional CPMVs, affine motion data inheritance from CUs of the upper CTU is handled differently from inheritance from regular neighboring CUs. If a candidate CU for affine motion data inheritance is in the upper CTU row, the lower left and lower right sub-block MVs in the row cache, instead of the CPMV, are used for affine MVP derivation. Thus, the CPMV is only stored in the local cache. If the candidate CU is a 6-parameter affine codec, the affine model is downgraded to a 4-parameter model. Figure 17 A schematic diagram illustrating the use of motion vectors for the proposed combination method is shown. For example... Figure 17 As shown, along the top CTU boundary, the motion vectors of the lower left and lower right sub-blocks of the CU are used for the affine inheritance of the CU in the bottom CTU.

[0140] 2.10.4. Refinement of Optical Flow Prediction for Affine Modes Compared to pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of prediction accuracy. To achieve finer-grained motion compensation, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth used for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the brightness prediction samples are refined by adding the difference derived from the optical flow equation. PROF is described in the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate sub-block predictions. .

[0141] Step 2) Calculate the spatial gradient of the sub-block prediction at each sample location using a 3-tap filter [-1, 0, 1]. The gradient calculation is exactly the same as the gradient calculation in BDOF.

[0142]

[0143] This is used to control the precision of the gradient. Each side of the sub-block (i.e., 4x4) prediction is expanded by one sample point for gradient computation. To avoid additional memory bandwidth and additional interpolation computation, those expanded samples on the expanded boundaries are copied from the nearest integer pixel position in the reference image.

[0144] Step 3) Brightness prediction refinement is calculated using the following optical flow equation.

[0145]

[0146] in It refers to the location of the sample points. The calculated sample MV (by (representation) and sample points The difference between the MVs of the sub-blocks of the same sub-block, such as Figure 18 As shown, Figure 18 The sub-block MV VSB and pixel Δv(i,j) are shown (small arrows). It is quantized in units of 1 / 32 brightness sample precision.

[0147] Since the affine model parameters and the sample point positions relative to the sub-block center do not change from one sub-block to another, therefore It can be computed for the first sub-block and reused for other sub-blocks in the same CU. This allows... From the sample point location To the center of the sub-block Horizontal and vertical offsets It can be derived from the following equation:

[0148] To maintain accuracy, the center of the sub-block Calculated as ( ( W SB - 1 ) / 2, ( H SB - 1) / 2), where W SB and H SB These are the width and height of the sub-block.

[0149] For a 4-parameter affine model

[0150] For a 6-parameter affine model

[0151] in These are the motion vectors of the control points at the top left, top right, and bottom left. These are the width and height of the CU.

[0152] Step 4) Finally, refine the brightness prediction. Added to sub-block prediction Final prediction I’ It is generated as the following equation.

[0153]

[0154] PROF is not applied in two cases for affine codecs: 1) all control points MV are the same, which indicates that the CU has only translational motion; 2) the affine motion parameters are greater than the specified limit, because the sub-block-based affine MC is downgraded to the CU-based MC to avoid large memory access bandwidth requirements.

[0155] Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. PROF is not applied during the affine motion estimation stage in the following two cases: a) if the CU is not the root block and its parent block does not choose an affine mode as its optimal mode, then PROF is not applied because the probability of the current CU choosing an affine mode as its optimal mode is low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold, and the current image is not a low-latency image, then PROF is not applied because the improvement introduced by PROF is small in this case. Thus, affine motion estimation using PROF can be accelerated.

[0156] 2.11. Sub-block-based temporal motion vector prediction (SbTMVP) VVC supports a sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-image to improve motion vector prediction and merge patterns for CUs in the current image. The same co-image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in two main aspects: TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level. TMVP obtains temporal motion vectors from co-occurring blocks in the co-occurring image (the co-occurring block is the lower right or center block relative to the current CU). SbTMVP applies motion shift before obtaining temporal motion information from the co-occurring image, where the motion shift is obtained from the motion vector of one of the spatial neighboring blocks from the current CU.

[0157] Figure 19A and Figure 19B The SbTMVP process in VVC is shown below. Figure 19AThe spatial neighbor block used by ATMVP is shown. Figure 19B This diagram illustrates how the motion field of a sub-CU is derived by applying motion shifts from spatial neighbors and scaling motion information from the corresponding co-located sub-CU. The SbTVMP process is as follows: Figure 19A and Figure 19B As shown. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, it checks... Figure 19A The spatial nearest neighbor A1 in the image is selected. If A1 has a motion vector that uses a co-located image as its reference image, then that motion vector is chosen as the motion shift to be applied. If no such motion is identified, the motion shift is set to (0, 0).

[0158] In the second step, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block) to the position of the current block. Figure 19B The corresponding image shown obtains motion information (motion vectors and reference indices) at the sub-CU level. Figure 19B The example assumes that motion shift is set as the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the center sample) in the co-location image is used to derive the motion information for the sub-CU. After the motion information of the co-location sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with the reference image of the current CU.

[0159] In VVC, a sub-block-based Merge list, including both SbTVMP candidates and affine Merge candidates, is used for signaling in sub-block-based Merge mode. SbTVMP mode is enabled / disabled via the Sequence Parameter Set (SPS) flag. If SbTVMP mode is enabled, the SbTVMP prediction is added as the first entry in the sub-block-based Merge candidate list, followed by the affine Merge candidate. The size of the sub-block-based Merge list is transmitted via signaling in the SPS, and the maximum allowed size of the sub-block-based Merge list in VVC is 5.

[0160] The sub-CU size used in SbTMVP is fixed at 8x8, and like the affine Merge pattern, the SbTMVP pattern only applies to CUs with a width and height greater than or equal to 8.

[0161] The encoding logic for the additional SbTMVP Merge candidate is the same as that for other Merge candidates, that is, for each CU in the P-strip or B-strip, an additional RD check is performed to determine whether to use the SbTMVP candidate.

[0162] 2.12. Adaptive Motion Vector Resolution (AMVR) In HEVC, when `use_integer_mv_flag` in the strip header is equal to 0, the motion vector difference (MVD) (between the CU's motion vector and the predicted motion vector) is transmitted through the signal in units of quarter-luminance samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the CU's MVD to be encoded and decoded with different precisions. Depending on the current CU's mode (normal AMVP mode or affine AVMP mode), the current CU's MVD can be adaptively selected as follows: Normal AMVP mode: quarter brightness sample, half brightness sample, integer brightness sample or four brightness sample.

[0163] Affine AMVP mode: quarter brightness sample, integer brightness sample, or 1 / 16 brightness sample.

[0164] If the current CU has at least one non-zero MVD component, the MVD resolution indication at the CU level is conditionally transmitted via signal transmission. If all MVD components (i.e., both the horizontal and vertical MVD of reference list L0 and reference list L1) are zero, the quarter-spot luminance sample MVD resolution is presumed.

[0165] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luminance sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required, and quarter-luminance sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate that half-luminance sample or other MVD precision (integer or quad-luminance sample) is used for the normal AMVP CU. In the case of half-luminance sample, a 6-tap interpolation filter is used instead of the default 8-tap interpolation filter for the half-luminance sample position. Otherwise, a third flag is signaled to indicate whether integer luminance sample MVD precision or quad-luminance sample MVD precision is used for the normal AMVP CU. In the case of an affine AMVP CU, the second flag is used to indicate whether integer luminance sample MVD precision or 1 / 16 luminance sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luminance sample, half-luminance sample, integer luminance sample, or quad-luminance sample), the CU's motion vector prediction value is rounded to the same precision as the MVD before being added to the MVD. The motion vector predictions are rounded to zero (i.e., negative motion vector predictions are rounded to positive infinity, and positive motion vector predictions are rounded to negative infinity).

[0166] The encoder uses RD checksums to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM11, RD checks for MVD accuracy other than quarter-luminance samples are only conditionally invoked. For normal AVMP mode, the RD costs for quarter-luminance sample MVD accuracy and integer luminance sample MVD accuracy are first calculated. Then, the RD costs for integer luminance sample MVD accuracy are compared with those for quarter-luminance sample MVD accuracy to determine if it is necessary to further check the RD costs for four-luminance sample MVD accuracy. When the RD cost for quarter-luminance sample MVD accuracy is significantly less than the RD cost for integer luminance sample MVD accuracy, the RD check for four-luminance sample MVD accuracy is skipped. Then, if the RD cost for integer luminance sample MVD accuracy is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for half-luminance sample MVD accuracy is skipped. For affine AMVP mode, if the affine inter-frame mode is not selected after checking the rate-distortion cost of affine Merge / Skip mode, Merge / Skip mode, quarter-lumen sample MVD precision normal AMVP mode, and quarter-lumen sample MVD precision affine AMVP mode, then the 1 / 16-lumen sample MVD precision and 1-pixel MVD precision affine inter-frame modes are not checked. Furthermore, in the 1 / 16-lumen sample and quarter-lumen sample MVD precision affine inter-frame modes, the affine parameters obtained in the quarter-lumen sample MVD precision affine inter-frame mode are used as the starting search point.

[0167] 2.13. Bidirectional prediction with CU-level weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference images and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended beyond simple averaging to allow for a weighted average of the two prediction signals.

[0168]

[0169] Five weights are allowed in weighted average two-way forecasting. For each bidirectional prediction CU, the weights w are determined in one of two ways: 1) for non-merge CUs, the weight index is transmitted via signal after the motion vector difference; 2) for merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW is applied only to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency images, all 5 weights are used. For non-low-latency images, only 3 weights (w∈{3,4,5}) are used.

[0170] At the encoder, a fast search algorithm is applied to find the weight indices without significantly increasing encoder complexity. These algorithms are summarized below. For more details, readers can refer to the VTM software and documentation JVET-L0646. When combined with AMVR, if the current image is a low-latency image, unequal weights are checked only conditionally for 1-pixel and 4-pixel motion vector precision.

[0171] When combined with an affine pattern, an affine ME will be performed for unequal weights if and only if the affine pattern is selected as the current best pattern.

[0172] When the two reference images in bidirectional prediction are the same, unequal weights are checked only under certain conditions.

[0173] Under certain conditions, unequal weights are not searched, depending on the POC distance between the current image and its reference image, the QP of the encoding / decoding algorithm, and the temporal level.

[0174] The BCW weight index is encoded using a context-coded bit followed by a bypass-coded bit. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is signaled via bypass coding to indicate which unequal weight was used.

[0175] Weighted Prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards for efficiently encoding and decoding video content with fade-in and fade-out effects. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be transmitted via signaling for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the corresponding weights and offsets of the reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if the CU uses WP, the BCW weight index is not transmitted via signaling, and w is presumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge patterns and inherited affine Merge patterns. For constructed affine Merge patterns, affine motion information is constructed based on motion information from up to 3 blocks. The BCW index of the CU using the constructed affine Merge pattern is simply set to be equal to the BCW index of the first control point MV.

[0176] In VVC, CIIP and BCW cannot be used together for a CU. When a CU is encoded or decoded in CIIP mode, the BCW index of the current CU is set to 2, for example, with equal weights.

[0177] 2.14. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is an encoding / decoding tool used to address the problem of local illumination variations between the current image and its temporal reference image. LIC is based on a linear model, where a scaling factor and offset are applied to the reference samples to obtain the predicted samples for the current block. Specifically, LIC can be mathematically modeled by the following equation:

[0178] in In coordinates The prediction signal for the current block at that location; It is composed of motion vectors The reference block it points to; These are the corresponding scaling factors and offsets applied to the reference block. Figure 20 Local lighting compensation is shown. Figure 20 The LIC procedure is shown. Figure 20 In this context, when LIC is applied to blocks, the Minimum Mean Square Error (LMSE) method is employed to minimize the neighboring samples of the current block (i.e., Figure 20 templates in T ) and its corresponding reference sample in the time-domain reference image (i.e., Figure 20 In T0 or T1 The difference between ) is used to derive the LIC parameters (i.e., The value of ). Furthermore, to reduce computational complexity, both the template sample and the reference template sample are downsampled (adaptive downsampling) to derive the LIC parameters; that is, only Figure 20 The shaded samples in the data were used for derivation. .

[0179] Figure 21 This shows the short side without downsampling. To improve encoding / decoding performance, such as... Figure 21 As shown, downsampling is not performed on the short side.

[0180] 2.15. Decoder-side motion vector refinement (DMVR) To improve the accuracy of motion vector refinement (MV) in Merge mode, a decoding-side refinement based on bilateral matching (BM) is applied in VVC. In the bidirectional prediction operation, the refined MV is searched around the initial MV in reference image lists L0 and L1. The BM method computes the distortion between two candidate blocks in reference image lists L0 and L1. Figure 22 The motion vector refinement on the decoding side is shown. For example... Figure 22 As shown, the SAD between two blocks is calculated based on each MV candidate (e.g., MV0' and MV1') around the initial MV. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal.

[0181] In VVC, the application of DMVR is limited and can only be applied to CUs encoded and decoded using the following modes and features: CU-level Merge pattern with bidirectional prediction of MV.

[0182] Relative to the current image, one reference image is from the past and the other is from the future.

[0183] The distances from the two reference images to the current image (i.e., the difference in point of view) are the same.

[0184] Both reference images are short-term reference images.

[0185] The CU has more than 64 luminance samples.

[0186] Both the CU height and CU width are greater than or equal to 8 luminance samples.

[0187] The BCW weight index indicates equal weights.

[0188] WP is disabled for the current block.

[0189] CIIP mode is not used for the current block.

[0190] The refined motion vector (MV) derived through the DMVR process is used to generate inter-frame prediction samples and also for temporal motion vector prediction in future image encoding and decoding. The original MV is used in the deblocking process and also for spatial motion vector prediction in future CU encoding and decoding.

[0191] Additional features of DMVR are mentioned in the following sub-entries.

[0192] 2.15.1. Search Scheme In DVMR, the search point revolves around the initial MV, and the MV offset follows the MV difference mirror rule. In other words, any point examined by DMVR (represented by the candidate MV pair (MV0, MV1)) obeys the following two equations:

[0193] in This represents the refinement offset between the initial MV and the refined MV in one of the reference images. The refinement search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase.

[0194] A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is calculated first. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the loss of uncertainty in DMVR refinement, a bias towards the forward MV is proposed during the DMVR process. The SAD between the reference blocks of the initial MV candidate references is reduced by 1 / 4 of the SAD value.

[0195] The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the surface equation of parameter error, rather than through an additional search using SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration with the minimum SAD at the center.

[0196] In subpixel offset estimation based on parametric error surfaces, the cost at the center location and the costs at the four nearest neighbor locations are used to fit a two-dimensional parabolic error surface equation of the following form.

[0197] in( This corresponds to the score position with the minimum cost, and C corresponds to the minimum cost. The above equation is solved by using the costs of the five search points. Calculated as:

[0198] The value is automatically constrained between -8 and 8 because all values ​​are positive, and the minimum value is... This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated score ( It is added to the integer distance thinning MV to obtain the subpixel accurate thinning increment MV.

[0199] 2.15.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 of a lumen sample. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search point surrounds the initial fractional pixel MV with an integer sample offset, so samples at those fractional positions need to be interpolated for the DMVR search process. To reduce computational complexity, a bilinear interpolation filter is used to generate fractional samples for the search process in DMVR. Another important effect of using a bilinear filter is that, utilizing a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal MC process, samples that are not needed by the interpolation process based on the original MV but are needed by the interpolation process based on the refined MV are filled from these available samples.

[0200] 2.15.3. Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luminance samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luminance samples. The maximum cell size for the DMVR search process is limited to 16x16.

[0201] 2.16. Multi-pass decoder-side motion vector refinement In this contribution, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector predictions.

[0202] 2.16.1. First pass – Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying the BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and the two reference blocks in L1.

[0203] BM performs a local search to derive the integer sample precision intDeltaMV and the half-pixel sample precision halfDeltaMv. The local search applies a 3x3 square search pattern, looping over the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0204] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of distortion between reference blocks. The local search of intDeltaMV or halfDeltaMV is terminated when bilCost at the center point of the 3x3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.

[0205] The existing fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as follows: MV0_pass1 = MV0 + deltaMV, MV1_pass1 = MV1 – deltaMV.

[0206] 2.16.2. Second pass – Sub-block-based bilateral matching MV refinement In the second pass, the refined MV is derived by applying the BM to 16x16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass for the reference image lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and the two reference sub-blocks in L1.

[0207] For each sub-block, BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values ​​of sHor and sVer are determined by the block dimension, and the maximum value of sHor and sVer is 8.

[0208] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference subblocks, as follows: bilCost = satdCost * costFactor. The search region (2 * sHor + 1) * (2 * sVer + 1) is divided into Figure 23 The diagram shows a maximum of five diamond-shaped search regions. Each search region is assigned a costFactor, determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond region is processed sequentially starting from the center of the search region. Within each region, search points are processed in raster scan order, starting from the top left corner and ending at the bottom right corner. A complete integer-pixel search terminates when the minimum bilCost within the current search region is less than a threshold (equal to sbW * sbH); otherwise, the complete integer-pixel search continues to the next search region until all search points have been checked.

[0209] BM performs a local search to derive the half-sample precision halfDeltaMv. The search pattern and cost function are the same as defined in 2.9.1.

[0210] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MV at the second pass is derived as follows: MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2), MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2).

[0211] 2.16.3. Third pass – Sub-block based bidirectional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8x8 grid sub-blocks. For each 8x8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy, without clipping from the refined MV of the parent block in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.

[0212] The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) at the third pass are derived as follows: MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv, MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv.

[0213] 2.17. Sample-based BDOF In sample-based BDOF, motion refinement (Vx, Vy) is not derived based on block derivation (Vx, Vy), but is performed on a per-sample basis.

[0214] The encoding / decoding block is divided into 8x8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD (Solution-Adjustment) between two reference sub-blocks and a threshold. If BDOF is applied to the sub-block, a sliding 5x5 window is used for each sample in the sub-block, and the existing BDOF process is applied to derive Vx and Vy for each sliding window. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample values ​​for the center sample of the window.

[0215] 2.18. Extended Merge Forecast In VVC, the Merge candidate list is constructed by including the following five types of candidates in sequence: (1) Airspace MVP from the airspace adjacent to the CU.

[0216] (2) Temporal MVP from the same CU.

[0217] (3) History-based MVP from FIFO table.

[0218] (4) Pair average MVP.

[0219] (5) Zero MV.

[0220] The size of the merge list is transmitted via signaling in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU encoded in Merge mode, the index of the best merge candidate is encoded using rounded unary binarization (TU). The first bit of the merge index is encoded using the context, and bypass encoding is used for the remaining bits.

[0221] This section provides the derivation process for each category of Merge candidates. Similar to HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a region of a specific size.

[0222] 2.18.1. Derivation of Airspace Candidates The derivation of spatial merge candidates in VVC is the same as that in HEVC, except that the positions of the first two merge candidates are swapped. Figure 24 The locations of the airspace merge candidates are shown. (In the area located...) Figure 24 At most four merged candidates are selected from the candidates at the indicated positions. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only if one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or if it is intra-frame encoded / decoded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates. This redundancy check ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Figure 25 The candidate pairs considered for redundancy checking of spatial merge candidates are shown. Conversely, only the candidate pairs considered are shown. Figure 25 The system uses arrow links to select pairs, and only adds candidates to the list if the corresponding candidates used for redundancy checks do not have the same motion information.

[0223] 2.18.2. Derivation of Time-Domain Candidates In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images to be used for the derivation of the co-located CU is explicitly transmitted via signal transmission in the strip header. Figure 26 A schematic diagram illustrating motion vector scaling for temporal merge candidates is shown. Figure 26 As shown by the dashed lines, the scaled motion vectors of the temporal merge candidate are obtained by scaling the motion vectors of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located reference image and the co-located image. The reference image index of the temporal merge candidate is set to 0.

[0224] like Figure 27 As shown, the position for the temporal candidate is selected between candidate C0 and C1. If the CU at position C0 is unavailable, intra-frame encoded or decoded, or outside the current line of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal merge candidate.

[0225] 2.18.3. Derivation of Merge Candidates Based on History Historically based MVP (HMVP) merge candidates are added to the merge list, following the spatial MVP and TMVP. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-frame encoding / decoding CU is present, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0226] The HMVP table size S is set to 6, indicating that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, the first-in, first-out (FIFO) rule of the constraints is utilized, where a redundancy check is first applied to find if a duplicate HMVP already exists in the table. If found, the duplicate HMVP is removed from the table, and all subsequent HMVP candidates are moved forward.

[0227] HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked sequentially and inserted into the candidate list after the TMVP candidates. Redundancy checks are applied to HMVP candidates for both spatial and temporal merge candidates.

[0228] To reduce the number of redundant check operations, the following simplifications are introduced: The number of HMPV candidates used in the Merge list generation is set to (N<= 4)? M: (8 (N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table.

[0229] Once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1, the process of building the Merge candidate list from the HMVP is terminated.

[0230] 2.18.4. Derivation of Pairwise Average Merge Candidates Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing Merge candidate list. These predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid.

[0231] If the Merge list is not full after adding pairwise average Merge candidates, a zero MVP will be inserted at the end until the maximum number of Merge candidates is reached.

[0232] 2.18.5. Merge Estimation Region The Merge Estimation Region (MER) allows for the independent derivation of Merge candidate lists for CUs within the same Merge Estimation Region (MER). Candidate blocks located within the same MER as the current CU are not included in the generation of the Merge candidate list for the current CU. Furthermore, the update process for the historical motion vector prediction candidate list is only updated when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left brightness sample position of the current CU in the image, and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and is transmitted via signaling as log2_parallel_merge_level_minus2 in the sequence parameter set.

[0233] 2.19. New Merge Candidates 2.19.1. Derivation of Non-Adjacent Merge Candidates Figure 28 This shows the VVC (Virtual Domain Controller) neighboring blocks of the current block. Within the VVC, Figure 28 The five spatial neighbor blocks and one temporal nearest neighbor shown were used to derive the Merge candidate.

[0234] We propose using the same pattern as in VVC to derive additional Merge candidates from positions not adjacent to the current block. To achieve this, for each search round i, the virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated using the following formula: Offsetx =-i×gridX, Offsety = -i×gridY Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.

[0235] Secondly, the width and height of the virtual block are calculated using the following formula: newWidth = i×2×gridX+ currWidthnewHeight = i×2×gridY +currHeight.

[0236] Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.

[0237] gridX and gridY are currently set to currWidth and currHeight, respectively.

[0238] Figure 29 This illustrates the relationship between the virtual block and the current block.

[0239] After the virtual block is generated, block A i B i C i D i and E i These can be considered VVC spatial neighbor blocks, and their positions are obtained using the same pattern as the pattern in the VVC. Clearly, if the search round i is 0, the virtual block is the current block. In this case, block A... i B i C i D i and E i It is a spatial neighbor block used in VVC Merge mode.

[0240] When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbor blocks are utilized.

[0241] Non-adjacent spatial merge candidates are inserted into the merge list after the temporal merge candidates in the order B1->A1->C1->D1->E1.

[0242] 2.19.2.STMVP We propose using three spatial merge candidates and one temporal merge candidate to derive the average candidate as the STMVP candidate.

[0243] STMVP is inserted before the Merge candidate in the upper left airspace.

[0244] STMVP candidates were deduplicated along with all previous Merge candidates in the Merge list.

[0245] For airspace candidates, the top three candidates in the current Merge candidate list are used.

[0246] For time-domain candidates, use the same position as the VTM / HEVC co-position.

[0247] For airspace candidates, the first, second, and third candidates inserted into the current Merge candidate list before STMVP are denoted as F, S, and T.

[0248] A time-domain candidate with the same position as the VTM / HEVC co-position used in TMVP is denoted as Col.

[0249] The motion vector (denoted as mvLX) of the STMVP candidate in the prediction direction X is derived as follows: 1) If all four Merge candidates have valid reference indices and are all equal to 0 in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_S + mvLX_T + mvLX_Col)>>2 2) If the reference indices of three of the four Merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F × 3 + mvLX_S × 3 + mvLX_Col × 2)>>3 or mvLX = (mvLX_F × 3 + mvLX_T × 3 + mvLX_Col × 2)>>3 or mvLX = (mvLX_S × 3 + mvLX_T × 3 + mvLX_Col × 2)>>3 3) If the reference indices of two of the four Merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_Col)>>1 or mvLX = (mvLX_S + mvLX_Col)>>1 or mvLX = (mvLX_T + mvLX_Col)>>1 Note: STMVP mode is turned off if time-domain candidates are not available.

[0250] 2.19.3. Merge list size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8.

[0251] 2.20. Geometric Partitioning (GPM) In VVC, geometric segmentation modes are supported for inter-frame prediction. Geometric segmentation modes are transmitted via signaling using a CU-level flag as a merge mode. Other merge modes include regular merge mode, MMVD mode, CIIP mode, and sub-block merge mode. For each possible CU size... (in Excluding 8x64 and 64x8, the geometric segmentation mode supports a total of 64 segments.

[0252] Figure 30 An example of GPM partitioning grouped at the same angle is shown. When using this mode, the CU is divided into two parts by geometrically positioned straight lines ( Figure 30 The position of the dividing line is mathematically derived from the angle and offset parameters of a specific segment. Each part of the geometric segmentation in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each segment, i.e., each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with conventional bidirectional prediction, only two motion-compensated predictions are needed per CU. The unidirectional prediction motion for each segment is derived using the process described in 2.20.1.

[0253] If a geometric segmentation mode is used for the current CU, the geometric segmentation index (angle and offset) and two merge indices (one for each segment) are further indicated via signal transmission. The number of maximum GPM candidate sizes is explicitly transmitted in SPS, and syntax binarization is specified for the GPM merge indices. After predicting each part of the geometric segmentation, a blending process with adaptive weights is used, and the sample values ​​along the geometric segmentation edges are adjusted, as shown in 2.20.2. This is the prediction signal for the entire CU, and the transformation and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric segmentation mode is stored, as shown in 2.20.3.

[0254] 2.20.1. Construction of One-Way Prediction Candidate List The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 2.18. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X equals the parity of n) is used as the nth unidirectional prediction motion vector for the geometric segmentation pattern. Figure 31 The unidirectional prediction MV selection for geometric segmentation mode is shown. These motion vectors are in Figure 31 The value is marked with "x". If the corresponding LX motion vector of the nth extended Merge candidate does not exist, then the L(1) of the same candidate... The X motion vector is used as a unidirectional predictive motion vector for the geometric segmentation pattern.

[0255] 2.20.2. Blending along the geometric segmentation edges After predicting each part of the geometric segment using its own motion, a blend is applied to the two predicted signals to derive samples around the segmentation edges. The blending weights at each location of the CU are derived based on the distance between the individual location and the segmentation edge.

[0256] Location The distance to the segmentation edge is derived as follows:

[0257] in It is an index for the angle and offset of the geometric segmentation, which depends on the geometric segmentation index emitted by the signal. The sign depends on the angle index. .

[0258] The weights of each part of the geometric segment are derived as follows:

[0259] partIdx depends on the angle index Weight An example in Figure 32 It is shown in the middle. Figure 32 The mixed weights using geometric segmentation patterns are shown. An example of generation.

[0260] 2.20.3. Motion field storage for geometric segmentation patterns Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded by the geometric segmentation pattern.

[0261] The type of motion vector stored for each individual location in the sports field is determined as follows:

[0262] Where motionIdx equals It is recalculated from equation (2-18). partIdx depends on the angle index. .

[0263] If sType equals 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field; otherwise, if sType equals 2, then the combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.

[0264] Otherwise, if Mv1 and Mv2 come from the same list, only the unidirectional predicted motion Mv2 is stored.

[0265] 2.21. Multiple Hypothesis Prediction In multiple hypothesis prediction (MHP), up to two additional predictions are transmitted over the signal, in addition to inter-frame AMVP mode, regular Merge mode, affine Merge mode, and MMVD mode. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.

[0266]

[0267] Weighting factor α As specified in Table 4 below: Table 4 – Weighting factors for MHP

[0268] For inter-frame AMVP mode, MHP is applied only when unequal weights in BCW are selected in bidirectional prediction mode.

[0269] Additional assumptions can be either Merge mode or AMVP mode. In Merge mode, motion information is indicated by the Merge index, and the Merge candidate list is the same as in the geometric segmentation mode. In AMVP mode, the reference index, MVP index, and MVD are transmitted via signals.

[0270] 2.22. Non-adjacent airspace candidates Non-adjacent airspace merge candidates are inserted after TMVP in the regular merge candidate list. Figure 33The spatial neighbor blocks used to derive spatial merge candidates are shown. The pattern of spatial merge candidates is shown in... Figure 33 The distance between non-adjacent spatial candidates and the current codec block is shown above. The distance between the current codec block and the candidate spatial domain is based on the width and height of the current codec block.

[0271] 2.23. Template Matching (TM) Template matching (TM) is a decoder-side MV derivation method used to refine the motion information of the current CU by finding the closest match between a template in the current image (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference image (i.e., the same size as the template). Figure 34 This illustrates template matching execution over the search area surrounding the initial MV. (Example) Figure 34 As shown, within the search range of [-8, +8] pixels, a better MV is searched around the initial motion of the current CU. The template matching previously proposed in JVET-J0021 is adopted with two modifications: the search step size is determined based on the AMVR mode, and in the Merge mode, the TM can be cascaded with the bilateral matching process.

[0272] In AMVP mode, MVP candidates are determined based on template matching error, selecting the one that minimizes the difference between the current block template and the reference block template. TM then performs MV refinement only on that specific MVP candidate. TM refines the MVP candidate using an iterative diamond search, starting with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. AMVP candidates can be further refined using a cross search with full-pixel MVD precision (or 4 pixels for 4-pixel AMVR mode), followed by half-pixels and quarter-pixels sequentially according to the AMVR mode specified in Table 5. This search process ensures that the MVP candidate maintains the same MV precision as indicated by the AMVR mode after the TM process.

[0273] Table 5 – Search patterns for AMVR and search patterns using AMVR's Merge mode

[0274] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 5, TM can proceed up to 1 / 8 pixel MVD precision, or skip those precisions beyond half-pixel MVD precision, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process or as an additional MV refinement process between block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enable condition check.

[0275] 2.24. Overlapping Block Motion Compensation (OBMC) Overlapping Block Motion Compensation (OBMC) was previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, it is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. Furthermore, it is applied to both luma and chroma components. In JEM, MC blocks correspond to codec blocks. When a CU is encoded and decoded using sub-CU modes (including sub-CU merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To handle CU boundaries uniformly, OBMC is performed at the sub-block level for all MC block boundaries, where the sub-block size is set to equal to 4×4, such as... Figure 35 As shown. Figure 35 A schematic diagram of a sub-block of an OBMC application is shown.

[0276] When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of the four connected neighboring sub-blocks (if available and different from the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block.

[0277] The predicted block based on the motion vectors of neighboring sub-blocks is represented as P N ,in N Instructions for neighboring Above , Down square , Left side and right side The index of the sub-block, and the predicted block based on the motion vector of the current sub-block is represented as: P C .when P N When based on the motion information of neighboring sub-blocks that include the same motion information as the current sub-block, the OBMC is not removed from... PN Execute. Otherwise, P N Each sample point was added P C The same points in, that is, P N Four rows / columns were added P C The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} were used. P N And the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} were used P C An exception is small MC blocks (i.e., when the height or width of the codec block is equal to 4 or the CU is encoded / decoded using subCU mode). For small MC blocks, P N Only two rows / columns were added P C In this case, the weighting factors {1 / 4, 1 / 8} are used. P N And the weighting factors {3 / 4, 7 / 8} were used P C For motion vectors generated based on vertical (horizontal) neighboring sub-blocks P N , P N Samples in the same row (column) are added with the same weighting factor. P C .

[0278] In JEM, for CUs with a size of 256 lumen samples or less, a CU level flag is transmitted via signaling to indicate whether OBMC is applied to the current CU. For CUs with a size greater than 256 lumen samples or not encoded / decoded using AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to the CU, its effect is considered during the motion estimation phase. The predicted signal formed by OBMC using motion information from the top and left neighbor blocks is used to compensate for the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.

[0279] 2.25. Multiple Transform Selection (MTS) for Core Transformation In addition to DCT-II, which is already used in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding and decoding of both inter-frame and intra-frame codec blocks. It uses multiple transforms selected from DCT8 / DST7. Newly introduced transform matrices are DST-VII and DCT-VIII. Table 6 shows the basis functions of the selected DST / DCT.

[0280] Table 6 – Transform basis functions for DCT-II / VIII and DSTVII for N-point inputs

[0281] To maintain the orthogonality of the transformation matrices, the transformation matrices are quantized more precisely than those in HEVC. To keep the intermediate values ​​of the transformation coefficients within the 16-bit range, all coefficients must have 10 bits after both the horizontal and vertical transformations.

[0282] To control the MTS scheme, separate enable flags are specified at the SPS level for intra-frame and inter-frame operations. When MTS is enabled at SPS, CU-level flags are signaled to indicate whether MTS is applied. Here, MTS is applied only to luminance. MTS signaling is skipped when one of the following conditions is met.

[0283] The position of the last effective coefficient of luminance TB is less than 1 (i.e., DC only).

[0284] The last effective coefficient of luminance TB is located in the MTS zero region.

[0285] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are transmitted via signaling to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 7. A unified transform selection for ISP and implicit MTS is used by removing intra-frame mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra-frame block and both intra-frame explicit MTS and inter-frame explicit MTS are enabled, only DST7 is used for both the horizontal and vertical transform cores. An 8-bit main transform core is used when transform matrix precision is involved. Therefore, all transform cores used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform cores, including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8, use an 8-bit main transform core.

[0286] Table 7 –Transformation and signaling mapping table

[0287] To reduce the complexity of large-sized DST-7 and DCT-8 blocks, the high-frequency transform coefficients are zeroed for DST-7 and DCT-8 blocks with dimensions (width or height, or both) equal to 32. Only the coefficients in the 16x16 low-frequency region are retained.

[0288] In HEVC, for example, block residuals can be encoded and decoded using transform skip mode. To avoid redundancy in syntax encoding and decoding, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Implicit MTS can also be enabled when MTS is enabled for inter-frame encoding / decoding blocks.

[0289] 2.26. Subblock Transformation (SBT) In VTM, a sub-block transform is introduced for inter-frame prediction CUs. In this transform mode, only a sub-part of the residual block is encoded / decoded for the CU. When the inter-frame prediction CU has cu_cbf equal to 1, the signal cu_sbt_flag can be used to indicate whether the entire residual block or a sub-part of the residual block is encoded / decoded. In the former case, the inter-frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is encoded / decoded using a presumed adaptive transform, and another portion of the residual block is zeroed out.

[0290] When SBTs are used in inter-frame encoding / decoding CUs, SBT type and SBT position information are transmitted via signals in the bitstream. For example... Figure 36 As shown, there are two SBT types and two SBT locations. For SBT-V (or SBT-H), the TU width (or height) can be equal to half the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. A 2:2 partition is like a binary tree (BT) partition, while a 1:3 / 3:1 partition is like an asymmetric binary tree (ABT) partition. In an ABT partition, only small regions include non-zero residuals. If one dimension of the CU is 8 in the luma samples, a 1:3 / 3:1 partition along that dimension is not allowed. A CU can have a maximum of 8 SBT modes.

[0291] Position-dependent transform core selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TB always uses DCT-2). Two positions in SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal and vertical transforms at each SBT position are... Figure 36The transformations are specified in the code. For example, the horizontal and vertical transformations at SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transformations in both dimensions are set to DCT-2. Therefore, the sub-block transformations jointly specify the TU slice, cbf, and the horizontal and vertical core transformation types of the residual block.

[0292] SBT is not applied to CUs encoded and decoded using intra-frame / inter-frame joint mode.

[0293] 2.27. Adaptive Merge Candidate Reordering Based on Template Matching To improve encoding and decoding efficiency, after constructing the merge candidate list, the order of each merge candidate is adjusted based on the template matching cost. The merge candidates are arranged in the list according to their ascending template matching cost. This is done in subgroups.

[0294] Template matching cost is measured by the sum of absolute differences (SAD) between the current CU's neighboring samples and their corresponding reference samples. If the merge candidate includes bidirectional predicted motion information, then... Figure 37 As shown, the corresponding reference sample points are the average values ​​of the corresponding reference sample points in reference list 0 and reference list 1. Figure 37 The neighboring samples used to calculate SAD are shown. If the Merge candidate includes motion information at the sub-CU level, then... Figure 38 As shown, the corresponding reference sample point is composed of the neighboring sample points of the corresponding reference sub-block. Figure 38 The neighboring samples used to calculate the SAD for sub-CU level motion information are shown.

[0295] like Figure 39 As shown, the sorting process is performed in subgroups. The first three merge candidates are sorted together. The last three merge candidates are sorted together.

[0296] The template size (width of the left template or height of the top template) is 1. The subgroup size is 3.

[0297] 2.28. Adaptive Merge Candidate List We can assume there are 8 merge candidates. We will take the first 5 merge candidates as the first subgroup and the last 3 merge candidates as the second subgroup (i.e., the last subgroup).

[0298] For the encoder, after constructing the Merge candidate list, as follows Figure 40 As shown, some Merge candidates are adaptively reordered in ascending order of Merge candidate cost. Figure 40 The reordering process in the encoder is shown.

[0299] More specifically, the template matching cost is calculated for the merge candidates in all subgroups except the last one; then the merge candidates in their own subgroups except the last one are reordered; finally, the final merge candidate list is obtained.

[0300] For the decoder, after constructing the Merge candidate list, such as Figure 41 As shown, some / no Merge candidates are adaptively reordered in ascending order at the Merge candidate cost. Figure 38 The reordering process in the decoder is illustrated. Figure 41 In this context, the subgroup containing the selected (transmitted via signal) Merge candidate is referred to as the selected subgroup.

[0301] More specifically, if the selected Merge candidate is in the last subgroup, the Merge candidate list construction process is terminated after deriving the selected Merge candidate, no reordering is performed, and the Merge candidate list remains unchanged; otherwise, the process is as follows: After deriving all Merge candidates in the selected subgroup, the Merge candidate list construction process is terminated; the template matching cost for the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained.

[0302] For both the encoder and decoder: The template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template.

[0303] When deriving reference samples for the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.

[0304] The reference samples (RT) for the template used for bidirectional prediction are obtained by using the reference samples of the template in reference list 0 as follows ( ) and reference samples of the template in reference list 1 ( It is derived by weighted averaging.

[0305]

[0306] The weights (8-w) of the reference templates in reference list 0 and the weights (w) of the reference templates in reference list 1 are determined by the BCW indices of the Merge candidates. The BCW indices equal to {0,1,2,3,4} correspond to w equal to {-2,3,4,5,10}, respectively.

[0307] If the Local Illumination Compensation (LIC) flag of the Merge candidate is true, the reference sample points of the template are derived using the LIC method.

[0308] Template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT.

[0309] The template size is 1. This means that the width of the left template and / or the height of the top template is 1.

[0310] If the encoding / decoding mode is MMVD, the Merge candidates used to derive the base Merge candidates are not reordered.

[0311] If the encoding / decoding mode is GPM, the Merge candidates used to derive the one-way prediction candidate list are not reordered.

[0312] 2.29. Geometric segmentation pattern with motion vector difference In the Geometry Segmentation Mode with Motion Vector Difference (GMVD), each geometric segment in the GPM can determine whether GMVD is used. If GMVD is selected for a geometric region, the region's MV is calculated as the sum of the MV of the merged candidates and the MVD. All other processing remains the same as in the GPM.

[0313] Using GMVD, MVD is transmitted as a pair of direction and distance via signals. Nine candidate distances are involved (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD in GMVD is shifted left by 2, just like in MMVD.

[0314] 2.30. Affine MMVD In affine MMVD, affine Merge candidates (called fundamental affine Merge candidates) are selected, and the MV of the control points is further refined by the MVD information transmitted by the signal.

[0315] The MVD information for the MV of all control points is the same in one prediction direction.

[0316] When the starting MV is a bidirectional prediction MV and the two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, while the other reference POC is less than the current image's POC), the MV offset of the list 0 MV component added to the starting MV has the opposite value to the MV offset of the list 1 MV; otherwise, when the starting MV is a bidirectional prediction MV and both lists point to the same side of the current image (i.e., both reference POCs are greater than the current image's POC, or both are less than the current image's POC), the MV offset of the list 0 MV component added to the starting MV has the same value as the MV offset of the list 1 MV.

[0317] 2.31. Adaptive Decoder-Side Motion Vector Refinement (ADMVR) In ECM-2.0, if the selected merge candidate satisfies the DMVR condition, a multi-pass decoder-side motion vector refinement (DMVR) method is applied in the regular merge mode. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16x16 sub-block within the codec block. In the third pass, the motion vectors (MVs) in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF).

[0318] The adaptive decoder-side motion vector refinement method consists of two new Merge modes, which are introduced to refine the motion vectors only in one direction (L0 or L1) of the bidirectional predictions of Merge candidates that satisfy the DMVR condition. A multi-pass DMVR process is applied to the selected Merge candidates to refine the motion vectors; however, in the first pass (i.e., PU level) of the DMVR, either MVD0 or MVD1 is set to zero.

[0319] Similar to the regular Merge pattern, the proposed Merge pattern derives its Merge candidates from spatially adjacent encoded / decoded blocks, TMVPs, non-adjacent blocks, HMVPs, and paired candidates. The difference lies in that only those satisfying the DMVR conditions are added to the candidate list. Both proposed Merge patterns use the same Merge candidate list (i.e., the ADMVR Merge list), and the encoding / decoding of the Merge index is the same as in the regular Merge pattern.

[0320] 2.32. IBC with Template Matching It is proposed to also use template matching with IBC in the IBC Merge pattern and the IBC AMVP pattern.

[0321] Compared to the IBC-TM Merge list used in the regular IBC Merge mode, the IBC-TM Merge list has been modified so that candidates are selected based on a deduplication method, where the motion distance between candidates is the same as in the regular TM Merge mode. The zero-motion satisfaction at the end (which is meaningless with respect to intra-frame encoding / decoding) has been replaced with motion vectors to the left (-W, 0), top (0, -H), and upper left (-W, -H) CUs. Then, if necessary, the list is satisfied with one of the leftmost CUs without deduplication.

[0322] In IBC-TM Merge mode, the selected candidates are refined using template matching methods before the RDO or decoding process. IBC-TM Merge mode competes with the regular IBC Merge mode, and the TM-Merge flag is transmitted via signaling.

[0323] In the IBC-TM AMVP mode, up to three candidates are selected from the IBC Merge list. Each of the three selected candidates is refined using a template matching method and ranked according to the template matching cost they produce. Then, typically only the top two are considered in the motion estimation process.

[0324] Template matching refinement for both IBC-TM Merge and AMVP modes is quite simple because the IBC motion vectors are constrained to integers and, as shown in the example... Figure 42 Within the reference area shown. Figure 42 The IBC reference region, which depends on the current CU position, is shown. Therefore, in IBC-TM Merge mode, all thinning is performed with integer precision, and in IBC-TM AMVP mode, they are performed with either integer precision or 4-pixel precision. In both cases, the motion vectors of the thinning in each thinning step must adhere to the constraints of the reference region.

[0325] 2.33. IBC Merge Mode with Block Vector Difference The IBC Merge mode with block vector difference is shown below.

[0326] The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 ​​pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD direction is two horizontal directions and two vertical directions.

[0327] The base candidates are selected from the top 5 candidates in the reordered IBC Merge list. And based on the SAD cost between the template (the row above and column to the left of the current block) and its reference for each base candidate, all possible MBVD refinement positions (20x4) are reordered. Finally, the top 8 refinement positions with the lowest template SAD cost are reserved as available positions for MBVD index encoding and decoding.

[0328] 2.34. Reconstructing the Reordered IBC (RR-IBC) Figure 43 An example of symmetry is shown in the screen content image.

[0329] Screen content codecs such as Intra-Block Copy (IBC) generate predicted blocks by directly copying previously encoded reference regions from the same image. Figure 43 As shown, symmetry is commonly observed in video content, particularly in text character regions and computer-generated graphics within sequences of screen content. Therefore, screen content encoding / decoding tools that take symmetry into account will effectively compress such video content.

[0330] For screen content video encoding and decoding, a Reconstruction Reordering IBC (RR-IBC) mode is proposed. When applied, samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is derived without flipping. On the decoder side, the reconstructed block is flipped back to recover the original block.

[0331] For blocks encoded and decoded using RR-IBC, two flipping methods are supported: horizontal flipping and vertical flipping. First, for blocks encoded and decoded using IBC AMVP, a syntax flag is signaled to indicate whether the reconstruction has been flipped. If flipped, another flag is further signaled, specifying the flipping type. For IBC Merge, the flipping type is inherited from the neighboring block, with no syntax signaling. Considering horizontal or vertical symmetry, the current block and the reference block are typically horizontally or vertically aligned. Therefore, when applying a horizontal flip, the vertical component of the BV is not signaled and is presumed to be equal to 0. Similarly, when applying a vertical flip, the horizontal component of the BV is not signaled and is presumed to be equal to 0.

[0332] To better utilize the symmetry property, a flip-aware BV adjustment method is applied to refine the block vector candidates. Figure 44A A schematic diagram of BV adjustment for horizontal flipping is shown. Figure 44B A schematic diagram of BV adjustment for vertical flipping is shown. For example, as Figure 44A and Figure 44B As shown, (x nbr , y nbr )and( x cur , y cur ) represent the coordinates of the center sample points of neighboring blocks and the current block, respectively. BV nbr and BV cur These represent the BV of the neighboring block and the current block, respectively. This is relevant when the neighboring block is encoded and decoded using a horizontal flipping method. BV cur The horizontal component does not inherit BV directly from neighboring blocks, but rather through... BV nbr The horizontal component (represented as) BV nbr h Add motion shift to calculate, that is, BV cur h =2( x nbr - x cur ) + BV nbr h Similarly, in the case where adjacent blocks are encoded and decoded using vertical flipping, BV cur The vertical component is through the direction BV nbr The vertical component (represented as) BV nbr v Add motion shift to calculate, that is, BV cur v =2( y nbr - y cur ) + BV nbr v .

[0333] 2.35. Automatic Relocation Block Vector Prediction (AR-BVP) Automatically relocated block vector prediction (AR-BVP) has been introduced into the construction of the IBC Merge / AMVP candidate list.

[0334] Figure 45 An example of how to derive AR-BVP is shown. For example... Figure 45As shown, the guiding block vector BV0,1 associated with the current block B0 points to the reference block B1. If B1 has a BV denoted as BV1,2 pointing to the reference block B2, then BV0,2, given by BV0,2 = BV0,1 + BV1,2, is defined as the AR-BVP guided by BV0,1. Similarly, BV0,n+1 can be derived by the following formula: BV0,n+1 =BV0,n+BVn,n+1 = BV0,1+BV1,2 +…+BVn-1,n +BVn,n+1.

[0335] Figure 46 Five positions in Bn are shown. When deriving BVn,n+1 under the guidance of BV0,n, all five positions of Bn (including the top left (e.g., Figure 46 LT in the middle), top right (e.g., Figure 46 RT in the middle), center (e.g., Figure 46 (Ctr in the middle), bottom left (e.g., Figure 46 (LB in the middle) and bottom right (e.g., Figure 46 The position of RB in the array is checked to find BVn,n+1.

[0336] In our implementation, the initial bootstrap vector BV0,1 is set to an existing BVP that is already in the IBC Merge / AMVP candidate list.

[0337] The AR-BVP candidate is inserted after the HBVP candidate.

[0338] 2.36. Intra-frame template matching Intra-frame template matching prediction (intra-frame TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template most similar to the current template and uses the corresponding block as the prediction block. The encoder then transmits the use of this mode via signaling, and the same prediction operation is performed on the decoder side.

[0339] Figure 47 The intra-frame template matching search region used is shown. The prediction signal is obtained by matching the L-shaped causal proximity of the current block with... Figure 47 It is generated by matching another block in a predefined search region, which consists of the following parts: R1: Current CTU R2: Top left CTU, R3: Above CTU, R4: Left CTU.

[0340] SAD was used as the cost function.

[0341] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.

[0342] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH in" " is a constant that controls the trade-off between gain and complexity. In practice, " "Equals 5."

[0343] Enable intra-frame template matching for CUs with width and height dimensions less than or equal to 64. The maximum CU size for intra-frame template matching is configurable.

[0344] When DIMD is not used in the current CU, the intra-template matching prediction mode is transmitted at the CU level via a dedicated flag.

[0345] 2.37. Non-local lighting compensation A nonlocal illumination compensation (NL-IC) scheme is proposed. Using this method, instead of template samples, samples from previously encoded and decoded CUs are used to derive a linear model for motion compensation of the current block. Specifically, after reconstruction of each inter-frame CU (except for GPM and SbTMVP CUs), a linear model is derived by minimizing the difference between the reconstructed and predicted samples of the block. Then, for both the regular Merge mode and the sub-block Merge mode, up to six NL-IC candidates obtained from the spatially adjacent and non-nearest neighbors of the current CU are inserted and reordered along with existing candidates in the Merge list. The top N candidates with the smallest SAD are retained in the list, with one index transmitted via signaling to indicate which candidate is selected. In this contribution, N remains the same as ECM-10.0, i.e., 10 for the regular Merge mode and 15 for the sub-block Merge mode. Furthermore, the same pattern used to obtain the spatially non-nearest neighbors in the regular Merge is reused to locate the corresponding non-nearest NL-IC candidates. When an IC-ILM candidate is selected, the associated linear model, along with its motion information (e.g., MV, CPMV, reference image index, etc.), is used to generate predicted samples for the CU.

[0346] 2.38. Local lighting compensation with slope adjustment A method for LIC with slope adjustment is proposed, where adjustment parameters are used to update LIC parameters in a slope adjustment similar to CCLM. For AMVP mode, the adjustment parameters are transmitted via signal transmission.

[0347] 2.39. Local lighting compensation with multiple templates A method for LIC with multiple templates (LIC-MTPL) is proposed. In addition to deriving parameters using a template that includes neighboring samples from both the top and left sides of the current block (LIC-TL), two new LIC modes are proposed: one with only the top template (LIC-T) and the other with only the left template (LIC-L). For the AMVP mode, the LIC index is signaled to indicate the template selection.

[0348] 2.40. Geometric segmentation pattern with affine prediction (GPM-affine) GPM is further extended to enable affine motion compensation (AMC). Therefore, GPM segmentation can be predicted via AMC inter-frame prediction, non-AMC inter-frame prediction, or intra-frame prediction. Furthermore, a GPM segmentation predicted via AMC can be combined with another GPM segmentation predicted via AMC, non-AMC, or intra-frame prediction.

[0349] When AMC is applied, similar to the construction of the one-way predictive merge candidate list for GPM in VVC, the one-way predictive affine merge candidate list is constructed from the sub-block-based merge candidate list after discarding sub-TMVP candidates. AMC is performed for GPM segmentation using the control point motion vectors (CPMVs) of the merge candidates in the one-way predictive affine merge candidate list.

[0350] For each GPM segment, a gpm_affine_flag is signaled to indicate whether AMC is applied to the GPM segment. Depending on whether AMC or non-AMC is applied, different arithmetic context models are used to signal merge candidate indices for the GPM segment.

[0351] In the current implementation, AMC is not allowed for GPM-MMVD and GPM-TM.

[0352] 2.41. Regression-based GPM mixture An additional implicit mode for GPM is proposed, in which two integer mixing matrices ( W 0 and W 1) Derived from the template (top row, left column). The blending matrix is ​​modeled as an affine ray function of the sample location (x, y) in the current CU:

[0353] The parameters (a, b, c) are derived from the reference template using the same solver used for CCCM, GLM, or GL-CCCM (MSE minimization). A list of paired candidates is constructed from the regular GPM candidates and reordered using template cost.

[0354] GPM implicit mode is indicated by CU-level flags ( gpm_implicit_flag ) is transmitted via signal. If gpm_ implicit_flag If true, then merge-idx The pair of GPM candidates to be used are encoded and decoded for signal transmission. If gpm_implicit_flag If false, then regular GPM syntax elements are transmitted via signals.

[0355] 3. Problem In the current GPM design, the combination of two sub-segments can be inter-frame prediction and inter-frame prediction, or inter-frame prediction and intra-frame prediction. In the SGPM design, the combination of two sub-segments can be intra-frame prediction and intra-frame prediction. In the IBC-GPM design, the combination of two sub-segments can be IBC and IBC, or IBC and intra-frame prediction. For geometric segmentation, the combination of inter-frame prediction and IBC is currently not allowed. Enabling the combination of inter-frame prediction and IBC for GPM can improve encoding and decoding performance.

[0356] 4. Detailed Solution The detailed embodiments described below should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow sense. Furthermore, these embodiments can be combined in any way.

[0357] In this disclosure, intra-block copying (IBC) may not be limited to current IBC techniques, but can be interpreted as a technique for obtaining a reference (or predicted) block using samples in the current strip / slice / sub-picture / image / other video unit (e.g., CTU line) in addition to conventional intra-prediction methods.

[0358] In the following discussion, IBC can be replaced by other codec tools that rely on encoded / decoded / reconstructed information within the same region, such as palettes or intra-frame template matching.

[0359] Geometric segmentation mode with intra-block copying 1. A method is proposed to use inter-frame prediction and intra-frame block copy (IBC) to obtain prediction / reconstruction samples of video units with at least two sub-segments / sub-blocks. The encoding / decoding mode is represented as GPM-IBC.

[0360] a. In one example, prediction samples for at least one sub-segment are obtained using inter-frame prediction and IBC is used to obtain prediction samples for at least one sub-segment.

[0361] b. In one example, a video unit can be divided into multiple sub-blocks with square or rectangular shapes.

[0362] c. In one example, a video unit can be geometrically divided into multiple sub-segments, such as GPM.

[0363] i. In one example, the GPM partitioning mode used for GPM-IBC can be the same as the GPM-related encoding / decoding tools.

[0364] 1) In one example, the codec tool can refer to regular GPM mode, GPM-TM, GPM-MMVD, GPM-intraframe, GPM-affine, SGPM, IBC-GPM.

[0365] ii. In one example, whether and / or how sub-segments are mixed can be the same as the encoding / decoding tools.

[0366] iii. In one example, regression-based GPM mixtures can be used for GPM-IBC.

[0367] d. In one example, when IBC is used to obtain prediction samples for one sub-segment, inter-frame prediction with unidirectional and / or bidirectional prediction can be used to obtain prediction samples for another sub-segment.

[0368] i. Alternatively, inter-frame prediction with unidirectional prediction is not allowed.

[0369] ii. Alternatively, inter-frame prediction with bidirectional prediction is not permitted.

[0370] e. In one example, when GPM-IBC is used, certain codec tools are not allowed to be used.

[0371] i. In one example, the codec tool may refer to OBMC, template matching-based reordering for GPM partitioning patterns, adaptive hybridization for GPM, ARMC, PROF, DMVR / multi-pass DMVR, BDOF or other decoder-side motion vector derivation methods, or codec tools that use template matching or template matching cost.

[0372] ii. Alternatively, GPM-IBC is not permitted when one or more of the above codecs are used.

[0373] iii. Alternatively, GPM-IBC can be used with one or more codecs.

[0374] 1) In one example, the OBMC can be performed on the inter-frame prediction signal before it is used to mix with the IBC prediction signal.

[0375] 2) In one example, whether and / or how adaptive blending for GPM is applied to GPM-IBC can depend on the codec information.

[0376] a) In one example, codec information can refer to video content.

[0377] b) In one example, when adaptive mixing for GPM is not used with GPM-IBC, the indication of the mixing index may not be transmitted via signaling.

[0378] c) In one example, whether and / or how adaptive blending for GPM is applied to GPM-IBC can be the same as a specific GPM mode, such as regular GPM mode, or GPM-MMVD, or GPM-TM, or GPM-intraframe, or GPM-affine.

[0379] f. In one example, when GPM-IBC is used, one or more IBC tools are not allowed to be used.

[0380] i. In one example, the IBC tool may refer to IBC AMVP mode, or IBC Merge mode, or RR-IBC, or IBC-TM, or IBC-GPM, or IBC-CIIP, or IBC-LIC, or filtered IBC, or bidirectional predictive IBC, or IBC-MBVD, or fractional IBC, or BVD prediction.

[0381] ii. Alternatively, one of the IBC tools mentioned above can be used for GPM-IBC.

[0382] g. In one example, the inter-frame prediction used in GPM-IBC can be the same as the inter-frame prediction used in other GPM modes, such as regular GPM mode, or GPM-TM, or GPM-MMVD, or GPM-intraframe, or GPM-affine.

[0383] i. Alternatively, the inter-frame prediction used in GPM-IBC may be different from the inter-frame prediction used in the GPM mode described above.

[0384] h. In one example, how two sub-segments used in GPM-IBC are mixed can be the same as the mixing method used in other GPM modes, such as regular GPM mode, or GPM-TM, or GPM-MMVD, or GPM-intraframe, or GPM-affine.

[0385] i. Alternatively, the mixing method in GPM-IBC may differ from the mixing method used in the GPM mode described above.

[0386] i. In one example, when GPM-IBC is used, one or more block vectors (BVs) can be used to obtain prediction / reconstruction samples.

[0387] i. In one example, BV can be transmitted by signal, or predefined, or derived.

[0388] 2. In one example, a BV list can be built for GPM-IBC.

[0389] a. In one example, one or more BV candidates can be used to construct the BV list.

[0390] i. In one example, spatially adjacent / non-adjacent neighbor BV candidates can be used.

[0391] ii. In one example, BV candidates from the HBVP table can be used.

[0392] iii. In one example, the constructed BV candidate can be used.

[0393] iv. In one example, paired candidates can be used.

[0394] v. In one example, the default BV candidate can be used.

[0395] vi. In one example, a time-domain BV candidate can be used.

[0396] vii. In one example, an AR-BVP candidate can be used.

[0397] b. In one example, reordering can be used for BV lists.

[0398] i. In one example, reordering could depend on template matching or template matching cost.

[0399] c. In one example, the BV refinement process can be used on one or more BV candidates in the BV list.

[0400] i. In one example, the refinement process can refer to template matching of the IBC.

[0401] ii. In one example, the refinement process can be applied after reordering.

[0402] 1) Alternatively, the refinement process can be applied before reordering.

[0403] iii. In one example, the refinement process can be applied to some specific BV candidates.

[0404] 1) In one example, the refinement process can be applied to the first X BV candidates.

[0405] d. In one example, BVP clustering can be used to construct a list of BVs.

[0406] e. In one example, the BV list can be constructed in the same way as it is for IBC tools.

[0407] i. In one example, the IBC tool may refer to IBC AMVP mode, or IBC Merge mode, or RR-IBC, or IBC-TM, or IBC-GPM, or IBC-CIIP, or IBC-LIC, or filtered IBC, or bidirectional predictive IBC, or IBC-MBVD, or fractional IBC, or BVD prediction.

[0408] ii. Alternatively, a separate BV list construction can be used for GPM-IBC.

[0409] f. In one example, IntraTMP can be used to construct a BV list.

[0410] i. In one example, one or more BV candidates can be generated using IntraTMP.

[0411] g. In one example, it can be shown which BV candidates from the BV list are used to obtain prediction / reconstruction samples in GPM-IBC.

[0412] i. In one example, one or more syntax elements can be used to indicate BV candidates.

[0413] h. In one example, which BV candidates in the BV list are used to obtain prediction / reconstruction samples in GPM-IBC can be predefined or derived.

[0414] 3. In one example, the encoding and decoding information of a video unit using GPM-IBC encoding and decoding can be stored and used by subsequent video units in the current frame and / or video units in subsequent frames.

[0415] a. In one example, the encoding / decoding information may refer to block vectors and / or intra-frame prediction modes and / or motion vectors.

[0416] b. In one example, one or more BVs can be used to build a Merge / AMVP candidate list for subsequent video units, or inserted into a historical BV cache for future reference.

[0417] i. Alternative location: BV is not stored.

[0418] 1) In one example, the motion information for the video unit is set to be equal to the inter-frame mode.

[0419] c. In one example, one or more IPMs may be used to build a list of MPMs for subsequent video units, or stored in an IPM cache, or used for chroma prediction.

[0420] i. In one example, the default IPM can be stored, such as a plane or a DC.

[0421] d. In one example, one or more MVs can be used to build a Merge / AMVP candidate list for subsequent video units, or stored in a motion information cache.

[0422] i. In one example, the MV in a subsegment that has a larger size than other subsegments can be used.

[0423] ii. In one example, the MV in the first / last sub-segment can be used.

[0424] iii. As an alternative, MVs in GPM-IBC are not allowed to be used.

[0425] 1) In one example, the motion information for the video unit is set to be equal to the IBC mode.

[0426] e. In one example, BV / IPM / MV can be stored at the sub-block level (e.g., 4×4 or 8×8).

[0427] i. In one example, the BV / MV of the prediction signal used to generate sub-blocks can be stored.

[0428] ii. In one example, when mixing is used to generate the prediction signal for a sub-block, either BV or MV can be stored.

[0429] f. In one example, the stored BV / IPM / MV can be used for the chromaticity components.

[0430] i. In one example, the BV / IPM / MV stored for the Luminance Video Unit can be used for the Chroma Video Unit.

[0431] ii. In one example, the BV / IPM / MV stored for a chroma video unit can be used for subsequent chroma video units.

[0432] g. In one example, video units encoded and decoded using GPM-IBC can be labeled as SCC blocks, and the labeling information is used in adaptive OBMC control.

[0433] i. Alternatively, video units encoded and decoded using GPM-IBC can be marked as non-SCC blocks, and the marked information is used in adaptive OBMC control.

[0434] ii. Alternatively, whether a video unit encoded or decoded using GPM-IBC is marked as an SCC block can be the same as in other GPM modes.

[0435] 4. In one example, the interaction between GPM-IBC and a specific codec tool can be the same as that between GPM-IBC and GPM-IBC.

[0436] a. In one example, a specific codec tool could refer to LMCS.

[0437] i. In one example, the demapping process can be performed on the IBC prediction signal before it is used to mix with the inter-frame prediction signal.

[0438] ii. In one example, LMCS can be performed when template-matching reordering for the GPM partitioning pattern is used for GPM-IBC.

[0439] b. In one example, a specific codec tool may refer to a loop filter.

[0440] i. In one example, the boundary strength setting of GPM-IBC can be the same as that of GPM-Intraframe.

[0441] 5. In one example, GPM-IBC can be applied to a specific GPM mode.

[0442] a. In one example, the codec tool may refer to GPM-TM and / or GPM-MMVD and / or GPM-Intraframe and / or GPM-Affine.

[0443] b. Alternatively, GPM-IBC may not be applied to one or more GPM modes.

[0444] 6. Whether and / or how to apply GPM-IBC can depend on the codec information, which can refer to: a. Whether specific codec methods, such as GPM or IBC, are allowed. b. Block dimensions and / or block size i. In one example, a block is not allowed to be encoded or decoded using GPM-IBC when the block size (W×H) is greater than or equal to the threshold (T1), where W and H represent the block width and block height, respectively.

[0445] 1) In one example, T1 = 64, or 128, or 256, or 512, or 1024, or 2048, or 4096.

[0446] ii. In one example, a block is not allowed to be encoded or decoded using GPM-IBC when the block size (W×H) is less than or equal to the threshold (T2), where W and H represent the block width and block height, respectively.

[0447] 1) In one example, T2 = 8, or 16, or 32, or 64, or 128, or 256, or 512.

[0448] iii. In one example, a block is not allowed to use GPM-IBC encoding / decoding when W is greater than or equal to the threshold (T3) and / or H is less than or equal to the threshold (T4).

[0449] 1) In one example, T3 = 4 / 8 / 16 / 32 / 64.

[0450] 2) In one example, T4 = 4 / 8 / 16 / 32 / 64.

[0451] iv. In one example, the block is not allowed to use GPM-IBC encoding / decoding when W is less than or equal to the threshold (T5) and / or H is greater than or equal to the threshold (T6).

[0452] 1) In one example, T5 = 8 / 16 / 32 / 64 / 128.

[0453] 2) In one example, T6 = 8 / 16 / 32 / 64 / 128.

[0454] v. In one example, blocks are not allowed to use GPM-IBC encoding / decoding when W / H and / or H / W are greater than or equal to a threshold.

[0455] 1) In one example, the threshold could be equal to 2 / 4 / 8 / 16 / 32.

[0456] vi. In one example, block size can refer to the brightness block size.

[0457] vii. In one example, block size can refer to chroma block size.

[0458] c. Block depth d. Strip / image type and / or segmentation tree type (single tree, dual tree, or local dual tree) e. Temporal layer identifier i. In one example, when the time layer identifier is less than or equal to TidTh1, the block is not allowed to use GPM-IBC encoding and decoding.

[0459] 1) In one example, TidTh1 = 0 / 1 / 2.

[0460] ii. In one example, when the time layer identifier is greater than or equal to TidTh2, the block is not allowed to use GPM-IBC encoding and decoding.

[0461] 1) In one example, TidTh2 = 3 / 4 / 5 / 6.

[0462] f. Block position g. Color format h. Color components i. In one example, GPM-IBC can be applied to all color components.

[0463] ii. In one example, when GPM-IBC is applied to the chromaticity component, it may be different from GPM-IBC applied to the luminance component.

[0464] iii. In one example, whether and / or how GPM-IBC is applied to the first component may depend on whether and / or how GPM-IBC is applied to the second component.

[0465] 1) In one example, the first component may refer to the chromaticity component (e.g., Cb and / or Cr), and the second component may refer to the luminance component (e.g., Y).

[0466] 2) In one example, GPM-IBC can be applied to the first component in the same way as the second component.

[0467] a) Alternatively, the way GPM-IBC is applied to the first component may differ from that of the second component.

[0468] iv. In one example, GPM-IBC can be applied to the luminance component but not to the chrominance component.

[0469] 1) In one example, the luminance component can refer to Y in the YCbCr color space or G in the RGB color space.

[0470] 2) In one example, the chromaticity components can refer to Cb and / or Cr in the YCbCr color space or R and / or B in the RGB color space.

[0471] i. Video content, such as video captured by a camera or video content displayed on a screen. i. In one example, GPM-IBC can be applied to screen content video.

[0472] 7. GPM-IBC indications can be conditionally transmitted via signaling, where conditions may include: a. Block dimensions and / or block size i. In one example, when the block size (W×H) is greater than or equal to the threshold (T7), the indication of GPM-IBC may not be transmitted via signaling, where W and H represent the block width and block height, respectively.

[0473] 1) In one example, T7 = 64, or 128, or 256, or 512, or 1024, or 2048, or 4096.

[0474] ii. In one example, when the block size (W×H) is less than or equal to the threshold (T8), the indication of GPM-IBC may not be transmitted via signaling, where W and H represent the block width and block height, respectively.

[0475] 1) In one example, T8 = 16, or 32, or 64, or 128, or 256, or 512.

[0476] iii. In one example, when W is less than or equal to the threshold (T9) and / or H is greater than or equal to the threshold (T10), the indication of GPM-IBC may not be transmitted via signaling.

[0477] 1) In one example, T9 = 4 / 8 / 16 / 32 / 64.

[0478] 2) In one example, T10 = 4 / 8 / 16 / 32 / 64.

[0479] iv. In one example, when W is greater than or equal to the threshold (T11) and / or H is less than or equal to the threshold (T12), the indication of GPM-IBC may not be transmitted via signaling.

[0480] 1) In one example, T11 = 8 / 16 / 32 / 64 / 128.

[0481] 2) In one example, T12 = 8 / 16 / 32 / 64 / 128.

[0482] v. In one example, when W / H and / or H / W are greater than or equal to a threshold, the indication of GPM-IBC may not be transmitted via signaling.

[0483] 1) In one example, the threshold could be equal to 2 / 4 / 8 / 16 / 32.

[0484] vi. In one example, block size can refer to the brightness block size.

[0485] vii. In one example, block size can refer to chroma block size.

[0486] b. Block depth c. Strip / image type and / or segmentation tree type (single tree, dual tree, or partial dual tree) d. Temporal layer identifier i. In one example, when the Time Layer Identifier (TID) is less than or equal to TidTh3, the indication of GPM-IBC may not be transmitted via signaling.

[0487] 1) In one example, TidTh3 = 0 / 1 / 2.

[0488] ii. In one example, when the Time Layer Identifier (TID) is greater than or equal to TidTh4, the indication of the GPM-IBC may not be transmitted via signaling.

[0489] 1) In one example, TidTh4 = 3 / 4 / 5 / 6.

[0490] e. Block location f. Color format g. Color components.

[0491] h. Video content, such as video captured by a camera or video content displayed on a screen. i. In one example, for video captured by a camera, the GPM-IBC indication may not be transmitted via signal.

[0492] 1) In one example, the GPM-IBC indication can be assumed to be the default value.

[0493] 8. Whether the current block is encoded or decoded in GPM-IBC mode can be transmitted via signals using one or more syntax elements (SE).

[0494] a. In one example, syntax elements may be binarized using fixed-length encoding, rounded unary encoding, unary encoding, EG encoding, or encoded flags.

[0495] b. In one example, syntax elements can be either bypassed or context-encoded.

[0496] i. The context may depend on encoded or decoded information, such as block dimensions and / or block size and / or stripe / picture type and / or information about neighboring blocks (adjacent or non-adjacent) and / or information about other encoding / decoding tools used for the current block and / or information about the temporal layer.

[0497] c. In one example, one or more syntax elements may be transmitted via signaling at the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0498] d. In one example, syntax elements can be encoded and decoded in a predictive manner.

[0499] e. In one example, syntax elements can be conditionally encoded or decoded.

[0500] i. For example, a second SE can be signaled to indicate whether GPM-IBC is used only when the first SE indicates that GPM-IBC is applicable.

[0501] 1) The first SE can be located at the sequence header / image header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0502] 2) The second SE can target blocks.

[0503] ii. For example, a third SE may be signaled to indicate how to perform GPM-IBC only when the second SE indicates the use of GPM-IBC.

[0504] 1) In one example, the third SE can be used to indicate which BV candidate is used.

[0505] f. In one example, when a particular codec tool is used, the SE indicating whether GPM-IBC is used may not be transmitted via signaling.

[0506] i. In one example, the codec tool could refer to GPM-TM, or GPM-MMVD, or GPM-intraframe, or GPM-affine.

[0507] ii. In one example, the SE indicating whether to use GPM-IBC can be presumed to be the default value, such as false.

[0508] iii. Alternatively, when GPM-IBC is used, the SE indicating whether a specific codec tool is used may not be transmitted via signal.

[0509] 1) In one example, the SE indicating whether a particular codec tool is used can be presumed to be a default value, such as false.

[0510] g. In one example, whether and / or how the SE is transmitted via signaling can depend on the specific codec tool.

[0511] i. In one example, the codec tool could refer to GPM - intra-frame.

[0512] 1) In one example, the SE can be conditionally transmitted via signaling, where the condition can refer to whether GPM-intraframe is used.

[0513] 2) In one example, the SE indicating whether GPM-IBC is used may not be transmitted via signaling.

[0514] 3) In one example, the SE indicating which BV is used can be transmitted via signaling together with the SE indicating which intra-prediction mode is used within the GPM-frame.

[0515] 9. The IBC in the methods disclosed above can be replaced by another codec tool.

[0516] a. In one example, the codec tool could refer to IntraTMP.

[0517] 10. It was proposed that intra-frame prediction in GPM can be combined with other GPM tools, such as GPM-MMVD, GPM-TM, or GPM-affine.

[0518] 11. In one example, a syntax element indicating which intra-prediction mode to use within a GPM-frame can be context-coded.

[0519] General aspects 12. In the above examples, a video unit can refer to a color component / sub-picture / strip / piece / code-decode tree unit (CTU) / CTU line / CTU group / code-decode unit (CU) / prediction unit (PU) / transform unit (TU) / code-decode tree block (CTB) / code-decode block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region including more than one sample or pixel.

[0520] 13. Whether and / or how the methods disclosed above can be applied to be transmitted via signaling at the sequence level / picture group level / picture level / strip level / piece group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / piece group header.

[0521] 14. Whether and / or how to apply the above methods may depend on the following information: a. Messages transmitted via signals in DPS / SPS / VPS / PPS / APS / Picture Header / Strip Header / Piece Group Header / Coder / Coder Tree Unit (CTU) / Coder / Coder Unit (CU) / CTU Line / CTU Group / TU / PU Block / Video Coder / Coder Unit; b. Location of CU / PU / TU / block / video codec unit; c. The block dimensions of the current block and / or its neighboring blocks; d. The block shape of the current block and / or the blocks adjacent to the current block; e. The encoding / decoding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode; f. Indication of color format (such as 4:2:0, 4:4:4); g. Encoding / decoding tree structure; h. Strip / panel type and / or image type; i. Color components (e.g., can be applied only to the chromaticity component or the luminance component); j. Temporal layer ID; k. Standard grade / level / tier.

[0522] Figure 48 A flowchart of a method 4800 for video processing according to an embodiment of the present disclosure is shown. Method 4800 is implemented during the conversion between video units of a video and a bitstream of a video.

[0523] At box 4810, for the conversion between video units and video bitstreams, a first target coding / decoding mode is applied to the video units. In this case, the first target coding / decoding mode includes inter-frame prediction and coding / decoding tools, the coding / decoding tools including at least one of the following: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks.

[0524] At box 4820, the conversion is performed based on a first target encoding / decoding mode. In some embodiments, the conversion may include encoding video units into a bitstream. Alternatively, the conversion may include decoding video units from a bitstream.

[0525] Method 4800 enables the application of the first target encoding / decoding mode. Compared to traditional solutions, Method 4800 significantly improves encoding / decoding efficiency and performance.

[0526] In some embodiments, a video unit can be geometrically divided into multiple sub-segments. For example, a video unit can be divided using a geometric segmentation mode (GPM). In some embodiments, the GPM segmentation mode used for the first target codec mode can be the same as the GPM segmentation mode used for a GPM-related codec tool. In some other embodiments, whether and / or how multiple sub-segments are mixed can be the same as the GPM-related codec tool. For example, a GPM-related codec tool can include at least one of the following: a regular GPM mode, GPM-template matching (TM), GPM-Merge mode with motion vector difference (MMVD), GPM-intraframe, GPM-affine, spatial geometric segmentation mode (SGPM), intraTMP-GPM, or IBC-GPM. In some embodiments, regression-based GPM mixing can be used for the first target codec mode.

[0527] In some embodiments, the first target codec mode may be used in conjunction with a target codec tool. For example, the target codec tool may include at least one of the following: Overlapping Block Motion Compensation (OBMC), Template Match-Based Reordering for GPM Partitioning, Adaptive Reordering with Merge Candidates (ARMC), Prediction Refinement with Optical Flow (PROF), Decoder-Side Motion Vector Refinement (DMVR), Multi-Pass DMVR, Bidirectional Optical Flow (BDOF), another decoder-side motion vector derivation method, a codec tool using template matching, or a codec tool using template matching cost. In some embodiments, OBMC may be performed on the inter-frame prediction signal before the inter-frame prediction signal is used to mix with the IBC prediction signal and / or the intraTMP prediction signal. In some other embodiments, whether and / or how adaptive mixing for GPM is applied to the first target codec mode may depend on codec information. For example, codec information may include video content. In some examples, if adaptive mixing for GPM is not used in conjunction with the first target codec mode, an indication of the mixing index may not be transmitted via signaling. In some embodiments, whether and / or how adaptive mixing for GPM is applied to the first target codec mode may be the same as the target GPM mode. For example, the target GPM mode may include at least one of the following: regular GPM mode, GPM-Merge mode with motion vector difference (MMVD), GPM-template matching (TM), GPM-intraframe, and GPM-affine.

[0528] In some embodiments, the block vector (BV) list may be constructed for a first target codec mode. In some embodiments, intra-template matching prediction (intraTMP) may be used to construct the BV list. For example, one or more BV candidates may be generated using intraTMP.

[0529] In some embodiments, the codec information of a video unit encoded using a first target codec mode can be stored and used by subsequent video units, which are at least one of the following: a subsequent video unit in the current frame or a video unit in a subsequent frame. In some embodiments, the codec information may include at least one of the following: block vectors, intra-frame prediction modes (IPM), or motion vectors. In some other embodiments, one or more block vectors (BVs) may be used to construct a merge candidate list and / or an advanced motion vector prediction (AMVP) candidate list for subsequent video units. Inter-frame modes that are non-skipped inter-frame modes or non-merge inter-frame modes can be considered as AMVP modes. AMVP can utilize the motion vectors of blocks that are temporally and / or spatially neighboring blocks to construct a candidate list for the current block.

[0530] In some embodiments, one or more BVs may be inserted into a historical BV cache for future reference. Alternatively, one or more BVs may not be stored. For example, motion information for a video unit may be set to equal the inter-frame mode.

[0531] In some embodiments, one or more intra-prediction modes (IPMs) can be used to construct a list of most probable modes (MPMs) for subsequent video units. In some other embodiments, one or more intra-prediction modes (IPMs) can be stored in an IPM cache. Alternatively, one or more intra-prediction modes (IPMs) can be used for chroma prediction. For example, a default IPM can be stored in an IPM cache. As an example, the default IPM may include at least one of a plane or a DC.

[0532] In some embodiments, one or more motion vectors (MVs) can be used to construct a Merge candidate list and / or an Advanced Motion Vector Prediction (AMVP) candidate list for subsequent video units. Alternatively, one or more motion vectors (MVs) can be stored in a motion information cache. In some embodiments, one or more MVs in a sub-segment can be used. In this case, the sub-segment has a larger size than another sub-segment. In some embodiments, one or more MVs in the first sub-segment can be used. Alternatively, one or more MVs in the last sub-segment can be used. In some other embodiments, MVs in the first target codec mode may not be allowed. For example, the motion information for a video unit may be set to be equal to the IBC mode and / or the intraTMP mode.

[0533] In some embodiments, at least one of the following may be stored at the sub-block level: BV, IPM, MV. For example, the sub-block level may include a 4×4 sub-block level or an 8×8 sub-block level. In some embodiments, at least one of the following may be stored: BV or MV of the prediction signal used to generate the sub-block. In some other embodiments, if the prediction signals used to generate the sub-block are mixed, at least one of the following may be stored: BV or MV.

[0534] In some embodiments, at least one of the following may be used for the chroma component: the stored BV, the stored IPM, or the stored MV. For example, at least one of the following may be used for the chroma video unit: the stored BV for the luma video unit, the stored IPM for the luma video unit, or the stored MV for the luma video unit. Alternatively, at least one of the following may be used for subsequent chroma video units: the stored BV for the chroma video unit, the stored IPM for the chroma video unit, or the stored MV for the chroma video unit.

[0535] In some embodiments, video units encoded using the first target codec mode can be marked as Screen Content Codec (SCC) blocks, and the marked information can be used in adaptive OBMC control. In some other embodiments, video units encoded using the first target codec mode can be marked as non-SCC blocks, and the marked information can be used in adaptive OBMC control. Alternatively, whether a video unit encoded using the first target codec mode is marked as an SCC block can be the same as another GPM mode.

[0536] In some embodiments, the interaction between the first target codec mode and the target codec tool can be the same as the interaction between the GPM-intra-frame target codec tool. In some embodiments, the target codec tool may include a luma mapping with chroma scaling (LMCS). In some examples, the inverse mapping process can be performed on the IBC prediction signal before it is used to mix with the inter-frame prediction signal. Alternatively, the inverse mapping process can be performed on the intraTMP prediction signal before it is used to mix with the inter-frame prediction signal. In some other embodiments, LMCS can be performed if template-match-based reordering for the GPM partitioning mode is used for the first target codec mode. In some embodiments, the target codec tool may include a loop filter. For example, the boundary strength setting of the first target codec mode may be the same as the boundary strength setting of the GPM-intra-frame.

[0537] In some embodiments, whether and / or how the first target codec mode is applied may depend on codec information. For example, codec information may include at least one of the following: whether a codec method, block dimension, block size, block depth, stripe type, image type, segmentation tree type, temporal layer identifier, block location, color format, color components, or video content is allowed. In some embodiments, if the block size is greater than or equal to a first threshold, the block may not be allowed to be encoded using the first target codec mode. In this case, the block size is equal to the block width multiplied by the block height. For example, the first threshold may include one of the following: 64, 128, 256, 512, 1024, 2048, or 4096. In some other embodiments, if the block size is less than or equal to a second threshold, the block may not be allowed to be encoded using the first target codec mode. In this case, the block size is equal to the block width multiplied by the block height. For example, the second threshold may include one of the following: 8, 16, 32, 64, 128, 256, or 512.

[0538] In some embodiments, a block may be denied encoding / decoding using the first target encoding / decoding mode if at least one of the following conditions is met: the block width is greater than or equal to a third threshold, or the block height is less than or equal to a fourth threshold. In some embodiments, the third threshold may include one of the following: 4, 8, 16, 32, or 64. In some embodiments, the fourth threshold may include one of the following: 4, 8, 16, 32, or 64.

[0539] In some embodiments, a block may be denied encoding / decoding using the first target encoding / decoding mode if at least one of the following conditions is met: the block width is less than or equal to a fifth threshold, or the block height is greater than or equal to a sixth threshold. In some embodiments, the fifth threshold may include one of the following: 8, 16, 32, 64, or 128. In some other embodiments, the sixth threshold may include one of the following: 8, 16, 32, 64, or 128.

[0540] In some embodiments, a block may be denied encoding / decoding using the first target encoding / decoding mode if at least one of the following conditions is met: the ratio of the block width to the block height is greater than or equal to a seventh threshold, or the ratio of the block height to the block width is greater than or equal to a seventh threshold. For example, the seventh threshold may include one of the following: 2, 4, 8, 16, or 32.

[0541] In some embodiments, if the temporal layer identifier is less than or equal to a first temporal layer identifier threshold, the block may be denied encoding / decoding using the first target codec mode. For example, the first temporal layer identifier threshold may be equal to one of the following: 0, 1, or 2. In some other embodiments, if the temporal layer identifier is greater than or equal to a second temporal layer identifier threshold, the block may be denied encoding / decoding using the first target codec mode. For example, the second temporal layer identifier threshold may be equal to one of the following: 3, 4, 5, or 6.

[0542] In some embodiments, the video content may include at least one of the following: video captured by a camera or screen content video. For example, a first target codec mode may be applied to screen content video.

[0543] In some embodiments, an indication of a first target codec mode may be transmitted via signaling if certain conditions are met. For example, the conditions may include at least one of the following: block dimension, block size, block depth, stripe type, picture type, segmentation tree type, temporal layer identifier, block location, color format, color components, or video content. In some embodiments, if the block size is greater than or equal to an eighth threshold, the indication of the first target codec mode may not be transmitted via signaling. In this case, the block size is equal to the block width multiplied by the block height. For example, the eighth threshold may include one of the following: 64, 128, 256, 512, 1024, 2048, or 4096. In some other embodiments, if the block size is less than or equal to a ninth threshold, the indication of the first target codec mode may not be transmitted via signaling. In this case, the block size is equal to the block width multiplied by the block height. For example, the ninth threshold may include one of the following: 16, 32, 64, 128, 256, or 512.

[0544] In some embodiments, the indication of the first target encoding / decoding mode may not be transmitted via signaling if at least one of the following conditions is met: the block width of the block is less than or equal to a tenth threshold, or the block height of the block is greater than or equal to an eleventh threshold. In some embodiments, the tenth threshold may include one of the following: 4, 8, 16, 32, or 64. In some embodiments, the eleventh threshold may include one of the following: 4, 8, 16, 32, or 64.

[0545] In some embodiments, the indication of the first target encoding / decoding mode may not be transmitted via signaling if at least one of the following conditions is met: the block width of the block is greater than or equal to a twelfth threshold, or the block height of the block is less than or equal to a thirteenth threshold. In some embodiments, the twelfth threshold may include one of the following: 8, 16, 32, 64, or 128. In some embodiments, the thirteenth threshold may include one of the following: 8, 16, 32, 64, or 128.

[0546] In some embodiments, the indication of the first target encoding / decoding mode may not be transmitted via signaling if at least one of the following conditions is met: the ratio of the block width to the block height is greater than or equal to a fourteenth threshold, or the ratio of the block height to the block width is greater than or equal to a fourteenth threshold. For example, the fourteenth threshold may include one of the following: 2, 4, 8, 16, or 32.

[0547] In some embodiments, if the time-domain layer identifier is less than or equal to a third time-domain layer identifier threshold, the indication of the first target codec mode may not be transmitted via signaling. For example, the third time-domain layer identifier threshold may be equal to one of the following: 0, 1, or 2. In some other embodiments, if the time-domain layer identifier is greater than or equal to a fourth time-domain layer identifier threshold, the indication of the first target codec mode may not be transmitted via signaling. For example, the fourth time-domain layer identifier threshold may be equal to one of the following: 3, 4, 5, or 6.

[0548] In some embodiments, the video content may include at least one of the following: video captured by a camera or video of screen content. In some examples, for video captured by a camera, an indication of the first target codec mode may not be transmitted via signaling. For example, the indication of the first target codec mode may be presumed to be a default value.

[0549] In some embodiments, whether the current block is encoded using a first target codec mode can be signaled by using one or more syntax elements (SEs). In some embodiments, whether and / or how one or more syntax elements are signaled can depend on the target codec tool. As an example, the target codec tool may include GPM-intra. In some embodiments, one or more syntax elements may be signaled if a condition is met. In this case, the condition may include whether GPM-intra is used. In some other embodiments, syntax elements indicating whether the first target codec mode can be used are not signaled. In some embodiments, syntax elements indicating which BV is used may be signaled together with syntax elements indicating which intra-prediction mode is used in GPM-intra. In some embodiments, syntax elements indicating which intra-prediction mode is used in GPM-intra can be context-coded.

[0550] In some embodiments, a video unit may include at least one of the following: color components, sub-pictures, strips, slices, codec tree units (CTUs), CTU rows, CTU groups, codec units (CUs), prediction units (PUs), transform units (TUs), codec tree blocks (CTBs), codec blocks (CBs), prediction blocks (PBs), transform blocks (TBs), blocks, sub-blocks of blocks, sub-regions within blocks, and regions including more than one sample point or pixel.

[0551] In some embodiments, an indication of whether and / or how to apply the first target codec mode may be indicated at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level. In some embodiments, an indication of whether and / or how to apply the first target codec mode may be indicated at one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.

[0552] In some embodiments, whether and / or how a first target codec mode is applied may depend on at least one of the following: messages transmitted via signaling in one of the following: DPS, SPS, VPS, PPS, APS, picture header, stripe header, slice group header, codec tree unit (CTU), codec unit (CU), CTU row, CTU group, TU, PU block or video codec unit, CU position, PU position, TU position, block position, video codec unit position, block dimension of the current block, block dimension of the current block's neighboring blocks, block shape of the current block, block shape of the current block's neighboring blocks, codec mode of the block, color format indication, codec tree structure, stripe group type, slice group type, picture type, color components, temporal layer ID, standard grade, standard level, or standard layer. In some embodiments, the codec mode of the block may include at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, non-IBC sub-block mode, intraTMP inter-frame mode, non-intraTMP inter-frame mode, or non-intraTMP sub-block mode. In some embodiments, the color format specification may include 4:2:0 or 4:4:4. In some other embodiments, color components may be applied to one of the following: chromaticity components or luminance components.

[0553] According to another embodiment of this disclosure, a non-transitory computer-readable recording medium is provided. This non-transitory computer-readable recording medium stores a bitstream of video generated by a method performed by an apparatus for video processing. The method includes: applying a first target encoding / decoding mode to video units of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; and generating a bitstream based on the first target encoding / decoding mode.

[0554] According to further embodiments of this disclosure, a method for storing a bitstream of video is provided. The method includes: applying a first target encoding / decoding mode to video units of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; generating a bitstream based on the first target encoding / decoding mode; and storing the bitstream in a non-transitory computer-readable recording medium.

[0555] The embodiments of this disclosure can be described according to the following entries, and their features can be combined in any reasonable manner.

[0556] Item 1. A method for video processing, comprising: a conversion between a video unit of a video and a bitstream of the video; applying a first target encoding / decoding mode to the video unit, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copy (IBC) or intra-template matching prediction (intraTMP), and the video unit includes a plurality of sub-segments or a plurality of sub-blocks; and performing the conversion based on the first target encoding / decoding mode.

[0557] Item 2. The method according to Item 1, wherein the video unit is geometrically divided into the plurality of sub-segments.

[0558] Item 3. The method according to Item 2, wherein the video unit is divided using a geometric segmentation pattern (GPM).

[0559] Item 4. The method according to Item 3, wherein the GPM partitioning mode used for the first target encoding / decoding mode is the same as the GPM partitioning mode used for GPM-related encoding / decoding tools.

[0560] Item 5. The method according to Item 3, wherein whether and / or how the plurality of sub-segments are mixed is the same as that of the GPM-related codec tools.

[0561] Item 6. The method according to Item 4 or 5, wherein the GPM-related encoding / decoding tool includes at least one of the following: regular GPM mode, GPM-template matching (TM), GPM-Merge mode with motion vector difference (MMVD), GPM-intraframe, GPM-affine, spatial geometry partitioning mode (SGPM), intraTMP-GPM, or IBC-GPM.

[0562] Item 7. The method according to Item 3, wherein regression-based GPM mixing is used for the first target encoding / decoding mode.

[0563] Item 8. The method according to Item 1, wherein the first target codec mode is used together with the target codec tool.

[0564] Item 9. The method according to Item 8, wherein the target encoding / decoding tool comprises at least one of the following: Overlapping Block Motion Compensation (OBMC), Template Matching-Based Reordering for GPM Partitioning Modes, Adaptive Reordering for Merge Candidates (ARMC), Predictive Refinement Using Optical Flow (PROF), Decoder-Side Motion Vector Refinement (DMVR), Multi-Pass DMVR, Bidirectional Optical Flow (BDOF), Another Decoder-Side Motion Vector Derivation Method, Encoding / Decoding Tool Using Template Matching, or Encoding / Decoding Tool Using Template Matching Cost.

[0565] Item 10. The method according to Item 9, wherein the OBMC is performed on the inter-frame prediction signal before the inter-frame prediction signal is used to mix with the IBC prediction signal and / or the intraTMP prediction signal.

[0566] Item 11. The method according to Item 9, wherein whether and / or how adaptive mixing for GPM is applied to the first target codec mode depends on the codec information.

[0567] Item 12. The method according to Item 11, wherein the encoding / decoding information includes video content.

[0568] Item 13. The method according to Item 11, wherein if the adaptive mixing for GPM is not used together with the first target codec mode, the indication of the mixing index is not transmitted via signaling.

[0569] Item 14. The method according to Item 11, wherein whether and / or how the adaptive mixing for GPM is applied to the first target codec mode is the same as the target GPM mode.

[0570] Item 15. The method according to Item 14, wherein the target GPM mode includes at least one of the following: regular GPM mode, GPM-Merge mode with motion vector difference (MMVD), GPM-template matching (TM), GPM-intraframe, and GPM-affine.

[0571] Item 16. The method according to Item 1, wherein a block vector (BV) list is constructed for the first target encoding / decoding mode.

[0572] Item 17. The method according to Item 16, wherein intra-template matching prediction (intraTMP) is used to construct the BV list.

[0573] Item 18. The method according to Item 17, wherein one or more BV candidates are generated using the intraTMP.

[0574] Item 19. The method according to Item 1, wherein the encoding and decoding information of the video unit encoded and decoded using the first target encoding and decoding mode is stored and used by a subsequent video unit, the subsequent video unit being at least one of the following: a subsequent video unit in the current frame, or a video unit in a subsequent frame.

[0575] Item 20. The method according to Item 19, wherein the encoding / decoding information includes at least one of the following: block vector, intra-frame prediction mode (IPM), or motion vector.

[0576] Item 21. According to the method of Item 19, one or more BVs are used to construct a Merge candidate list and / or an Advanced Motion Vector Prediction (AMVP) candidate list for the subsequent video units, or one or more BVs are inserted into a historical BV cache for future reference.

[0577] Item 22. The method described in Item 19, wherein one or more BVs are not stored.

[0578] Item 23. The method according to Item 22, wherein the motion information for the video unit is set to be equal to the inter-frame mode.

[0579] Item 24. The method according to Item 19, wherein one or more intra-frame prediction modes (IPMs) are used to construct a list of most probable modes (MPMs) for the subsequent video unit, or one or more intra-frame prediction modes (IPMs) are stored in an IPM cache, or one or more intra-frame prediction modes (IPMs) are used for chroma prediction.

[0580] Item 25. The method according to Item 24, wherein a default IPM is stored in the IPM cache, wherein the default IPM includes at least one of a plane or a DC.

[0581] Item 26. The method according to Item 19, wherein one or more motion vectors (MVs) are used to construct a Merge candidate list and / or an Advanced Motion Vector Prediction (AMVP) candidate list for the subsequent video units, or wherein one or more motion vectors (MVs) are stored in a motion information cache.

[0582] Item 27. The method according to Item 26, wherein one or more MVs in a sub-segment are used, wherein the sub-segment has a larger size than another sub-segment.

[0583] Item 28. The method according to Item 26, wherein one or more MVs in the first sub-segment are used, or wherein one or more MVs in the last sub-segment are used.

[0584] Item 29. According to the method described in Item 19, none of the MVs in the first target codec mode are allowed to be used.

[0585] Item 30. The method according to Item 29, wherein the motion information for the video unit is set to be equal to IBC mode and / or intraTMP mode.

[0586] Item 31. The method according to Item 19, wherein at least one of the following is stored at the sub-block level: BV, IPM, MV.

[0587] Item 32. The method according to Item 31, wherein the sub-block level includes a 4×4 sub-block level or an 8×8 sub-block level.

[0588] Item 33. The method according to Item 32, wherein at least one of the following is stored: BV used to generate the prediction signal for the sub-block or MV used to generate the prediction signal for the sub-block.

[0589] Item 34. The method according to Item 32, wherein if the mixing is used to generate the prediction signal for the sub-block, at least one of the following is stored: BV or MV.

[0590] Item 35. The method according to Item 19, wherein at least one of the following is used for the chromaticity component: the stored BV, the stored IPM, or the stored MV.

[0591] Item 36. The method according to Item 35, wherein at least one of the following is used for the chroma video unit: BV stored for the luma video unit, IPM stored for the luma video unit, or MV stored for the luma video unit.

[0592] Item 37. The method according to Item 35, wherein at least one of the following is used for a subsequent chroma video unit: a BV stored for the chroma video unit, an IPM stored for the chroma video unit, or a MV stored for the chroma video unit.

[0593] Item 38. The method according to Item 19, wherein the video unit encoded using the first target encoding / decoding mode is marked as a screen content encoding / decoding (SCC) block, and the marked information is used in adaptive OBMC control.

[0594] Item 39. The method according to Item 19, wherein the video unit encoded using the first target encoding / decoding mode is marked as a non-SCC block, and the marked information is used in adaptive OBMC control.

[0595] Item 40. The method according to Item 19, wherein the video unit encoded and decoded using the first target encoding / decoding mode is marked as an SCC block in the same way as another GPM mode.

[0596] Item 41. The method according to Item 1, wherein the interaction between the first target codec mode and the target codec tool is the same as the interaction between GPM-intraframe and the target codec tool.

[0597] Item 42. The method according to Item 41, wherein the target codec tool includes a luminance map with chroma scaling (LMCS).

[0598] Item 43. The method according to Item 42, wherein the demapping process is performed on the IBC prediction signal before the IBC prediction signal is used to mix with the inter-frame prediction signal, or wherein the demapping process is performed on the intraTMP prediction signal before the intraTMP prediction signal is used to mix with the inter-frame prediction signal.

[0599] Item 44. The method according to Item 42, wherein the LMCS is performed if template-matching-based reordering for the GPM partitioning mode is used for the first target encoding / decoding mode.

[0600] Item 45. The method according to Item 41, wherein the target encoding / decoding tool includes a loop filter.

[0601] Item 46. The method according to Item 45, wherein the boundary strength setting of the first target encoding / decoding mode is the same as the boundary strength setting within the GPM-frame.

[0602] Item 47. The method according to Item 1, wherein whether and / or how the first target codec mode is applied depends on codec information, wherein the codec information includes at least one of the following: whether a codec method, block dimension, block size, block depth, stripe type, picture type, segmentation tree type, temporal layer identifier, block location, color format, color components, or video content is allowed.

[0603] Item 48. The method according to Item 47, wherein if the block size is greater than or equal to a first threshold, the block is not allowed to be encoded or decoded using the first target encoding / decoding mode, wherein the block size is equal to the block width multiplied by the block height.

[0604] Item 49. The method according to Item 48, wherein the first threshold comprises one of the following: 64, 128, 256, 512, 1024, 2048, or 4096.

[0605] Item 50. The method according to Item 47, wherein if the block size is less than or equal to a second threshold, the block is not allowed to be encoded or decoded using the first target encoding / decoding mode, wherein the block size is equal to the block width multiplied by the block height.

[0606] Item 51. The method according to Item 50, wherein the second threshold comprises one of the following: 8, 16, 32, 64, 128, 256 or 512.

[0607] Item 52. The method according to Item 47, wherein a block is not permitted to be encoded or decoded using the first target encoding / decoding mode if at least one of the following conditions is met: the block width of the block is greater than or equal to a third threshold, or the block height of the block is less than or equal to a fourth threshold.

[0608] Item 53. The method according to Item 52, wherein the third threshold includes one of the following: 4, 8, 16, 32 or 64.

[0609] Item 54. The method according to Item 52, wherein the fourth threshold comprises one of the following: 4, 8, 16, 32 or 64.

[0610] Item 55. The method according to Item 47, wherein a block is not allowed to be encoded or decoded using the first target encoding / decoding mode if at least one of the following conditions is met: the block width of the block is less than or equal to a fifth threshold, or the block height of the block is greater than or equal to a sixth threshold.

[0611] Item 56. The method according to Item 55, wherein the fifth threshold comprises one of the following: 8, 16, 32, 64 or 128.

[0612] Item 57. The method according to Item 55, wherein the sixth threshold includes one of the following: 8, 16, 32, 64 or 128.

[0613] Item 58. The method according to Item 47, wherein a block is not allowed to be encoded or decoded using the first target encoding / decoding mode if at least one of the following conditions is met: the ratio of the block width to the block height of the block is greater than or equal to a seventh threshold, or the ratio of the block height to the block width is greater than or equal to the seventh threshold.

[0614] Item 59. The method according to Item 58, wherein the seventh threshold comprises one of the following: 2, 4, 8, 16 or 32.

[0615] Item 60. The method according to Item 47, wherein if the temporal layer identifier is less than or equal to a first temporal layer identifier threshold, the block is not allowed to be encoded or decoded using the first target encoding / decoding mode.

[0616] Item 61. The method according to Item 60, wherein the first time-domain layer identifier threshold is equal to one of the following: 0, 1 or 2.

[0617] Item 62. The method according to Item 47, wherein if the time-domain layer identifier is greater than or equal to the second time-domain layer identifier threshold, the block is not allowed to be encoded or decoded using the first target encoding / decoding mode.

[0618] Item 63. The method according to Item 62, wherein the second time-domain layer identifier threshold is equal to one of the following: 3, 4, 5 or 6.

[0619] Item 64. The method according to Item 47, wherein the video content includes at least one of the following: video captured by a camera or video of screen content.

[0620] Item 65. The method according to Item 64, wherein the first target encoding / decoding mode is applied to the screen content video.

[0621] Item 66. The method according to Item 1, wherein an indication of the first target encoding / decoding mode is transmitted via signaling if conditions are met, wherein the conditions include at least one of the following: block dimension, block size, block depth, stripe type, picture type, segmentation tree type, temporal layer identifier, block location, color format, color components, or video content.

[0622] Item 67. The method according to Item 66, wherein if the block size is greater than or equal to an eighth threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling, wherein the block size is equal to the block width multiplied by the block height.

[0623] Item 68. The method according to Item 67, wherein the eighth threshold comprises one of the following: 64, 128, 256, 512, 1024, 2048 or 4096.

[0624] Item 69. The method according to Item 66, wherein if the block size is less than or equal to a ninth threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling, wherein the block size is equal to the block width multiplied by the block height.

[0625] Item 70. The method according to Item 69, wherein the ninth threshold comprises one of the following: 16, 32, 64, 128, 256, or 512.

[0626] Item 71. The method according to Item 66, wherein the indication of the first target encoding / decoding mode is not transmitted by signal if at least one of the following conditions is met: the block width of the block is less than or equal to the tenth threshold, or the block height of the block is greater than or equal to the eleventh threshold.

[0627] Item 72. The method according to Item 71, wherein the tenth threshold includes one of the following: 4, 8, 16, 32 or 64.

[0628] Item 73. The method according to Item 71, wherein the eleventh threshold includes one of the following: 4, 8, 16, 32 or 64.

[0629] Item 74. The method according to Item 66, wherein the indication of the first target encoding / decoding mode is not transmitted by signal if at least one of the following conditions is met: the block width of the block is greater than or equal to the twelfth threshold, or the block height of the block is less than or equal to the thirteenth threshold.

[0630] Item 75. The method according to Item 74, wherein the twelfth threshold comprises one of the following: 8, 16, 32, 64, or 128.

[0631] Item 76. The method according to Item 74, wherein the thirteenth threshold includes one of the following: 8, 16, 32, 64 or 128.

[0632] Item 77. The method according to Item 66, wherein the indication of the first target encoding / decoding mode is not transmitted by signal if at least one of the following conditions is met: the ratio of the block width to the block height of the block is greater than or equal to a fourteenth threshold, or the ratio of the block height to the block width is greater than or equal to the fourteenth threshold.

[0633] Item 78. The method according to Item 77, wherein the fourteenth threshold includes one of the following: 2, 4, 8, 16 or 32.

[0634] Item 79. The method according to Item 66, wherein if the time-domain layer identifier is less than or equal to a third time-domain layer identifier threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling.

[0635] Item 80. The method according to Item 79, wherein the third time-domain layer identifier threshold is equal to one of the following: 0, 1 or 2.

[0636] Item 81. The method according to Item 66, wherein if the time-domain layer identifier is greater than or equal to the fourth time-domain layer identifier threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling.

[0637] Item 82. The method according to Item 81, wherein the fourth time-domain layer identifier threshold is equal to one of the following: 3, 4, 5 or 6.

[0638] Item 83. The method according to Item 66, wherein the video content includes at least one of the following: video captured by a camera or video of screen content.

[0639] Item 84. The method according to Item 83, wherein, for video captured by the camera, the indication of the first target encoding / decoding mode is not transmitted via signal.

[0640] Item 85. The method according to Item 84, wherein the indication of the first target encoding / decoding mode is presumed to be a default value.

[0641] Item 86. The method according to Item 1, wherein whether the current block is encoded or decoded using the first target encoding / decoding mode is transmitted via signaling using one or more syntax elements (SE).

[0642] Item 87. The method according to Item 86, wherein whether and / or how the one or more syntax elements are transmitted by signal depends on the target codec tool.

[0643] Item 88. The method according to Item 87, wherein the target encoding / decoding tool includes GPM-intraframe.

[0644] Item 89. The method according to Item 88, wherein the one or more syntax elements are transmitted via signaling if conditions are met, wherein the conditions include whether the GPM-intraframe is used.

[0645] Item 90. The method according to Item 88, wherein the syntax element indicating whether the first target encoding / decoding mode is used is not transmitted via signal.

[0646] Item 91. The method according to Item 88, wherein a syntax element indicating which BV is used is transmitted via signaling together with a syntax element indicating which intra-prediction mode is used within the GPM-frame.

[0647] Item 92. The method according to Item 1, wherein the syntax element indicating which intra-prediction mode is used in the GPM-frame is context-coded.

[0648] Item 93. The method according to any one of items 1-92, wherein the video unit comprises at least one of the following: color component, sub-picture, strip, slice, codec tree unit (CTU), CTU row, CTU group, codec unit (CU), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec block (CB), prediction block (PB), transform block (TB), block, sub-block of block, sub-region within block, region comprising more than one sample point or pixel.

[0649] Item 94. The method according to any one of items 1-92, wherein an indication of whether and / or how the first target encoding / decoding mode is applied is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.

[0650] Item 95. The method according to any one of items 1-92, wherein an indication of whether and / or how to apply the first target encoding / decoding mode is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header or slice header.

[0651] Item 96. The method according to any one of items 1-92, wherein whether and / or how the first target codec mode is applied depends on at least one of the following information: a message transmitted via signaling in one of the following: DPS, SPS, VPS, PPS, APS, picture header, stripe header, slice group header, codec tree unit (CTU), codec unit (CU), CTU row, CTU group, TU, PU block or video codec unit, the position of CU, the position of PU, the position of TU, ​​the position of block, the position of video codec unit, the block dimension of the current block, the block dimension of the neighboring blocks of the current block, the block shape of the current block, the block shape of the neighboring blocks of the current block, the codec mode of the block, an indication of the color format, the codec tree structure, stripe group type, slice group type, picture type, color components, the ID of the temporal layer, the standard grade, the standard level, or the standard layer.

[0652] Item 97. The method according to Item 96, wherein the encoding / decoding mode of the block includes at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, non-IBC sub-block mode, intraTMP inter-frame mode, non-intraTMP inter-frame mode, or non-intraTMP sub-block mode.

[0653] Item 98. The method according to Item 96, wherein the indication of the color format includes 4:2:0 or 4:4:4.

[0654] Item 99. The method according to Item 96, wherein the color component is applied to one of: a chromaticity component or a luminance component.

[0655] Item 100. The method according to any one of items 1-99, wherein the conversion includes encoding the video unit into the bitstream.

[0656] Item 101. The method according to any one of items 1-99, wherein the conversion includes decoding the video unit from the bitstream.

[0657] Item 102. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of items 1-101.

[0658] Item 103. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of items 1-101.

[0659] Item 104. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method comprises: applying a first target encoding / decoding mode to video units of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes a plurality of sub-segments or a plurality of sub-blocks; and generating the bitstream based on the first target encoding / decoding mode.

[0660] Item 105. A method for storing a bitstream of video, comprising: applying a first target encoding / decoding mode to video units of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes a plurality of sub-segments or a plurality of sub-blocks; generating the bitstream based on the first target encoding / decoding mode; and storing the bitstream in a non-transitory computer-readable recording medium.

[0661] Example device Figure 49A block diagram of a computing device 4900 in which various embodiments of the present disclosure may be implemented is shown. The computing device 4900 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).

[0662] It should be understood that, Figure 49 The computing device 4900 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0663] like Figure 49 As shown, computing device 4900 includes general-purpose computing device 4900. Computing device 4900 may include at least one or more processors or processing units 4910, memory 4920, storage unit 4930, one or more communication units 4940, one or more input devices 4950, and one or more output devices 4960.

[0664] In some embodiments, the computing device 4900 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server provided by a service provider, a large computing device, etc. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, and includes accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 4900 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0665] Processing unit 4910 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 4920. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 4900. Processing unit 4910 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0666] Computing device 4900 typically includes various computer storage media. Such media can be any media accessible by computing device 4900, including but not limited to volatile and non-volatile media, or removable and non-removable media. Memory 4920 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory) or any combination thereof. Storage cell 4930 can be any removable or non-removable media and may include machine-readable media, such as memory, flash drives, disks, or other media that can be used to store information and / or data and can be accessed within computing device 4900.

[0667] The computing device 4900 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although in Figure 49 Not shown, but may provide disk drives for reading from and / or writing to removable non-volatile disks, and optical disc drives for reading from and / or writing to removable non-volatile optical discs. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.

[0668] Communication unit 4940 communicates with another computing device via a communication medium. Furthermore, the functionality of the components in computing device 4900 can be implemented by a single computing cluster or by multiple computing machines communicating via communication connections. Therefore, computing device 4900 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.

[0669] Input device 4950 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, etc. Output device 4960 can be one or more of various output devices, such as a monitor, speaker, printer, etc. With the aid of communication unit 4940, computing device 4900 can also communicate with one or more external devices (not shown), such as storage devices and display devices. Computing device 4900 can also communicate with one or more devices that enable a user to interact with computing device 4900, or any device that enables computing device 4900 to communicate with one or more other computing devices (e.g., network card, modem, etc.), if needed. Such communication can be performed via an input / output (I / O) interface (not shown).

[0670] In some embodiments, some or all components of computing device 4900 may be arranged in a cloud computing architecture rather than integrated into a single device. In a cloud computing architecture, components may be provided remotely and may work together to achieve the functionality described herein. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (WAN), such as the Internet, using suitable protocols. For example, a cloud computing provider provides applications accessible via a web browser or any other computing component over a WAN. The software or components of the cloud computing architecture, along with corresponding data, may be stored on servers at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations in remote data centers. Cloud computing infrastructure may provide services through shared data centers, although they may appear as a single access point for the user. Therefore, cloud computing architectures can be used to provide the components and functionality described herein from service providers at remote locations. Alternatively, they may be provided from conventional servers or installed directly or otherwise on client devices.

[0671] In embodiments of this disclosure, computing device 4900 may be used to implement video encoding / decoding. Memory 4920 may include one or more video codec modules 4925 having one or more program instructions. These modules can be accessed and executed by processing unit 4910 to perform the functions of the various embodiments described herein.

[0672] In an example embodiment of performing video encoding, input device 4950 may receive video data as input 4970 to be encoded. The video data may be processed, for example, by video codec module 4925 to generate an encoded bitstream. The encoded bitstream may be provided as output 4980 via output device 4960.

[0673] In an example embodiment of performing video decoding, input device 4950 may receive an encoded bitstream as input 4970. The encoded bitstream may be processed, for example, by a video codec module 4925 to generate decoded video data. The decoded video data may be provided as output 4980 via output device 4960.

[0674] While this disclosure has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of this application as defined by the appended claims. These changes are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of this application is not intended to be limiting.

Claims

1. A method for video processing, comprising: For the conversion between video units and the bitstream of the video, a first target encoding / decoding mode is applied to the video unit, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of the following: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; and The conversion is performed based on the first target encoding / decoding mode.

2. The method according to claim 1, wherein the video unit is geometrically divided into the plurality of sub-segments.

3. The method of claim 2, wherein the video unit is divided using a geometric segmentation pattern (GPM).

4. The method of claim 3, wherein the GPM partitioning mode used for the first target encoding / decoding mode is the same as the GPM partitioning mode used for GPM-related encoding / decoding tools.

5. The method of claim 3, wherein whether and / or how the plurality of sub-segments are mixed is the same as the GPM-related encoding / decoding tool.

6. The method according to claim 4 or 5, wherein the GPM-related encoding / decoding tool comprises at least one of the following: regular GPM mode, GPM-template matching (TM), GPM-Merge mode with motion vector difference (MMVD), GPM-intraframe, GPM-affine, spatial geometry partitioning mode (SGPM), intraTMP-GPM, or IBC-GPM.

7. The method of claim 3, wherein regression-based GPM mixing is used in the first target encoding / decoding mode.

8. The method of claim 1, wherein the first target codec mode is used together with the target codec tool.

9. The method of claim 8, wherein the target encoding / decoding tool comprises at least one of the following: Overlapping Block Motion Compensation (OBMC), Template Matching-Based Reordering for GPM Partitioning Modes, Merge Candidate Adaptive Reordering (ARMC), Predictive Refinement Using Optical Flow (PROF), Decoder-Side Motion Vector Refinement (DMVR), Multi-Pass DMVR, Bidirectional Optical Flow (BDOF), another decoder-side motion vector derivation method, an encoding / decoding tool using template matching, or an encoding / decoding tool using template matching cost.

10. The method of claim 9, wherein the OBMC is performed on the inter-frame prediction signal before the inter-frame prediction signal is used to mix with the IBC prediction signal and / or the intraTMP prediction signal.

11. The method of claim 9, wherein whether and / or how adaptive mixing for GPM is applied to the first target codec mode depends on the codec information.

12. The method of claim 11, wherein the encoding / decoding information includes video content.

13. The method of claim 11, wherein if the adaptive mixing for GPM is not used in conjunction with the first target codec mode, the indication of the mixing index is not transmitted via signaling.

14. The method of claim 11, wherein whether and / or how the adaptive mixing for GPM is applied to the first target codec mode is the same as the target GPM mode.

15. The method of claim 14, wherein the target GPM mode comprises at least one of the following: a regular GPM mode, a GPM-Merge mode with motion vector difference (MMVD), a GPM-template matching (TM) mode, a GPM-intraframe mode, or a GPM-affine mode.

16. The method of claim 1, wherein a block vector (BV) list is constructed for the first target encoding / decoding mode.

17. The method of claim 16, wherein intra-template matching prediction (intraTMP) is used to construct the BV list.

18. The method of claim 17, wherein one or more BV candidates are generated using the intraTMP.

19. The method of claim 1, wherein the encoding and decoding information of the video unit encoded and decoded using the first target encoding and decoding mode is stored and used by subsequent video units, the subsequent video units being at least one of the following: a subsequent video unit in the current image, or a video unit in a subsequent image.

20. The method of claim 19, wherein the encoding / decoding information includes at least one of the following: block vector, intra-frame prediction mode (IPM), or motion vector.

21. The method of claim 19, wherein one or more BVs are used to construct a Merge candidate list and / or an Advanced Motion Vector Prediction (AMVP) candidate list for the subsequent video units, or One or more BVs are inserted into the historical BV cache for future reference.

22. The method of claim 19, wherein one or more BVs are not stored.

23. The method of claim 22, wherein the motion information for the video unit is set to be equal to the inter-frame mode.

24. The method of claim 19, wherein one or more intra-frame prediction modes (IPMs) are used to construct a most probable mode (MPM) list for the subsequent video unit, or One or more of the intra-prediction modes (IPMs) are stored in the IPM cache, or One or more of the intra-frame prediction modes (IPMs) are used for chroma prediction.

25. The method of claim 24, wherein a default IPM is stored in the IPM cache, wherein the default IPM comprises at least one of a plane or a DC.

26. The method of claim 19, wherein one or more motion vectors (MVs) are used to construct a Merge candidate list and / or an Advanced Motion Vector Prediction (AMVP) candidate list for the subsequent video units, or One or more motion vectors (MVs) are stored in the motion information cache.

27. The method of claim 26, wherein one or more MVs in the sub-segment are used, wherein the sub-segment has a larger size than another sub-segment.

28. The method of claim 26, wherein one or more MVs in the first sub-segment are used, or One or more MVs from the last sub-segment are used.

29. The method of claim 19, wherein none of the MVs in the first target encoding / decoding mode are allowed to be used.

30. The method of claim 29, wherein the motion information for the video unit is set to be equal to IBC mode and / or intraTMP mode.

31. The method of claim 19, wherein at least one of the following is stored at the sub-block level: BV, IPM, MV.

32. The method of claim 31, wherein the sub-block level includes a 4×4 sub-block level or an 8×8 sub-block level.

33. The method of claim 32, wherein at least one of the following is stored: BV used to generate the prediction signal for the sub-block or MV used to generate the prediction signal for the sub-block.

34. The method of claim 32, wherein if the mixture is used to generate the prediction signal for the sub-block, at least one of the following is stored: BV or MV.

35. The method of claim 19, wherein at least one of the following is used for the chromaticity component: the stored BV, the stored IPM, or the stored MV.

36. The method of claim 35, wherein at least one of the following is used for the chroma video unit: BV stored for the luma video unit, IPM stored for the luma video unit, or MV stored for the luma video unit.

37. The method of claim 35, wherein at least one of the following is used for a subsequent chroma video unit: a BV stored for the chroma video unit, an IPM stored for the chroma video unit, or a MV stored for the chroma video unit.

38. The method of claim 19, wherein the video unit encoded using the first target encoding / decoding mode is marked as a screen content encoding / decoding (SCC) block, and the marked information is used in adaptive OBMC control.

39. The method of claim 19, wherein the video unit encoded and decoded using the first target encoding / decoding mode is marked as a non-SCC block, and the marked information is used in adaptive OBMC control.

40. The method of claim 19, wherein whether the video unit encoded and decoded using the first target encoding / decoding mode is marked as an SCC block is the same as another GPM mode.

41. The method of claim 1, wherein the interaction between the first target codec mode and the target codec tool is the same as the interaction between the GPM-intraframe and the target codec tool.

42. The method of claim 41, wherein the target encoding / decoding tool includes a luminance map with chroma scaling (LMCS).

43. The method of claim 42, wherein the demapping process is performed on the IBC prediction signal before the IBC prediction signal is used to mix with the inter-frame prediction signal, or The inverse mapping process is performed on the intraTMP prediction signal before it is used to mix with the inter-frame prediction signal.

44. The method of claim 42, wherein the LMCS is performed if template-matching-based reordering for the GPM partitioning mode is used for the first target encoding / decoding mode.

45. The method of claim 41, wherein the target encoding / decoding tool comprises a loop filter.

46. ​​The method of claim 45, wherein the boundary strength setting of the first target encoding / decoding mode is the same as the boundary strength setting within the GPM-frame.

47. The method of claim 1, wherein whether and / or how the first target codec mode is applied depends on codec information, wherein the codec information includes at least one of the following: Whether encoding / decoding methods are allowed Block dimension, Block size, Block depth, Strip type, Image type, Segmentation tree type, Temporal layer identifier, Block location, Color format, Color components, or Video content.

48. The method of claim 47, wherein if the block size is greater than or equal to a first threshold, the block is not permitted to be encoded or decoded using the first target encoding / decoding mode, wherein the block size is equal to the block width multiplied by the block height.

49. The method of claim 48, wherein the first threshold comprises one of the following: 64, 128, 256, 512, 1024, 2048, or 4096.

50. The method of claim 47, wherein if the block size is less than or equal to a second threshold, the block is not permitted to be encoded or decoded using the first target encoding / decoding mode, wherein the block size is equal to the block width multiplied by the block height.

51. The method of claim 50, wherein the second threshold comprises one of the following: 8, 16, 32, 64, 128, 256 or 512.

52. The method of claim 47, wherein a block is not permitted to be encoded or decoded using the first target encoding / decoding mode if at least one of the following conditions is met: the block width of the block is greater than or equal to a third threshold, or the block height of the block is less than or equal to a fourth threshold.

53. The method of claim 52, wherein the third threshold comprises one of the following: 4, 8, 16, 32 or 64.

54. The method of claim 52, wherein the fourth threshold comprises one of the following: 4, 8, 16, 32 or 64.

55. The method of claim 47, wherein a block is not permitted to be encoded or decoded using the first target encoding / decoding mode if at least one of the following conditions is met: the block width of the block is less than or equal to a fifth threshold, or the block height of the block is greater than or equal to a sixth threshold.

56. The method of claim 55, wherein the fifth threshold comprises one of the following: 8, 16, 32, 64, or 128.

57. The method of claim 55, wherein the sixth threshold comprises one of the following: 8, 16, 32, 64, or 128.

58. The method of claim 47, wherein a block is not permitted to be encoded or decoded using the first target encoding / decoding mode if at least one of the following conditions is met: the ratio of the block width to the block height of the block is greater than or equal to a seventh threshold, or the ratio of the block height to the block width is greater than or equal to the seventh threshold.

59. The method of claim 58, wherein the seventh threshold comprises one of the following: 2, 4, 8, 16 or 32.

60. The method of claim 47, wherein if the time-domain layer identifier is less than or equal to a first time-domain layer identifier threshold, the block is not allowed to be encoded or decoded using the first target encoding / decoding mode.

61. The method of claim 60, wherein the first time-domain layer identifier threshold is equal to one of the following: 0, 1 or 2.

62. The method of claim 47, wherein if the time-domain layer identifier is greater than or equal to the second time-domain layer identifier threshold, the block is not allowed to be encoded or decoded using the first target encoding / decoding mode.

63. The method of claim 62, wherein the second time-domain layer identifier threshold is equal to one of the following: 3, 4, 5 or 6.

64. The method of claim 47, wherein the video content includes at least one of the following: video captured by a camera or video of screen content.

65. The method of claim 64, wherein the first target encoding / decoding mode is applied to the screen content video.

66. The method of claim 1, wherein an indication of the first target encoding / decoding mode is transmitted via signaling if a condition is met, wherein the condition includes at least one of the following: Block dimension, Block size, Block depth, Strip type, Image type, Segmentation tree type, Temporal layer identifier, Block location, Color format, Color components, or Video content.

67. The method of claim 66, wherein if the block size is greater than or equal to an eighth threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling, wherein the block size is equal to the block width multiplied by the block height.

68. The method of claim 67, wherein the eighth threshold comprises one of the following: 64, 128, 256, 512, 1024, 2048, or 4096.

69. The method of claim 66, wherein if the block size is less than or equal to a ninth threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling, wherein the block size is equal to the block width multiplied by the block height.

70. The method of claim 69, wherein the ninth threshold comprises one of the following: 16, 32, 64, 128, 256, or 512.

71. The method of claim 66, wherein the indication of the first target encoding / decoding mode is not transmitted by signal if at least one of the following conditions is met: the block width of the block is less than or equal to a tenth threshold, or the block height of the block is greater than or equal to an eleventh threshold.

72. The method of claim 71, wherein the tenth threshold comprises one of the following: 4, 8, 16, 32 or 64.

73. The method of claim 71, wherein the eleventh threshold comprises one of the following: 4, 8, 16, 32 or 64.

74. The method of claim 66, wherein the indication of the first target encoding / decoding mode is not transmitted by signal if at least one of the following conditions is met: the block width of the block is greater than or equal to the twelfth threshold, or the block height of the block is less than or equal to the thirteenth threshold.

75. The method of claim 74, wherein the twelfth threshold comprises one of the following: 8, 16, 32, 64, or 128.

76. The method of claim 74, wherein the thirteenth threshold comprises one of the following: 8, 16, 32, 64, or 128.

77. The method of claim 66, wherein the indication of the first target encoding / decoding mode is not transmitted via signaling if at least one of the following conditions is met: the ratio of the block width to the block height of the block is greater than or equal to a fourteenth threshold, or the ratio of the block height to the block width is greater than or equal to the fourteenth threshold.

78. The method of claim 77, wherein the fourteenth threshold comprises one of the following: 2, 4, 8, 16 or 32.

79. The method of claim 66, wherein if the time-domain layer identifier is less than or equal to a third time-domain layer identifier threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling.

80. The method of claim 79, wherein the third time-domain layer identifier threshold is equal to one of the following: 0, 1 or 2.

81. The method of claim 66, wherein if the time-domain layer identifier is greater than or equal to a fourth time-domain layer identifier threshold, the indication of the first target encoding / decoding mode is not transmitted via signaling.

82. The method of claim 81, wherein the fourth time-domain layer identifier threshold is equal to one of the following: 3, 4, 5 or 6.

83. The method of claim 66, wherein the video content includes at least one of the following: video captured by a camera or video of screen content.

84. The method of claim 83, wherein, for video captured by the camera, the indication of the first target encoding / decoding mode is not transmitted via signal transmission.

85. The method of claim 84, wherein the indication of the first target encoding / decoding mode is presumed to be a default value.

86. The method of claim 1, wherein whether the current block is encoded or decoded using the first target encoding / decoding mode is transmitted via signaling using one or more syntax elements (SE).

87. The method of claim 86, wherein whether and / or how the one or more syntax elements are transmitted via signal depends on the target codec tool.

88. The method of claim 87, wherein the target encoding / decoding tool includes GPM-intraframe.

89. The method of claim 88, wherein the one or more syntax elements are transmitted via signaling if a condition is met, wherein the condition includes whether the GPM-intraframe is used.

90. The method of claim 88, wherein the syntax element indicating whether the first target encoding / decoding mode is used is not transmitted via signal.

91. The method of claim 88, wherein the syntax element indicating which BV is used is transmitted via signaling together with the syntax element indicating which intra-prediction mode is used within the GPM-frame.

92. The method of claim 1, wherein the syntax element indicating which intra-prediction mode is used within a GPM-frame is context-encoded.

93. The method according to any one of claims 1-92, wherein the video unit comprises at least one of the following: Color components, Sub-images, strip, piece, Code-decode tree unit (CTU) CTU line, CTU group, Codec Unit (CU) Prediction Unit (PU) Transformer Unit (TU) Code-decode tree block (CTB). Code Block (CB), Predicted Block (PB), Transform block (TB), piece, Sub-blocks of a block Sub-regions within the block This includes regions containing more than one sample point or pixel.

94. The method according to any one of claims 1-92, wherein the indication of whether and / or how to apply the first target codec mode is indicated in one of the following places: sequence level, Image group level, Image quality, strip level, or Film series level.

95. The method according to any one of claims 1-92, wherein the indication of whether and / or how to apply the first target codec mode is indicated in one of the following: Sequence header, Image header, Sequence Parameter Set (SPS) Video Parameter Set (VPS) Dependency Parameter Set (DPS) Decoding Capability Information (DCI) Image Parameter Set (PPS) Adaptive Parameter Set (APS) strip head, or The beginning of the film.

96. The method according to any one of claims 1-92, wherein whether and / or how the first target encoding / decoding mode is applied depends on at least one of the following: Messages transmitted via signaling in one of the following formats: DPS, SPS, VPS, PPS, APS, image header, strip header, slice header, codec tree unit (CTU), codec unit (CU), CTU line, CTU group, TU, PU block, or video codec unit. The location of CU The location of PU The location of TU The location of the block The location of the video encoding / decoding unit. The block dimension of the current block. The block dimensions of the neighboring blocks of the current block. The shape of the current block. The block shape of the neighboring blocks of the current block. Block encoding / decoding modes, Indicators of color format, Encoder tree structure, Strip group type, Film set type, Image type, Color components, Time-domain ID, Standard level Standard level, or Standard layer.

97. The method of claim 96, wherein the encoding / decoding mode of the block includes at least one of the following: IBC inter-frame mode, non-IBC inter-frame mode, non-IBC sub-block mode, intraTMP inter-frame mode, non-intraTMP inter-frame mode, or non-intraTMP sub-block mode.

98. The method of claim 96, wherein the color format indication includes 4:2:0 or 4:4:

4.

99. The method of claim 96, wherein the color component is applied to one of: a chromaticity component or a luminance component.

100. The method according to any one of claims 1-99, wherein the conversion comprises encoding the video unit into the bitstream.

101. The method according to any one of claims 1-99, wherein the conversion comprises decoding the video unit from the bitstream.

102. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-101.

103. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1-101.

104. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by means of a video processing apparatus, wherein the method includes: A first target encoding / decoding mode is applied to a video unit of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools including at least one of the following: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; and The bitstream is generated based on the first target encoding / decoding mode.

105. A method for storing a bitstream of video, comprising: The first target encoding / decoding mode is applied to the video unit of the video, wherein the first target encoding / decoding mode includes inter-frame prediction and encoding / decoding tools, the encoding / decoding tools include at least one of the following: intra-block copying (IBC) or intra-template matching prediction (intraTMP), and the video unit includes multiple sub-segments or multiple sub-blocks; The bitstream is generated based on the first target encoding / decoding mode; and The bitstream is stored in a non-transitory computer-readable recording medium.