Method and device for video processing and medium
By adopting the combination method of intra-block copying and geometric segmentation mode in video encoding and decoding technology, the problem of insufficient encoding and decoding efficiency in the prior art is solved, and more efficient video encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202380072689.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-14
- Filing Date
- 2023-10-13
- Publication Date
- 2025-05-30
AI Technical Summary
The existing video encoding and decoding technology has shortcomings in improving the encoding and decoding efficiency, especially in the design of intra prediction modes and motion vector prediction.
The method of intra-block copy (IBC) combined with geometric segmentation mode (GPM) is adopted to obtain sub-segment prediction of the video unit through intra-block copy, and the video unit is encoded and coded based on the geometric segmentation mode.
Improves the performance and efficiency of video encoding and decoding, especially when processing video frames containing complex edges and motion.
Smart Images

Figure CN120077649A_ABST
Abstract
Description
Technical Field
[0002] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to intra block copy with geometric partitioning. Background Art
[0003] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is generally a desire to further improve the encoding and decoding efficiency of video coding and decoding technologies. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for video processing.
[0005] In a first aspect, a method for video processing is proposed. The method includes: for the conversion between a video unit of a video and a bitstream of the video, dividing the video unit into a plurality of sub-partitions using a predefined method, where the video unit is encoded and decoded using an intra block copy (IBC)-geometric partitioning mode (GPM); obtaining a prediction of at least one sub-partition of the video unit using intra block copy; and performing the conversion based on the prediction of at least one sub-partition of the video unit. This can improve the encoding and decoding performance and efficiency.
[0006] In a second aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method includes: dividing a video unit of the video into a plurality of sub-partitions using a predefined method, where the video unit is encoded and decoded using an intra block copy (IBC)-geometric partitioning mode (GPM); obtaining a prediction of at least one sub-partition of the video unit using intra block copy; and generating a bitstream based on the prediction of at least one sub-partition of the video unit. This can improve the encoding and decoding performance and efficiency.
[0009] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes: dividing a video unit of the video into a plurality of sub-divisions using a predefined method, wherein the video unit is encoded and decoded using an Intra Block Copy (IBC)-Geometry Partitioning Mode (GPM); obtaining a prediction of at least one sub-division of the video unit using Intra Block Copy; generating a bitstream based on the prediction of at least one sub-division of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium. This can improve the encoding and decoding performance and efficiency.
[0010] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify the key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram showing an example video encoding and decoding system according to some embodiments of the present disclosure is shown;
[0013] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure is shown;
[0014] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure is shown;
[0015] Figure 4 An example of an encoder block diagram is shown;
[0016] Figure 5 67 intra prediction modes are shown;
[0017] Figure 6 Reference sample points for wide-angle intra prediction are shown;
[0018] Figure 7 The problem of discontinuity in the case of a direction exceeding 45° is shown;
[0019] Figure 8 MMVD search points are shown;
[0020] Figure 9 It is an illustration for a symmetric MVD mode;
[0021] Figure 10 An extended CU region used in BDOF is shown;
[0022] Figure 11 Shows the top and left neighboring blocks used in CIIP weight derivation;
[0023] Figure 12 Shows an affine motion model based on control points;
[0024] Figure 13 Shows the affine MVF of each sub-block;
[0025] Figure 14 Shows the position of the inherited affine motion prediction value;
[0026] Figure 15 Shows control point motion vector inheritance;
[0027] Figure 16 Shows the positioning of candidate positions for constructing the affine Merge mode;
[0028] Figure 17 Is an illustration of the motion vector usage of the proposed combination method;
[0029] Figure 18 Shows sub-block MV VSB and pixel Δv(i,j);
[0030] Figure 19A Shows the spatial neighboring blocks used in ATVMP;
[0031] Figure 19B Shows the derivation of the sub-CU motion field by applying motion displacements from spatial neighbors and scaling the motion information from the corresponding co-located CUs;
[0032] Figure 20 Shows position illumination compensation;
[0033] Figure 21 Shows short-side no subsampling;
[0034] Figure 22 Shows decoder-side motion vector refinement;
[0035] Figure 23 Shows the diamond region in the search area;
[0036] Figure 24 Shows the positions of spatial Merge candidates;
[0037] Figure 25 Shows the candidate pairs considered for redundancy check of spatial Merge candidates;
[0038] Figure 26 Is an illustration of motion vector scaling for temporal Merge candidates;
[0039] Figure 27 Shows the candidate positions of the temporal Merge candidates C0 and C1;
[0040] Figure 28 Shows the VVC spatial neighborhood blocks of the current block;
[0041] Figure 29 Is an illustration of the virtual block in the i-th search round;
[0042] Figure 30 Shows an example of the GPM division grouped by the same angle;
[0043] Figure 31 Shows the unidirectional prediction MV selection for the geometric partitioning mode;
[0044] Figure 32 Shows the exemplary generation of the bending weight w0 using the geometric partitioning mode;
[0045] Figure 33 Shows the spatial neighborhood blocks used to derive the spatial Merge candidates;
[0046] Figure 34 Shows performing template matching on the search area around the initial MV;
[0047] Figure 35 Is an illustration of the sub-blocks of the OBMC application;
[0048] Figure 36 Shows the SBT position, type, and transform type;
[0049] Figure 37 Shows the neighboring samples for calculating the SAD;
[0050] Figure 38 Shows the neighboring samples for calculating the SAD of the sub-CU level motion information;
[0051] Figure 39 Shows the sorting process;
[0052] Figure 40 Shows the reordering process in the encoder;
[0053] Figure 41 Shows the reordering process in the decoder;
[0054] Figure 42 Shows the IBC reference region depending on the current CU position;
[0055] Figure 43 Shows an example of the symmetry in the screen content picture;
[0056] Figure 44AIllustration of BV adjustment for horizontal flipping;
[0057] Figure 44B Illustration of BV adjustment for vertical flipping;
[0058] Figure 45 Shows the in-frame template matching search area used;
[0059] Figure 46 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and
[0060] Figure 47 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.
[0061] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description of the Invention
[0062] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein may be implemented in various ways other than those described below.
[0063] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.
[0064] References in the present disclosure to "one embodiment", "an embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0065] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0066] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes" and / or "including" when used herein specify the presence of the stated features, elements and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0067] Example environment
[0068] Figure 1 is a block diagram showing an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0069] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.
[0070] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that forms an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 through the network 130A. The encoded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0071] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from a source device 110 or a storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.
[0072] The video encoder 114 and the video decoder 124 may operate according to video compression standards (such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards).
[0073] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.
[0074] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 an example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0075] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0076] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0077] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for purposes of explanation, these components are Figure 2is shown separately in the example of
[0078] The splitting unit 201 may split the picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0079] The mode selection unit 203 may select, for example, one coding mode from multiple coding modes (intra coding or inter coding) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block to be used as a reference picture. In some examples, the mode selection unit 203 may select a combined intra and inter prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0080] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the cache 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of a picture from the cache 213 other than the picture associated with the current video block.
[0081] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P slice" and a "B slice" may refer to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture.
[0082] In some examples, the motion estimation unit 204 may perform uni-directional prediction on the current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, the reference index indicating the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicating the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, a prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0083] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find a reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indices and a plurality of motion vectors, the plurality of reference indices indicating the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicating the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0084] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0085] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.
[0086] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0087] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0088] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0089] The residual generation unit 207 can generate residual data for a current video block by subtracting (e.g., indicated by a minus sign) a (plurality of) predicted video blocks of the current video block from the current video block. The residual data of the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.
[0090] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0091] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0092] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0093] The inverse quantization unit 210 and the inverse transform unit 211 can respectively apply inverse quantization and inverse transform to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. The reconstruction unit 212 can add the reconstructed residual video blocks to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0094] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the block effect artifacts in the video block.
[0095] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream including the entropy encoded data.
[0096] Figure 3 is a block diagram showing an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0097] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3In an example, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in this disclosure.
[0098] In Figure 3 an example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, video decoder 300 may perform a decoding process generally opposite to the encoding process described with respect to video encoder 200.
[0099] Entropy decoding unit 301 may retrieve an encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 may decode the entropy-coded video data, and motion compensation unit 302 may determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. Motion compensation unit 302 may determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.
[0100] Motion compensation unit 302 may produce a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision may be included in a syntax element.
[0101] Motion compensation unit 302 may use the interpolation filter used by video encoder 200 during encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. Motion compensation unit 302 may determine the interpolation filter used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filter to produce a prediction block.
[0102] The motion compensation unit 302 may use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information that describes how each macroblock of the pictures of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" may refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding / decoding, signal prediction, and residual signal reconstruction. A slice may be the entire picture or may also be a region of the picture.
[0103] The intra prediction unit 303 may use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0104] The reconstruction unit 306 may obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter may also be applied to filter the decoded block to remove blocking artifact. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also produces the decoded video for presentation on a display device.
[0105] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. Additionally, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. Further, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates.
[0106] 1. Brief Overview
[0107] This disclosure relates to video coding and decoding technologies. Specifically, it relates to Intra Block Copy (IBC), how and / or whether to combine IBC with geometric partitioning, and other coding and decoding tools in image / video coding and decoding. It can be applied to existing video coding and decoding standards such as HEVC or Versatile Video Coding (VVC). It can also be applicable to future video coding and decoding standards or video codecs.
[0108] 2. Introduction
[0109] Video coding and decoding standards have evolved mainly through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Coding (AVC), and the H.265 / HEVC standard. Since H.262, video coding and decoding standards have been based on a hybrid video coding and decoding structure, in which temporal prediction plus transform coding is utilized. To explore future video coding and decoding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, many new methods have been adopted by JVET and incorporated into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard, with the goal of reducing the bitrate by 50% compared to HEVC.
[0110] 2.1. Coding and Decoding Processes of Typical Video Codecs
[0111] Figure 4 Shows an example of the encoder block diagram of VVC, which includes three in-loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Different from DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets and by applying a Finite Impulse Response (FIR) filter respectively, where the coding and decoding side information is signaled by the offsets and filter coefficients. ALF is located at the last processing stage of each picture and can be regarded as a tool to attempt to capture and fix the artifacts generated in the previous stage.
[0112] 2.2. Intra Mode Coding and Decoding with 67 Intra Prediction Modes
[0113] To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65, as Figure 5 shown, and the planar mode and DC mode remain unchanged. These dense directional intra prediction modes are applicable to all block sizes and applicable to both luma intra prediction and chroma intra prediction.
[0114] In HEVC, each intra-coded block has a square shape and the length of each of its sides is a power of 2. Therefore, when generating the intra prediction value using the DC mode, no division operation is required. In VVC, the block can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks.
[0115] 2.2.1. Wide-angle Intra Prediction
[0116] Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index further depends on the block shape. The conventional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VVC, several conventional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode coding and decoding method remains unchanged.
[0117] To support these prediction directions, a top reference of length 2W + 1 and a left reference of length 2H + 1 are defined, as Figure 6 shown.
[0118] The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2-1.
[0119] Table 2-1 - Intra Prediction Modes Replaced by Wide-Angle Modes
[0120] As Figure 7As shown, in the case of wide-angle frame prediction, two vertically adjacent predicted sample points can use two non-adjacent reference sample points. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap Δpα. If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that satisfy this condition, which are [-14, -12, -10, -6, 72, 76, 78, 80]. When the block is predicted through these modes, the sample points in the reference cache are directly copied without applying any interpolation. With this modification, the number of sample points that need to be smoothed is reduced. In addition, it aligns the design of the regular prediction mode and the non-fractional mode in the wide-angle mode.
[0121] In VVC, in addition to supporting the 4:2:0 chroma format, the 4:2:2 chroma format and the 4:4:4 chroma format are also supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format was originally ported from HEVC, and the number of entries was extended from 35 to 67 to align with the extension of the intra prediction mode. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes in the range of 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values in the mapping table entries to more accurately convert the prediction angles of chroma blocks.
[0122] 2.3. Inter-frame prediction
[0123] For each inter-frame predicted CU, the motion parameters include the motion vector, the reference picture index and the reference picture list usage index, and additional information required for the new codec features of VVC for sample generation to be used for inter-frame prediction. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in the skip mode, the CU is associated with a PU and does not have significant residual coefficients, no coded motion vector delta or reference picture index. The Merge mode is specified, where the motion parameters for the current CU are obtained from neighboring CUs (including spatial candidates and temporal candidates), and additional mechanisms introduced in VVC are adopted. The Merge mode can be applied to any inter-frame predicted CU, not just the skip mode. An alternative to the Merge mode is the explicit transmission of motion parameters, where the motion vector, the corresponding reference picture index and the reference picture list usage flag for each reference picture list, and other required information are explicitly signaled for each CU.
[0124] 2.4. Intra block copy (IBC)
[0125] Intra Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. As is well known, it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has been reconstructed within the current picture. The luminance block vector of the CU encoded and decoded by IBC is of integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precisions. The CU encoded and decoded by IBC is regarded as a third prediction mode in addition to the intra or inter prediction modes. The IBC mode is applicable to CUs with a width and height both less than or equal to 64 luminance samples.
[0126] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height not greater than 16 luminance samples. For non-Merge modes, block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a local search based on block matching will be performed.
[0127] In the hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For a current block with a larger size, the hash key is determined to match the hash key of the reference block when the hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost of each matching reference is calculated, and the one with the minimum cost is selected.
[0128] In the block matching search, the search range is set to cover the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using flags, and it can be signaled in IBC AMVP mode or IBC skip / Merge mode as follows:
[0129] – IBC skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC-encoded blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and paired candidates.
[0130] – IBC AMVP Mode: The block vector difference is coded and decoded in the same way as the motion vector difference. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the upper neighbor (if it is IBC coding / decoding). When either neighbor is not available, the default block vector will be used as the prediction value. A flag is signaled to indicate the block vector prediction value index.
[0131] 2.5. Merge Mode with MVD (MMVD)
[0132] In addition to the Merge mode that directly uses the implicitly derived motion information for the prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag is sent to specify which MMVD mode is used for the CU.
[0133] In MMVD, after a Merge candidate is selected, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, the index specifying the motion magnitude, and the index for the indication of the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected as the MV basis. The Merge candidate flag is signaled to specify which one of the first Merge candidate and the second Merge candidate is used.
[0134] The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2-2.
[0135] Table 2-2: Relationship between Distance Index and Predefined Offset Distance Index 0 1 2 3 4 5 6 7 Offset (in luminance samples) 1 / 4 1 / 2 1 2 4 8 16 32
[0136] The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions, as shown in Table 2-3. Note that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a uni-directional prediction MV or a bi-directional prediction MV and both lists point to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture or both are less than the POC of the current picture), the symbol specification in Table 2-3 is added to the sign of the MV offset of the starting MV. When the starting MV is a bi-directional prediction MV and the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture and the POC of the other reference is less than the POC of the current picture), and the difference in POC in list 0 is greater than the difference in POC in list 1, the symbol specification in Table 2-3 is added to the sign of the MV offset of the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. In other cases, if the difference in POC in list 1 is greater than list 0, the symbol specification in Table 2-3 is added to the sign of the MV offset of the list 1 MV component of the starting MV, and the sign of the list 0 MV has the opposite value.
[0137] The MVD is scaled according to the difference in POC in each direction. If the difference in POC in both lists is the same, no scaling is required. Otherwise, if the difference in POC in list 0 is greater than the difference in POC in list 1, then the MVD for list 1 is scaled by defining the difference in POC of L0 as td and the difference in POC of L1 as tb, as Figure 26 described. If the difference in POC of L1 is greater than L0, then the MVD of list 0 is scaled in the same way. If the starting MV is uni-directional prediction, then the MVD is added to the available MV.
[0138] Table 2-3: Signs of MV offset specified by the direction index Direction Index 00 01 10 11 x - axis + – N / A N / A y - axis N / A N / A + –
[0139] 2.6. Symmetric MVD encoding and decoding
[0140] In VVC, in addition to the regular uni-directional prediction mode MVD signaling and bi-directional prediction mode MVD signaling, the symmetric MVD mode for the bi-directional prediction MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not signaled but is derived.
[0141] The decoding process of the symmetric MVD mode is as follows:
[0142] 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows:
[0143] — If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0.
[0144] — Otherwise, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise BiDirPredFlag is set to 0.
[0145] 2) At the CU level, if the CU is coded in bi-directional prediction and BiDirPredFlag is equal to 1, then the symmetry mode flag indicating whether the symmetry mode is used is signaled explicitly.
[0146] When the symmetry mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indices of list 0 and list 1 are set to be equal to the reference picture pair respectively. MVD1 is set to be equal to (-MVD0). The final motion vector is as shown in the following formula.
[0147] In the encoder, the symmetric MVD motion estimation starts from the initial MV estimation. A set of initial MV candidates includes the MVs obtained from the uni-directional prediction search, the MVs obtained from the bi-directional prediction search, and the MVs from the AMVP list. The one with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.
[0148] 2.7. Bidirectional Optical Flow (BDOF)
[0149] The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, is included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version and requires much less computation, especially in terms of the number of multiplications and the size of the multiplier.
[0150] BDOF is used to refine the bi-directional prediction signal of the CU at the 4×4 sub-block level. BDOF is applied to the CU if all of the following conditions are met:
[0151] — The CU is coded using the "true" bi-directional prediction mode, i.e., one of the two reference pictures is before the current picture in the display order, and the other of the two reference pictures is after the current picture in the display order;
[0152] — The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same;
[0153] — Both reference pictures are short-term reference pictures;
[0154] — The CU is coded without using the affine mode or the SbTMVP Merge mode;
[0155] — The CU has more than 64 luma samples;
[0156] — Both the CU height and the CU width are greater than or equal to 8 luma samples;
[0157] — The BCW weight index indicates equal weights;
[0158] — The current CU does not have WP enabled;
[0159] — The CIIP mode is not used for the current CU.
[0160] BDOF is only applied to the luma component. As its name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of objects is smooth. For each 4×4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bi-predicted sample values in the 4x4 sub-block. The following steps are applied during BDOF.
[0161] First, by directly calculating the differences between two neighboring samples, the horizontal and vertical gradients of the two prediction signals, and are calculated, i.e.,
[0162] where I (k) (i, j) is the sample value at the coordinates (i, j) of the prediction signal in the list k, k = 0, 1, and shift1 is calculated based on the luma bit depth bitDepth as shift1 = max(6, bitDepth - 6).
[0163] Then, the autocorrelations and cross-correlations of the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated as follows:
[0164] where, θ(i, j) = (I (1)(i,j) >> n b )-(I (0) (i,j) >> n b )
[0165] where Ω is a 6×6 window surrounding a 4×4 sub-block, and n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively.
[0166] Then, using the cross-correlation terms and auto-correlation terms, the motion refinement (v x , v y ) is derived using the following method:
[0167] where th′ BIO = 2 max(5,BD-7) . is the floor function, and
[0168] Based on the motion refinement and gradient, the following adjustment is calculated for each sample point in the 4×4 sub-block:
[0169] Finally, the BDOF sample points of the CU are calculated by adjusting the bi-predicted sample points in the following manner: pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + o offset ) >> shift (2 - 7)
[0170] These values are chosen such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits.
[0171] To derive the gradient values, some predicted sample points I (k) (i,j) in list k (k = 0, 1) outside the current CU boundary need to be generated. As Figure 10As shown, the BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating out-of-boundary prediction samples, the prediction samples in the extended region (white positions) are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without using interpolation, and the regular 8-tap motion compensation interpolation filter is used to generate the prediction samples within the CU (gray positions). These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e., repeated) from their nearest neighbors.
[0172] When the width and / or height of the CU is greater than 16 luma samples, it is divided into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries are considered as the CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than the threshold, the BDOF process is not applied to that sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD calculated in the DVMR process between the initial L0 prediction samples and the L1 prediction samples is reused here.
[0173] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., the luma_weight_lx_flag of any one of the two reference pictures is 1, then the BDOF is also disabled. When the CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled.
[0174] 2.8. Combined Inter-Frame and Intra-Frame Prediction (CIIP)
[0175] In VVC, when a CU is encoded or decoded in the Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter-frame / intra-frame prediction (CIIP) mode is applied to the current CU. As its name implies, CIIP prediction combines the inter-frame prediction signal with the intra-frame prediction signal. The inter-frame prediction signal P in the CIIP mode inter is derived using the same inter-frame prediction process applied to the regular Merge mode; and the intra-frame prediction signal P intraIt is derived after the conventional intra prediction process using the planar mode. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weight values are calculated depending on the coding / decoding modes of the top and left neighboring blocks (as Figure 11 depicted) as follows:
[0176] — If the top neighbor is available and is intra-coded / decoded, set isIntraTop to 1, otherwise set isIntraTop to 0;
[0177] — If the left neighbor is available and is intra-coded / decoded, set isIntraLeft to 1, otherwise set isIntralLeft to 0;
[0178] — If (isIntraLeft + isIntraTop) equals 2, set wt to 3;
[0179] — Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2;
[0180] — Otherwise, set wt to 1.
[0181] The CIIP prediction is established as follows: P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2 (2 - 1).
[0182] 2.9. Affine Motion Compensation Prediction
[0183] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many kinds of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As Figure 12 shown, the affine motion field of a block is described by the motion information of two control point motion vectors (4 parameters) or three control point motion vectors (6 parameters).
[0184] For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as:
[0185] For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as:
[0186] where (mv 0x , mv 0y(mv 1x is the motion vector of the upper left control point, 1x , mv 1y 1y ) is the motion vector of the upper right control point, and (mv 2x , mv 2y 2x , mv 2y ) is the motion vector of the lower left control point.
[0187] To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector of each 4×4 luma sub-block, the motion vector of the center sample of each sub-block (as shown in Figure 13 ) is calculated according to the above equation and rounded to 1 / 16 fractional precision. Then a motion compensation interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size of the chrominance component is also set to 4×4. The MV of the 4×4 chrominance sub-block is calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks.
[0188] Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode.
[0189] 2.9.1. Affine Merge Prediction
[0190] The AF_MERGE mode can be applied to a CU whose width and height are both greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of the spatial neighboring CUs. There can be up to five CPMV candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPVM candidates are used to form the affine Merge candidate list:
[0191] - Inherited affine Merge candidates inferred from the CPMV of neighboring CUs;
[0192] - Constructed affine Merge candidate CPMV derived using the translational MVs of neighboring CUs;
[0193] - Zero MV.
[0194] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as shown in Figure 14 Figure 14 As shown. For the predicted values on the left side, the scanning order is A0 -> A1, and for the predicted values on the upper side, the scanning order is B0 -> B1 -> B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When adjacent affine CUs are identified, their control point motion vectors are used to derive the CPMV candidates in the affine Merge list of the current CU. As shown in the figure, if the neighboring lower-left block A is coded in affine mode, the motion vectors v 2 , v 3 and v 4 of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated according to v 2 and v 3 . When block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated according to v 2 , v 3 and v 4 .
[0195] The constructed affine candidates refer to constructing candidates by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the specified spatial neighbors and temporal neighbors shown in Figure 16 . CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV 1 , the B2 -> B3 -> A2 block is checked, and the MV of the first available block is used. For CPMV 2 , the B1 -> B0 block is checked, and for CPMV 3 , the A1 -> A0 block is checked. TMVP is used as CPMV 4 (if it is available).
[0196] After the MVs of the four control points are obtained, the affine Merge candidates are constructed based on this motion information. The following combinations of control point MVs are used to construct in sequence: {CPMV 1 , CPMV 2 , CPMV 3}, {CPMV 1 , CPMV 2 , CPMV 4}, {CPMV 1 , CPMV 3 , CPMV 4}, {CPMV 2 , CPMV 3 , CPMV 4}, {CPMV 1, CPMV 2},{CPMV 1 , CPMV 3}。
[0197] Combinations of 3 CPMVs are used to construct 6-parameter affine Merge candidates, and combinations of 2 CPMVs are used to construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of the control point MVs are discarded.
[0198] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, zero MVs are inserted at the end of the list.
[0199] 2.9.2. Affine AMVP Prediction
[0200] The affine AMVP mode can be applied to CUs with widths and heights greater than or equal to 16. The affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list has a size of 2, and it is generated by sequentially using the following four types of CPVM candidates:
[0201] — Inherited affine AMVP candidates inferred from the CPMVs of neighboring CUs;
[0202] — Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs;
[0203] — Translational MVs from neighboring CUs;
[0204] — Zero MVs.
[0205] The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting the inherited affine motion prediction values into the candidate list, the deduplication process is not applied. The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 16 . The same checking order as in the construction of affine Merge candidates is used. In addition, the reference picture indices of neighboring blocks are also checked. The first block in the checking order is used, which is inter-frame coded and has the same reference picture as in the current CU. There is only one. When the current CU is coded using the 4-parameter affine mode, and mv 0 and mv 1When all are available, they are added as a candidate in the affine AMVP list. When the current CU is coded / decoded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0206] If, after checking the inherited affine AMVP candidates and the constructed AMVP candidates, the affine AMVP list candidates are still less than 2, mv 0 , mv 1 and mv 2 will be added in sequence as translational MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the list.
[0207] 2.9.3. Affine Motion Information Storage
[0208] In VVC, the CPMVs of an affine CU are stored in a separate cache. The stored CPMVs are only used to generate the inherited CPMVs in the affine Merge mode and the affine AMVP mode for the most recently coded CU. The sub-block MVs derived from the CPMVs are used for motion compensation, MV derivation for the Merge / AMVP list of translational MVs, and deblocking.
[0209] To avoid the picture line cache for additional CPMVs, the inheritance of affine motion data from the CU above the CTU is processed differently from the inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the upper CTU row, the left-bottom and right-bottom sub-block MVs in the line cache are used instead of the CPMVs for affine MVP derivation. In this way, the CPMVs are only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is degraded to a 4-parameter model. As Figure 17 shown, along the top CTU boundary, the left-bottom and right-bottom sub-block motion vectors of the CU are used for affine inheritance of the CU in the lower CTU.
[0210] 2.9.4. Prediction Refinement Using Optical Flow for Affine Mode
[0211] Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a more refined motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the differences derived from the optical flow equations. PROF is described in the following four steps:
[0212] Step 1) Sub - block - based affine motion compensation is performed to generate sub - block prediction I(i,j).
[0213] Step 2) Using a 3 - tap filter [-1,0,1], the spatial gradients g x (i,j) and g y (i,j) are calculated at each sample position. The gradient calculation is exactly the same as that in BDOF. g x (i,j)=(I(i + 1,j)>>shift1)-(I(i - 1,j)>>shift1) (2 - 11) g y (i,j)=(I(i,j + 1)>>shift1)-(I(i,j - 1)>>shift1) (2 - 12)
[0214] shift1 is used to control the precision of the gradient. The sub - block (i.e., 4x4) prediction is extended by one sample for each side of the gradient calculation. To avoid extra memory bandwidth and extra interpolation calculation, those extended samples on the extended boundaries are copied from the nearest integer - pixel positions in the reference picture.
[0215] Step 3) Luminance prediction refinement is calculated through the following optical flow equation. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (2 - 13)
[0216] where as Figure 18 shown, Δv(i,j) is the difference between the sample MV (denoted as v(i,j)) calculated for the sample position (i,j) and the sub - block MV of the sub - block to which the sample (i,j) belongs. Δv(i,j) is quantized in units of 1 / 32 luminance - sample precision.
[0217] Since the affine model parameters and the sample positions relative to the sub - block center do not change from sub - block to sub - block, Δv(i,j) can be calculated for the first sub - block and reused for other sub - blocks in the same CU. Let dx(i,j) and dy(i,j) be the horizontal and vertical offsets of the sample position (i,j) to the sub - block center (x SB ,y SB ), Δv(x,y) can be derived through the following equation:
[0218] To maintain accuracy, the sub - block center (x SB,y SB ) is calculated as ((W SB - 1) / 2, (H SB - 1) / 2), where W SB and H SB are the width and height of the sub - block respectively.
[0219] For the 4 - parameter affine model,
[0220] For the 6 - parameter affine model,
[0221] where (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the motion vectors of the control points at the upper - left, upper - right, and lower - left, and w and h are the width and height of the CU.
[0222] Step 4) Finally, the luminance prediction refinement ΔI(i, j) is added to the sub - block prediction I(i, j). The final prediction I’ is generated by the following equation. I′(i, j) = I(i, j)+ΔI(i, j) (2 - 18)
[0223] PROF is not applicable to the affine - coded CUs in two cases: 1) all control - point MVs are the same, which indicates that the CU only has translational motion; 2) the affine motion parameters are greater than the specified limit, because the sub - block - based affine MC is degraded to CU - based MC to avoid large memory - access bandwidth requirements.
[0224] Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. In the following two cases, PROF is not applied in the affine motion estimation stage: a) If the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) If the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low - latency picture, then PROF is not applied because the improvement introduced by PROF for this case is small. In this way, the affine motion estimation using PROF can be accelerated.
[0225] 2.10. Sub - block - based Temporal Motion Vector Prediction (SbTMVP)
[0226] VVC supports the sub - block - based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co - located picture to improve the motion vector prediction and Merge mode of the CUs in the current picture. The same co - located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects:
[0227] — TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub - CU level;
[0228] — While TMVP prefetches the temporal motion vector from the co - located block (the co - located block is the bottom - right block or the center block relative to the current CU) in the co - located picture, SbTMVP applies a motion displacement before prefetching the temporal motion information from the co - located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0229] The SbTVMP process is shown in Figure 19A and Figure 19B SbTMVP predicts the motion vectors of the sub - CUs within the current CU in two steps. In the first step, Figure 19A the spatial neighbor A1 in
[0230] is checked. If A1 has a motion vector using the co - located picture as its reference picture, that motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). Figure 19B In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current picture) to obtain the sub - CU level motion information (motion vector and reference index) from the co - located picture as shown in Figure 19B The example in
[0231] In VVC, a combined sub-block based Merge list containing both SbTMVP candidates and affine Merge candidates is used for signaling the sub-block based Merge mode. The SbTMVP mode is enabled / disabled by a Sequence Parameter Set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry of the list of sub-block based Merge candidates, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5.
[0232] The sub-CU size used in SbTMVP is fixed to 8x8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs with width and height both greater than or equal to 8.
[0233] The coding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P or B slice, additional RD checks are performed to decide whether to use the SbTMVP candidate.
[0234] 2.11. Adaptive Motion Vector Resolution (AMVR)
[0235] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the Motion Vector Difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in quarter-luma samples. In VVC, a CU-level Adaptive Motion Vector Resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be encoded / decoded with different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AVMP mode), the MVD of the current CU can be adaptively selected as follows:
[0236] — Normal AMVP mode: quarter-luma samples, half-luma samples, integer-luma samples, or four-luma samples.
[0237] — Affine AMVP mode: quarter-luma samples, integer-luma samples, or 1 / 16-luma samples.
[0238] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both the horizontal MVD and the vertical MVD of reference list L0 and reference list L1) are zero, a quarter-luma sample MVD resolution is assumed.
[0239] For a CU with at least one non-zero MVD component, the first flag is signaled to indicate whether quarter-luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luma sample MVD precision is used for the current CU. Otherwise, the second flag is signaled to indicate whether half-luma sample or other MVD precision (integer or quarter-luma sample) is used for normal AMVP CUs. In the case of half-luma samples, a 6-tap interpolation filter is used for the half-luma sample positions instead of the default 8-tap interpolation filter. Otherwise, the third flag is signaled to indicate whether integer-luma sample MVD precision or quarter-luma sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer-luma sample MVD precision or 1 / 16-luma sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luma sample, half-luma sample, integer-luma sample or quarter-luma sample), the predicted motion vector of the CU is rounded to the same precision as the MVD before being added to the MVD. The predicted motion vector is rounded towards zero (i.e., a negative predicted motion vector is rounded towards positive infinity and a positive predicted motion vector is rounded towards negative infinity).
[0240] The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM11, the RD check for MVD accuracy outside the quarter-luma samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for MVD accuracy of quarter-luma samples and the RD cost for MV accuracy of integer-luma samples are calculated. Then, the RD cost for MVD accuracy of integer-luma samples is compared with the RD cost for MVD accuracy of quarter-luma samples to decide whether it is necessary to further check the RD cost for MVD accuracy of four-luma samples. When the RD cost for MVD accuracy of quarter-luma samples is much smaller than the RD cost for MVD accuracy of integer-luma samples, the RD check for MVD accuracy of four-luma samples is skipped. Then, if the RD cost for MVD accuracy of integer-luma samples is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for MVD accuracy of half-luma samples is skipped. For the affine AMVP mode, if the affine inter prediction mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, the Merge / skip mode, the normal AMVP mode with MVD accuracy of quarter-luma samples, and the affine AMVP mode with MVD accuracy of quarter-luma samples, the affine inter prediction modes with MV accuracy of 1 / 16-luma samples and 1-pixel MV accuracy are not checked. Additionally, in the affine inter prediction modes with 1 / 16-luma samples and quarter-luma samples MV accuracy, the affine parameters obtained in the affine inter prediction mode with quarter-luma samples MV accuracy are used as the starting search points.
[0241] 2.12. Bi-directional prediction with CU-level weighting (BCW)
[0242] In HEVC, the bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals: P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (2 - 19)
[0243] Five weights are allowed in weighted bi - prediction, \(w\in\{- 2,3,4,5,10\}\). For each bi - predicted CU, the weight \(w\) is determined in one of two ways: 1) For non - Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (\(w\in\{3,4,5\}\)) are used.
[0244] — At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. For further details, the reader may refer to the VTM software and the document JVET - L0646. When combined with AMVR, if the current picture is a low - latency picture, unequal weights are only conditionally checked for 1 - pixel and 4 - pixel motion vector precisions.
[0245] — When combined with affine, affine ME is performed for unequal weights if and only if the affine mode is selected as the current best mode.
[0246] — When the two reference pictures in bi - prediction are the same, unequal weights are only conditionally checked.
[0247] — When certain conditions are met, unequal weights are not searched, depending on the POC distance between the current picture and its reference pictures, the coding - decoding QP, and the temporal level.
[0248] The BCW weight index is decoded using a context - coded bit followed by a bypass - coded bit. The first context - coded bit indicates whether equal weights are used; and if unequal weights are used, the additional bits are signaled using bypass coding to indicate which unequal weight is used.
[0249] Weighted Prediction (WP) is an encoding / decoding tool supported by the H.264 / AVC and HEVC standards for efficient encoding / decoding of video content in fading situations. Support for WP has also been added in the VVC standard. WP allows signaling of weighted parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For a Merge CU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV.
[0250] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is encoded / decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights.
[0251] 2.13. Local Illumination Compensation (LIC)
[0252] Local Illumination Compensation (LIC) is an encoding / decoding tool that addresses the problem of local illumination variations between the current picture and its temporal reference pictures. LIC is based on a linear model where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples of the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x,y) = α·P r (x + v x , y + v y ) + β
[0253] where P(x,y) is the predicted signal of the current block at coordinates (x,y); P r (x + v x , y + v y ) is the reference block pointed to by the motion vector (v x , v y ); and α and β are the corresponding scaling factor and offset applied to the reference block. Figure 20 shows the LIC process. In Figure 20 , when LIC is applied to a block, the Least Mean Square Error (LMSE) method is adopted to minimize the neighboring samples of the current block (i.e., Figure 20the difference between the template T) in it and its corresponding reference sample points in the time-domain reference picture (i.e., Figure 20 T0 or T1 in it) to derive the values of the LIC parameters (i.e., α and β). In addition, to reduce the computational complexity, both the template samples and the reference template samples are subsampled (adaptive subsampling) to derive the LIC parameters, i.e., only Figure 20 the shaded samples in it are used to derive α and β.
[0254] To improve the coding and decoding performance, as Figure 21 shown, subsampling of the short side is not performed.
[0255] 2.14. Decoder-side Motion Vector Refinement (DMVR)
[0256] To improve the accuracy of the MVs in the Merge mode, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bi-predictive operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. As Figure 22 shown, the SAD between two blocks is calculated based on each MV candidate (e.g., MV0’ and MV1’) around the initial MV. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predictive signal.
[0257] In VVC, the application of DMVR is restricted and is only applied to the CUs encoded and decoded using the following modes and features:
[0258] — CU-level Merge mode with bi-predictive MVs;
[0259] — For the current picture, one reference picture is past and the other reference picture is future;
[0260] — The distances (i.e., POC differences) from the two reference pictures to the current picture are the same;
[0261] — Both reference pictures are short-term reference pictures;
[0262] — The CU has more than 64 luma samples;
[0263] — Both the CU height and the CU width are greater than or equal to 8 luma samples;
[0264] — The BCW weight index indicates equal weights;
[0265] — WP is not enabled for the current block;
[0266] — The CIIP mode is not used for the current block.
[0267] The refined MVs derived from the DMVR process are used to generate inter - prediction samples and are also used for temporal motion - vector prediction for future picture coding. The original MVs are used for the de - blocking process and are also used for spatial motion - vector prediction for future CU coding.
[0268] Additional functions of DMVR are mentioned in the following sub - articles.
[0269] 2.14.1. Search Scheme
[0270] In DVMR, the search points are around the initial MV, and the MV offset follows the MV - difference mirroring rule. In other words, any point examined by DMVR, represented by a candidate MV pair (MV0, MV1), follows the following two equations: MV0′ = MV0+MV_offset (2 - 20) MV1′ = MV1 - MV_offset (2 - 21)
[0271] Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer - luminance samples starting from the initial MV. The search includes an integer - sample offset search stage and a fractional - sample refinement stage.
[0272] The integer - sample offset search uses a 25 - point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer - sample stage of DMVR is terminated. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster - scan order. The point with the minimum SAD is selected as the output of the integer - sample offset search stage. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates reduces the SAD value by 1 / 4.
[0273] After the integer - sample search, there is fractional - sample refinement. To save computational complexity, the fractional - sample refinement is derived using the parametric error - surface equation instead of performing an additional search using SAD comparison. The fractional - sample refinement is conditionally invoked based on the output of the integer - sample search stage. When the integer - sample search stage is terminated with the minimum SAD at the center in the first - iteration search or the second - iteration search, the fractional - sample refinement is further applied.
[0274] In the sub - pixel offset estimation based on the parametric error - surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two - dimensional parabolic error - surface equation of the following form E(x,y)=A(x - x min ) 2+B(y - y min ) 2 +C(2 - 22)
[0275] where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1, 0) - E(1, 0)) / (2(E(-1, 0) + E(1, 0) - 2E(0, 0))) (2 - 23) y min = (E(0, -1) - E(0, 1)) / (2((E(0, -1) + E(0, 1) - 2E(0, 0))) (2 - 24).
[0276] The values of x min and y min are automatically limited between -8 and 8 because all cost values are positive and the minimum value is E(0, 0). This corresponds to a half - pixel offset with 1 / 16 - pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer - distance refined MV to obtain a sub - pixel - accurate refined incremental MV.
[0277] 2.14.2. Bilinear Interpolation and Sample Padding
[0278] In VVC, the resolution of the MV is 1 / 16 luminance samples. Samples at fractional positions are interpolated using an 8 - tap interpolation filter. In DMVR, the search points are around the initial fractional - pixel MV with integer - sample offsets, so interpolation of samples at these fractional positions is required for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that, by using the bilinear filter, within the 2 - sample search range, compared with the normal motion - compensation process, DVMR does not access more reference samples. After obtaining the refined MV using the DMVR search process, the normal 8 - tap interpolation filter is applied to generate the final prediction. To not access more reference samples of the normal MC process, samples will be padded from those available samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV.
[0279] 2.14.3. Maximum DMVR Processing Unit
[0280] When the width and / or height of a CU is greater than 16 luma samples, it is further partitioned into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16.
[0281] 2.15. Multi-pass decoder-side motion vector refinement
[0282] In this paper, instead of DMVR, multi-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coded / decoded block. In the second pass, BM is applied to each 16x16 sub-block within the coded / decoded block. In the third pass, the MVs in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction.
[0283] 2.15.1. First pass - Block-based bilateral matching MV refinement
[0284] In the first pass, the refined MVs are derived by applying BM to the coded / decoded block. Similar to decoder-side motion vector refinement (DMVR), the refined MVs are searched around two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. Based on the minimum bilateral matching cost between two reference blocks in L0 and L1, the refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs.
[0285] BM performs a local search to derive the integer sample accuracy intDeltaMV and the half-pixel sample accuracy halfDeltaMV. The local search applies a 3×3 square search pattern, cycling within the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block scale, and the maximum values of sHor and sVer are 8.
[0286] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the distorted DC effect between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the local search for intDeltaMV or halfDeltaMV is terminated. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.
[0287] Existing fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MVs after the first pass are derived as:
[0288] ·MV0_pass1 = MV0 + deltaMV
[0289] ·MV1_pass1 = MV1 – deltaMV
[0290] 2.15.2. Second Pass - Sub - block - based Bilateral Matching MV Refinement
[0291] In the second pass, refined MVs are derived by applying BM to 16×16 grid sub - blocks. For each sub - block, refined MVs are searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass for the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between two reference sub - blocks in L0 and L1.
[0292] For each sub - block, BM performs a full search to derive the integer - sample accuracy intDeltaMV. The full search has a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVert are 8.
[0293] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub - blocks, i.e., bilCost = satdCost*costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into 5 diamond - shaped search areas as Figure 23 shown. Each search area is assigned a cost factor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in raster - scan order, starting from the upper - left corner of the area and going to the lower - right corner of the area. The integer - pixel full search is terminated when the minimum bilCost within the current search area is less than or equal to the threshold of sbW*sbH; otherwise, the integer - pixel full search continues to the next search area until all search points are checked.
[0294] BM performs a local search to derive the half - sample accuracy halfDeltaMv. The search pattern and cost function both follow the definitions in Section 2.9.1.
[0295] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MV for the second pass is derived as:
[0296] ·MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2)
[0297] ·MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2).
[0298] 2.15.3. Third Pass - Sub - block - based Bidirectional Optical Flow MV Refinement
[0299] In the third pass, the refined MV is derived by applying BDOF to 8×8 grid sub - blocks. For each 8×8 sub - block, BDOF refinement is applied to derive scaled Vx and Vy without clipping starting from the refined MV of the second - pass parent - child blocks. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped to be between - 32 and 32.
[0300] The refined MV for the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as:
[0301] ·MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv
[0302] ·MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) – bioMv.
[0303] 2.16. Sample - based BDOF
[0304] In sample - based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, it is performed for each sample.
[0305] The coded - decoded block is divided into 8×8 sub - blocks. For each sub - block, whether to apply BDOF is determined by checking the SAD between two reference sub - blocks against a threshold. If it is decided to apply BDOF to the sub - block, for each sample in the sub - block, a sliding 5×5 window is used and the existing BDOF process is applied to each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional predicted sample value of the central sample of the window.
[0306] 2.17. Extended Merge Prediction
[0307] In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates:
[0308] (1) Spatial MVPs from spatially neighboring CUs
[0309] (2) Temporal MVPs from co-located CUs
[0310] (3) History-based MVPs from the FIFO table
[0311] (4) Paired-average MVPs
[0312] (5) Zero MVs.
[0313] The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coding in the Merge mode, the index of the best Merge candidate is encoded using truncated unary binary (TU). The first binary bit of the Merge index is coded using context, while bypass coding is used for the other binary bits.
[0314] The derivation process of Merge candidates for each category is provided in this section. Similar to the operation in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain-sized region.
[0315] 2.17.1. Spatial candidate derivation
[0316] The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates at the indicated positions, up to four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, and A1 are unavailable (e.g., because it belongs to another stripe or slice) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thereby improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Only the pairs connected by the arrows in Figure 25 are considered, and a candidate is added to the list only when the corresponding candidates used for redundancy checking do not have the same motion information.
[0317] 2.17.2 Temporal candidate derivation
[0318] In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on the co-located CUs belonging to the co-located reference picture. The reference picture list to be used for deriving the co-located CUs is signaled explicitly in the slice header. As Figure 26 shown by the dashed line in Figure 26 , the scaled motion vector of the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero.
[0319] The position of the temporal candidate is selected between candidates C0 and C1, as Figure 27 shown. If the CU at position C0 is not available, is intra-coded / decoded, or is outside the current row of the CTU, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate.
[0320] 2.17.3. History-based Merge candidate derivation
[0321] The history-based MVP (HMVP) Merge candidates are added to the Merge list after the spatial MVP and TMVP. In this method, the motion information of the previously coded blocks is stored in a table and is used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a CU that is non-sub-block inter-coded, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0322] The HMVP list size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is the same HMVP in the table. If found, the same HMVP is removed from the table, and then all the subsequent HMVP candidates are moved forward, and the HMVP candidates can be used for the Merge candidate list construction process. The several most recent HMVP candidates in the table are checked in order and are inserted into the candidate list after the TMVP candidates. A redundancy check to the spatial or temporal Merge candidates is applied to the HMVP candidates.
[0323] To reduce the number of redundancy check operations, the following simplification is introduced:
[0324] The number of HMPV candidates used for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table.
[0325] Once the total number of available Merge candidates reaches one less than the maximum allowed Merge candidates, the process of constructing the Merge candidate list from HMVP is terminated.
[0326] 2.17.4. Pairwise average Merge candidate derivation
[0327] Pairwise average candidates are generated by averaging predefined pairs of candidates in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices of the Merge candidate list. The average motion vectors are calculated separately for each reference list. If both motion vectors are available in a list, the two motion vectors are averaged even if they point to different reference pictures; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, this list is kept invalid.
[0328] When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum Merge candidate number is reached.
[0329] 2.17.5. Merge estimation region
[0330] The Merge estimation region (MER) allows for independent derivation of the Merge candidate list for CUs within the same Merge estimation region (MER). Candidate blocks within the same MER as the current CU are not included for generating the Merge candidate list for the current CU. Additionally, the update process for the list of candidates for history-based motion vector prediction values is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luminance sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2.
[0331] 2.18. New Merge candidates
[0332] 2.18.1. Non - adjacent Merge Candidate Derivation
[0333] In VVC, Figure 28 the five spatial neighboring blocks and one temporal neighboring block shown are used to derive Merge candidates.
[0334] It is proposed to derive additional Merge candidates from positions non - adjacent to the current block using the same pattern as in VVC. To achieve this, for each search round i, virtual blocks are generated based on the current block as follows:
[0335] First, for the current block, the relative position of the virtual block is calculated by the following formula:
[0336] Offsetx = -i×gridX, Offsety = -i×gridY
[0337] where Offsetx and Offsetty represent the offset of the upper - left corner of the virtual block relative to the upper - left corner of the current block, and gridX and gridY are the width and height of the search grid.
[0338] Second, the width and height of the virtual block are calculated by the following formula:
[0339] newWidth = i×2×gridX + currWidth newHeight = i×2×gridY + currHeight.
[0340] where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.
[0341] gridX and gridY are currently set to currWidth and currHeight respectively.
[0342] Figure 29 The relationship between the virtual block and the current block is shown.
[0343] After generating the virtual block, blocks A i , B i , C i , D i and E i can be regarded as the VVC spatial neighboring blocks of the virtual block, and their positions are obtained using the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, blocks A i , B i , C i , Di and E i It is the spatial neighboring block used in VVC Merge mode.
[0344] When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbor blocks are used.
[0345] Press B for non-adjacent spatial merge candidates 1 ->A 1 ->C 1 ->D 1 ->E 1 The order is inserted into the Merge list after the time domain Merge candidates.
[0346] 2.18.2.STMVP
[0347] It is proposed to use three spatial domain Merge candidates and one temporal domain Merge candidate to derive the average candidate as the STMVP candidate.
[0348] The STMVP is inserted before the spatial merge candidate in the upper left corner.
[0349] The STMVP candidate is deduplicated along with all previous merge candidates in the merge list.
[0350] For spatial candidates, the first three candidates in the current Merge candidate list are used.
[0351] For the temporal candidates, the same positions as the VTM / HEVC co-location positions are used.
[0352] For spatial candidates, the first, second, and third candidates inserted into the current Merge candidate list before STMVP are denoted as F, S, and T.
[0353] The temporal candidate having the same position as the VTM / HEVC co-location position used in TMVP is denoted as Col.
[0354] The motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:
[0355] 1) If the reference indexes of the four Merge candidates are all valid and equal to zero in the prediction direction X (X = 0 or 1),
[0356] mvLX=(mvLX_F+mvLX_S+mvLX_T+mvLX_Col)>>2
[0357] 2) If the reference indices of three out of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1),
[0358] mvLX = (mvLX_F × 3 + mvLX_S × 3 + mvLX_Col × 2) >> 3 or
[0359] mvLX = (mvLX_F × 3 + mvLX_T × 3 + mvLX_Col × 2) >> 3 or
[0360] mvLX = (mvLX_S × 3 + mvLX_T × 3 + mvLX_Col × 2) >> 3
[0361] 3) If the reference indices of two out of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1),
[0362] mvLX = (mvLX_F + mvLX_Col) >> 1 or
[0363] mvLX = (mvLX_S + mvLX_Col) >> 1 or
[0364] mvLX = (mvLX_T + mvLX_Col) >> 1
[0365] Note: If the temporal candidate is not available, the STMVP mode is turned off.
[0366] 2.18.3. Merge List Size
[0367] If both non - adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8.
[0368] 2.19. Geometric Partitioning Mode (GPM)
[0369] In VVC, the geometric partitioning mode is supported for inter - frame prediction. The CU - level flag is used as a kind of Merge mode to signal the geometric partitioning mode. Other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub - block Merge mode. The geometric partitioning mode supports a total of 64 partitions, applicable to all possible CU sizes w × h = 2 m ×2 n (where m, n ∈ {3…6}), excluding the two special cases of 8 × 64 and 64 × 8.
[0370] When using this mode, the CU is divided into two parts by a geometrically - located line (Figure 30 )。The position of the dividing line is derived mathematically from the perspective of a specific partition and an offset parameter. Each part of the geometric partition in the CU is inter - frame predicted using its own motion; each partition allows only uni - directional prediction, i.e., each part has a motion vector and a reference index. The uni - directional prediction motion constraint is applied to ensure that, like in traditional bi - directional prediction, each CU only requires two motion - compensated predictions. The uni - directional prediction motion for each partition is derived using the process described in 2.20.1.
[0371] If the geometric partition mode is used for the current CU, the geometric partition index (angle and offset) indicating the partition mode of the geometric partition and two Merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization for the GPM Merge indices is specified. After predicting each part of the geometric partition, the sample values along the geometric partition edge are adjusted using a blending process with adaptive weights as in 2.20.2. This is the prediction signal for the entire CU, and the transform and quantization processes are applied to the entire CU as in other prediction modes. Finally, as in 2.20.3, the motion field of the CU predicted using the geometric partition mode is stored.
[0372] 2.19.1 Uni - directional prediction candidate list construction
[0373] The uni - directional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 2.18. Let n denote the index of the uni - directional prediction motion in the geometric uni - directional prediction candidate list. The LX motion vector (where X is the parity of n) of the n - th extended Merge candidate is used as the n - th uni - directional prediction motion vector for the geometric partition mode. These motion vectors are marked with "x" in Figure 31 . If the corresponding LX motion vector of the n - th extended Merge candidate does not exist, the L(1 - X) motion vector of the same candidate is instead used as the uni - directional prediction motion vector for the geometric partition mode.
[0374] 2.19.2. Blending along the geometric partition edge
[0375] After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive the samples around the geometric partition edge. The blending weight for each position in the CU is derived based on the distance between the individual position and the partition edge.
[0376] The distance from the position (x, y) to the partition edge is derived as:
[0377] where i and j are indices for the angle and offset of the geometric partition, which depend on the geometric partition index transmitted through the signal. ρ x,j and ρ y,j The sign of depends on the angle index i.
[0378] The weight of each part of the geometric partition is derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) (2-29) w 1 (x,y) = 1 - w 0 (x,y) (2-31).
[0379] partIdx depends on the angle index i. Figure 32 An example of the weight w 0 is shown.
[0380] 2.19.3. Motion Field Storage for Geometric Partition Mode
[0381] Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded / decoded in the geometric partition mode.
[0382] The type of motion vector stored for each individual position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx ≤ 0? (1 - partIdx) : partIdx) (2-32)
[0383] where motionIdx is equal to d(4x + 2, 4y + 2), which is recalculated from Equation (2-18). partIdx depends on the angle index i.
[0384] If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, then the combination Mv from Mv0 and Mv2 is stored. The combination Mv is generated using the following procedure:
[0385] 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional predicted motion vector.
[0386] Otherwise, if Mv1 and Mv2 are from the same list, then only the unidirectional predicted motion Mv2 is stored.
[0387] 2.20. Multiple Hypothesis Prediction
[0388] In multi-hypothesis prediction (MHP), up to two additional prediction values are signaled on top of the inter-frame AMVP mode, the regular Merge mode, the affine mode, and the MMVD mode. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 = (1 - α n+1 )p n + α n+1 h n+1
[0389] The weighting factor α is specified according to Table 2-4 below.
[0390] Table 2-4—Weighting Factors for MHP add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0391] For the inter-frame AMVP mode, MHP is applied only if non-equal weights in BCW are selected in the bi-predictive mode.
[0392] The additional hypothesis can be the Merge mode or the AMVP mode. In the case of the Merge mode, the motion information is indicated by the Merge index, and the Merge candidate list is the same as in the geometric partitioning mode. In the case of the AMVP mode, the reference index, the MVP index, and the MVD are signaled.
[0393] 2.21. Non-Adjacent Spatial Candidates
[0394] Non-adjacent spatial Merge candidates are inserted after the TMVP in the regular Merge candidate list. The style of the spatial Merge candidate is as Figure 33 shown. The distance between the non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block.
[0395] 2.22. Template Matching (TM)
[0396] Template matching (TM) is a decoder-side MV derivation method for refining the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference picture (i.e., of the same size as the template). As Figure 34 shown, within the [-8, +8] pixel search range, a better MV is searched around the initial motion of the current CU. The template matching previously proposed in JVET-J0021 is adopted herein with two modifications: determining the search step size based on the AMVR mode, and TM can be cascaded using the bilateral matching process in the Merge mode.
[0397] In the AMVP mode, the MVP candidate is determined based on the template matching error by selecting the one that achieves the minimum difference between the current block template and the reference block template, and then the TM performs MV refinement only for this specific MVP candidate. The TM refines this MVP candidate by using an iterative diamond search starting from the full-pixel MVD accuracy within the [-8, +8] pixel search range (or 4 pixels for the 4-pixel AMVR mode). The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for the 4-pixel AMVR mode), and then successively using half-pixel MVD accuracy and quarter-pixel MVD accuracy as specified in Table 2-5 for the AMVR mode. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process.
[0398] Table 2-5 - Search Patterns of AMVR and Merge Mode Utilizing AMVR
[0399] In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 2-5, the TM can be performed all the way to 1 / 8 pixel MVD accuracy, or skip those accuracies beyond half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (which is used when AMVR is in the half-pixel mode). Additionally, when the TM mode is enabled, the template matching can work in an independent process or an additional MV refinement process between the block-based bilateral matching (BM) method and the sub-block-based bilateral matching method, depending on whether the BM can be enabled according to its enabling condition check.
[0400] 2.23. Overlapped Block Motion Compensation (OBMC)
[0401] Overlapped block motion compensation (OBMC) has been used in H.263 before. In JEM, different from H.263, OBMC can be turned on and off using the CU-level syntax. When OBMC is used in JEM, OBMC is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. In addition, it applies to both the luminance component and the chrominance component. In JEM, the MC block corresponds to the coding / decoding block. When the CU is coded / decoded using the sub-CU mode (including sub-CU Merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To handle the CU boundary in a unified way, OBMC is performed at the sub-block level for all MC block boundaries, where the sub-block size is set to be equal to 4×4, as Figure 35 shown.
[0402] When OBMC is applied to the current sub-block, in addition to the current motion vector, if the motion vectors of four connected neighboring sub-blocks are available and different from the current motion vector, they are also used to derive the prediction block of the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal of the current sub-block.
[0403] Represent the prediction block based on the motion vector of the neighboring sub-block as P N , where N indicates the indices of the neighboring upper, lower, left, and right sub-blocks, and the prediction block based on the motion vector of the current sub-block is represented as P C . When P N is based on the motion information of neighboring sub-blocks that contains the same motion information as the current sub-block, OBMC is not performed starting from P N . Otherwise, each sample point of P N is added to the same sample point in P C , that is, the four rows / columns of P N are added to P C . The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . Except for small MC blocks (i.e., when the height or width of the coded / decoded block is equal to 4, or when the CU is coded using the sub-CU mode), where only two rows / columns of P N are added to P C . In this case, the weighting factors {1 / 4, 1 / 8} are used for P N , while the weighting factors {3 / 4, 7 / 8} are used for P C . For P N generated based on the motion vectors of vertical (horizontal) neighboring sub-blocks, the sample points in the same row (column) of P N are added to P C with the same weighting factor.
[0404] In JEM, for CUs with a size less than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU. For CUs with a size greater than 256 luma samples or coded without using the AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account during the motion estimation stage. The prediction signal formed using the motion information of the top neighboring block and the left neighboring block by OBMC is used to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied.
[0405] 2.24. Multiple Transform Selection (MTS) for Kernel Transform
[0406] In addition to DCT-II that has been adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of both inter-coded blocks and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the basis functions of the selected DST / DCT.
[0407] Table 2-6 - Transform basis functions of DCT-II / VIII and DST VII for N-point input
[0408] To maintain the orthogonality of the transform matrix, the transform matrix is quantized more accurately than the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms.
[0409] To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter respectively. When MTS is enabled at the SPS, CU-level flags are signaled to indicate whether MTS is applied. Here, MTS is only applicable to luma. The MTS signaling is skipped when one of the following conditions is applied:
[0410] - The position of the last significant coefficient of the luma TB is less than 1 (i.e., only DC);
[0411] - The last significant coefficient of the luma TB is within the MTS zeroing region.
[0412] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types for the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 2-7. A unified transform selection for ISP and implicit MTS is used by eliminating the intra-mode and block-shape dependencies. If the current block is in the ISP mode or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 is used for both the horizontal and vertical transform kernels. In terms of the transform matrix precision, 8-bit primary transform kernels are used. Thus, all the transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, for other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8), 8-bit primary transform kernels are used.
[0413] Table 2-7 - Transformation and Signaling Mapping Table
[0414] To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are zeroed out. Only the coefficients within the 16×16 low-frequency region are retained.
[0415] Similar to HEVC, the residual of a block can be coded and decoded using the transform skip mode. To avoid redundancy in syntax coding and decoding, when the MTS_CU_flag at the CU level is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Additionally, when MTS is enabled for an inter-coded block, the implicit MTS can still be enabled.
[0416] 2.25. Sub-Block Transform (SBT)
[0417] In VTM, a sub-block transform is introduced for CUs in inter prediction. In this transform mode, only a sub-part of the residual block is coded for the CU. When an inter-predicted CU has cu_cbf equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is coded. For the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is coded using a presumed adaptive transform while the other part of the residual block is zeroed out.
[0418] When SBT is used for an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, as Figure 36 shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to a binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to an asymmetric binary tree (ABT) partition. In the ABT partition, only a small region contains non-zero residuals. If one dimension of the CU is 8 (in terms of luminance samples), the 1:3 / 3:1 partition along that dimension is not allowed. A CU can have up to 8 SBT modes.
[0419] A location-dependent transform kernel selection is applied to the luminance transform blocks in SBT-V and SBT-H (chrominance TBs always use DCT-2). The two locations, SBT-H and SBT-V, are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT location are specified in Figure 36 For example, the horizontal and vertical transforms for SBT-V location 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transforms in both dimensions are set to DCT-2. Thus, the sub-block transform jointly specifies the horizontal and vertical kernel transform types for TU slicing, cbf, and the residual block.
[0420] SBT is not applied to the CUs encoded / decoded using the combined inter-intra mode.
[0421] 2.26. Adaptive Merge Candidate Reordering Based on Template Matching
[0422] To improve the encoding / decoding efficiency, after the Merge candidate list is constructed, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list in ascending order of the template matching cost. It is operated on a subgroup basis.
[0423] The template matching cost is measured by the SAD (Sum of Absolute Differences) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes bi-predicted motion information, then as Figure 37 shown, the corresponding reference sample is the average of the corresponding reference samples in reference list 0 and the corresponding reference samples in reference list 1. If the Merge candidate includes motion information at the sub-CU level, then as Figure 38 shown, the corresponding reference sample consists of the neighboring samples of the corresponding reference sub-block.
[0424] As Figure 39 shown, the sorting process is operated on a subgroup basis. The first three Merge candidates are sorted together. The subsequent three Merge candidates are sorted together. The template size (width of the left template or height of the upper template) is 1. The subgroup size is 3.
[0425] 2.27. Adaptive Merge Candidate List
[0426] We assume that the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first subgroup, and the subsequent 3 Merge candidates are taken as the second subgroup (i.e., the last subgroup).
[0427] For the encoder, as Figure 40As shown, after the Merge candidate list is constructed, some of the Merge candidates are adaptively reordered in ascending order of the Merge candidate cost. More specifically, the template matching cost of the Merge candidates in all subgroups except the last subgroup is calculated; then, the Merge candidates in their own subgroups are reordered except for the last subgroup; finally, the final Merge candidate list is obtained.
[0428] For the decoder, after the Merge candidate list is constructed, as Figure 41 shown, some / none of the Merge candidates are adaptively reordered in ascending order of the Merge candidate cost. In Figure 41 , the subgroup where the selected (for signal transmission) Merge candidate is located is called the selected subgroup.
[0429] More specifically, if the selected Merge candidate is in the last subgroup, the Merge candidate list construction process is terminated after the selected Merge candidate is derived, no reordering is performed, and the Merge candidate list remains unchanged; otherwise, the process is as follows:
[0430] After all the Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; the template matching cost of the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained.
[0431] For both the encoder and the decoder, the template matching cost is derived as a function of T and RT, where T is a set of sample points in the template and RT is a set of reference points for the template.
[0432] When deriving the reference points of the template of a Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.
[0433] The reference points (RT) of the template for bidirectional prediction are derived by weighted averaging the reference points (RT 0 ) of the template in reference list 0 and the reference points (RT 1 ) of the template in reference list 1. RT = ((8 - w) * RT 0 + w * RT 1 + 4) >> 3 (2 - 33)
[0434] The weights (8 - w) of the reference templates in reference list 0 and the weight (w) of the reference templates in reference list 1 are determined by the BCW index of the Merge candidate. The BCW indices equal to {0, 1, 2, 3, 4} correspond to w equal to {-2, 3, 4, 5, 10} respectively.
[0435] If the local illumination compensation (LIC) flag of the Merge candidate is true, the reference sample points of the template are derived using the LIC method.
[0436] The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT.
[0437] The template size is 1. This means that the width of the left template and / or the height of the upper template is 1.
[0438] If the codec mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered.
[0439] If the codec mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered.
[0440] 2.28 Geometric prediction mode with motion vector difference
[0441] In the geometric prediction mode with motion vector difference (GMVD), each geometric partition in GPM can decide whether to use GMVD. If GMVD is selected for a geometric region, the MV of that region is calculated as the sum of the MV of the Merge candidate and the MVD. All other processing remains the same as in GPM.
[0442] Using GMVD, the MVD is transmitted in the form of a direction and distance pair by signal. Nine candidate distances (1 / 4 - pixel, 1 / 2 - pixel, 1 - pixel, 2 - pixel, 3 - pixel, 4 - pixel, 6 - pixel, 8 - pixel, 16 - pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) are included. Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD in GMVD is also left - shifted by 2 bits as in MMVD.
[0443] 2.29 Affine MMVD
[0444] In affine MMVD, the affine Merge candidate (referred to as the base affine Merge candidate) is selected, and the MVs of the control points are further refined by the MVD information transmitted by signal. The MVD information of all control points' MVs is the same in one prediction direction.
[0445] When the starting MV is a bi - directional prediction MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture while the other reference POC is less than the POC of the current picture), the MV offsets added to the list 0 MV component of the starting MV and the MV offset of the list 1 MV have opposite values; otherwise, when the starting MV is a bi - directional prediction MV and both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the MV offsets added to the list 0 MV component of the starting MV and the MV offset of the list 1 MV are the same.
[0446] 2.30 Adaptive Decoder - side Motion Vector Refinement (ADMVR)
[0447] In ECM - 2.0, if the selected Merge candidate satisfies the DMVR condition, the multi - pass decoder - side motion vector refinement (DMVR) method is applied in the regular Merge mode. In the first pass, bilateral matching (BM) is applied to the coded - decoded block. In the second pass, BM is applied to each 16x16 sub - block within the coded - decoded block. In the third pass, the MVs in each 8x8 sub - block are refined by applying bidirectional optical flow (BDOF).
[0448] The adaptive decoder - side motion vector refinement method consists of two new Merge modes that are introduced to refine the MV only in one direction (L0 or L1) of the bi - directional prediction of the Merge candidates that satisfy the DMVR condition. The multi - pass DMVR process is applied to the selected Merge candidate to refine the motion vector. However, in the first pass (i.e., PU - level) DMVR, MVD0 or MVD1 is set to zero.
[0449] Similar to the regular Merge mode, the Merge candidates for the proposed Merge modes are derived from spatially neighboring coded - decoded blocks, TMVP, non - adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The same Merge candidate list (i.e., the ADMVR Merge list) is used by the two proposed Merge modes, and the Merge index is coded - decoded in the regular Merge mode.
[0450] 2.31. IBC with Template Matching
[0451] It is proposed to use template matching with IBC for both the IBC Merge mode and the IBC AMVP mode.
[0452] Compared with the Merge list used by the conventional IBC Merge mode, the IBC-TM Merge list has been modified such that candidates are selected according to a deduplication method that uses the motion distance between candidates as in the conventional TM Merge mode. The trailing zero motion fills (which is meaningless for intra coding / decoding) has been replaced by motion vectors (-W, 0), (0, -H), and (-W, -H) pointing to the left CU, top CU, and top-left CU, and then, if necessary, the list is filled with left motion vectors without performing deduplication.
[0453] In the IBC-TM Merge mode, before the RDO or decoding process, the selected candidates are refined using a template matching method. The IBC-TM Merge mode has been made competitive with the conventional IBC Merge mode, and the TM Merge flag is signaled.
[0454] In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC Merge list. Each of these 3 selected candidates is refined using a template matching method and sorted according to its resulting template matching cost. Then usually only the first 2 are considered during the motion estimation process.
[0455] The template matching refinement for both the IBC-TM Merge mode and the AMVP mode is very simple because the IBC motion vectors are constrained to be integers and within the reference region, as Figure 42 shown. So, in the IBC-TM Merge mode, all refinements are performed with integer precision, and in the IBC-TM AMVP mode, it is performed with integer or 4-pixel precision. In both cases, the refined motion vectors in each refinement step must comply with the constraints of the reference region.
[0456] 2.32 IBC Merge Mode with Block Vector Difference
[0457] The IBC Merge mode with block vector difference is as follows.
[0458] The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions.
[0459] The basic candidates are selected from the top five candidates in the IBC Merge list reordering. And for all possible MBVD refinement positions (20x4) of each base candidate, they are reordered based on the SAD cost between the template (one row above and one column to the left of the current block) and its reference for each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are kept as available positions and thus used for MBVD index encoding and decoding.
[0460] 2.33. Reconstruction-Reordered IBC (RR-IBC)
[0461] Screen content coding and decoding tools similar to Intra Block Copy (IBC) generate prediction blocks by directly copying previously coded reference regions in the same picture. Symmetry is often observed in video content, especially in text character regions and computer-generated graphics in screen content sequences, such as Figure 43 shown. Therefore, specific screen content coding and decoding tools that consider symmetry will effectively compress such video content.
[0462] The Reconstruction-Reordered IBC (RR-IBC) mode for screen content video coding and decoding is proposed. When it is applied, the samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is derived without flipping. On the decoder side, the reconstructed block is flipped back to restore the original block.
[0463] For blocks coded and decoded by RR-IBC, two flipping methods (horizontal flip and vertical flip) are supported. First, for blocks coded and decoded by IBC AMVP, syntax flags are signaled to indicate whether the reconstruction is flipped, and if it is flipped, another flag specifying the flip type is further signaled. For IBC Merge, the flip type is inherited from neighboring blocks and no syntax is signaled. Considering horizontal symmetry or vertical symmetry, the current block and the reference block are usually horizontally aligned or vertically aligned. Therefore, when the horizontal flip is applied, the vertical component of the BV is not signaled and is presumed to be equal to 0. Similarly, when the vertical flip is applied, the horizontal component of the BV is not signaled and is presumed to be equal to 0.
[0464] To better utilize the symmetry property, a flip-aware BV adjustment method is applied to refine block vector candidates. For example, as Figure 44A - Figure 44B shown, (x nbr , y nbr ) and (x cur , y cur ) represent the coordinates of the central samples of the neighboring block and the current block respectively, BV nbr and BV curRepresent the BV of the neighboring block and the current block respectively. Instead of directly inheriting the BV from the neighboring block, the BV cur 's horizontal component, in the case where the neighboring block is encoded and decoded with horizontal flipping, is calculated by adding the motion displacement to the BV nbr 's horizontal component (denoted as BV nbr h ), that is, BV cur h = 2(x nbr - x cur ) + BV nbr h . Similarly, the vertical component of BV cur , in the case where the neighboring block is encoded and decoded with vertical flipping, is calculated by adding the motion displacement to the BV nbr 's vertical component (denoted as BV nbr v ), that is, BV cur v = 2(y nbr - y cur ) + BV nbr v .
[0465] 2.34. Intra-frame template matching
[0466] Intra-frame template matching prediction (Intra TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed part of the current frame, and its L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode, and the same prediction operation is performed on the decoder side.
[0467] The prediction signal is generated by matching the L-shaped causal neighbor of the current block with Figure 45 another block in the predefined search area in
[0468] R1: The current CTU
[0469] R2: The upper left of the CTU
[0470] R3: Above the CTU
[0471] R4: To the left of the CTU.
[0472] SAD is used as the cost function.
[0473] Within each region, the decoder searches for the template with the minimum SAD relative to the current template and uses its corresponding block as the prediction block.
[0474] The sizes of all regions (SearchRange_w, SearchRange_h) are set proportionally to the block sizes (BlkW, BlkH) so that each pixel has a fixed number of SAD comparisons. That is: SearchRange_w = a * BlkW SearchRange_h = a * BlkH
[0475] where "a" is a constant that controls the gain / complexity trade-off. In fact, "a" is equal to 5. The intra-template matching tool is enabled for CUs with dimensions less than or equal to 64 in width and height. The maximum CU size for intra-template matching is configurable.
[0476] When DIMD is not used for the current CU, the intra-template matching prediction mode is signaled at the CU level via a dedicated flag.
[0477] 3. Problem
[0478] In the current design of IBC, the entire block is copied from the reconstructed region in the current picture. The content in the block can come from two or more different objects, and the way of obtaining the prediction of the entire block is not feasible for this case.
[0479] 4. Detailed Solution
[0480] The following embodiments should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow way. In addition, these embodiments can be combined in any way.
[0481] In the present disclosure, intra-block copy (IBC) may not be limited to the current IBC technology, but can be interpreted as a technology in which the reference (or prediction) block is obtained using samples in the current strip / slice / sub-picture / picture / other video units (e.g., CTU rows), excluding conventional intra-prediction methods.
[0482] In the present disclosure, CIBCIP (or IBC-CIIP) may refer to a codec tool that combines intra-block copy (IBC) and intra-prediction. It is a codec tool that obtains the prediction of a block using both IBC and intra-prediction.
[0483] In the present disclosure, IBC-LIC may refer to a codec tool in which local illumination compensation is used to refine video units coded using IBC.
[0484] In the following discussion, the IBC may be replaced by other coding / decoding tools that rely on coded / decoded / reconstructed information (e.g., palette, intra-template matching) within the same region. Intra Block Copy with Geometric Partitioning 1. It is proposed that intra-block copy can be used to obtain the prediction of at least one sub-partition within a video unit when the video unit is geometrically divided into more than one sub-partition. The coding / decoding mode is denoted as IBC-GPM. a. Alternatively, it is proposed that when a video unit is predicted by more than one hypothesis, the IBC can be used to obtain the prediction of at least one hypothesis within the video unit, and the hypothesis predictions are geometrically weighted and summed to generate the final prediction. b. In one example, the way of dividing a video unit into sub-partitions can be the same as that of GPM or GPM-intra. i. Alternatively, the way of dividing a video unit into sub-partitions can be different from that of GPM or GPM-intra. 1) In one example, the geometric angle or geometric offset can be different. a) In one example, a geometric offset equal to 0 may not be allowed to be used in IBC-GPM. 2) In one example, for video units with different block sizes / dimensions, how to divide the video unit can be different. ii. In one example, one or more geometric partitioning modes used for GPM or GPM-intra may not be allowed to be used for IBC-GPM. iii. In one example, the maximum number of geometric partitioning modes can be predefined or signaled in the bitstream. iv. In one example, the sub-partitions adjacent to one or more sides of the video unit can be predicted by intra prediction. 1) In one example, the side can refer to the left / top / right / bottom side of the video unit. 2) In one example, the sub-partitions adjacent to one or more sides of the video unit may not be allowed to be predicted by intra prediction. a) In one example, the side can refer to the right / bottom side of the video unit. c. In one example, a video unit can be divided into N sub-partitions, where N is an integer greater than 1. i. In one example, N = 2. ii. In one example, when N = 2, the prediction signal of the first sub-partition is derived using the first method, and the prediction signal of the second sub-partition is derived using the second method. 1) In one example, the first (second) method may refer to IBC / palette / intra template matching, and the second (first) method may refer to IBC / palette / intra template matching / intra prediction / inter - prediction. 2) In one example, the two methods may be different. a) In one example, the first (second) method is IBC, and the second (first) method is intra prediction. iii. In one example, the determination of which method to use to derive the prediction signal for a sub - partition can be made by whether the signal is transmitted, or is predefined or derived in real - time. 1) In one example, one or more syntax elements can be used for the determination. a) In one example, when N = 2, the syntax element (ibc_gpm_intra_flag) is used to indicate whether intra prediction is used to derive the prediction signal for the first (second) sub - partition, where ibc_gpm_intra_flag being equal to X indicates that intra prediction is used. i. In one example, when ibc_gpm_intra_flag is equal to X for the first (second) sub - partition, a specific prediction method is used for the second (first) sub - partition, such as IBC. 2) In one example, the determination can depend on the coding - decoding information. a) In one example, the coding - decoding information can refer to the block size / dimension. b) In one example, the coding - decoding information can refer to the segmentation mode. d. In one example, when IBC is used to obtain a prediction, the IBC Merge mode and / or the IBC AMVP mode can be used. i. In one example, how to use IBC to obtain a prediction for at least one sub - partition can be the same as the way of obtaining the prediction for the entire block coded using IBC. ii. Alternatively, how to use IBC to obtain a prediction for at least one sub - partition or a hypothesis can be different from the way of obtaining the prediction for the entire block coded using IBC. 1) In one example, the maximum allowed number of Merge / AMVP candidates can be different. 2) In one example, the construction of the IBC Merge / AMVP candidate list can be different. 3) In one example, the IBC Merge candidate list may not be reordered. a) Alternatively, the IBC Merge candidate list can be reordered in a different way. iii. In one example, one or more specific modes may not be used to obtain a prediction. 1) In one example, a specific Merge mode may refer to the IBC TM Merge / AMVP mode, or the IBC Merge mode with BVD, or the RR-IBC Merge / AMVP mode, AMVR for IBC. e. In one example, intra prediction may be used to obtain predictions for one or more sub - partitions / hypotheses. i. In one example, intra prediction may refer to specific codec tools, such as conventional intra prediction, DIMD, TIMD, MRL, ISP, MIP, intra TMP, CCLM, MMLM, CCCM, GLM. ii. In one example, an intra prediction mode (IPM) candidate list is used, and one or more IPMs from the IPM candidate list may be used to obtain predictions for one or more sub - partitions. 1) In one example, the IPM candidate list may include one or more IPMs from the primary MPM list or the secondary MPM list. 2) In one example, the IPM candidate list may be constructed using TIMD and / or DIMD. 3) In one example, block vectors may be used to construct the IPM candidate list. a) For example, block vectors may be used by IBC to obtain predictions for one or more sub - partitions / hypotheses. 4) In one example, whether and / or how to construct the IPM candidate list may depend on codec information. a) In one example, codec information may refer to block size / dimension. b) In one example, codec information may refer to the partitioning mode that geometrically divides a block into sub - partitions. c) In one example, the size of the IPM candidate list may be the same for a block size / dimension or all partitioning modes. 5) In one example, the size of the IPM candidate list or the maximum number of IPMs used to derive intra prediction for sub - partitions may be predefined, or signaled, or derived in real - time. f. In one example, inter prediction may be used to obtain predictions for one or more sub - partitions / hypotheses. i. In one example, inter-frame prediction may refer to specific codec tools such as CIIP (e.g., CIIP-plane, CIIP-TIMD, CIIP-TM), BCW (e.g., BCW index derived from TM), MMVD (e.g., MMVD or TM-based reordering for MMVD), template matching (TM), IBC (e.g., IBC-TM, IBC with block vector difference, IBC with reconstruction reordering), affine (e.g., affine-MMVD, TM-based reordering for affine MMVD), DMVR / multi-pass DMVR, PROF, BDOF / sample-based BDOF, adaptive decoder-side motion vector refinement (ADMVR), OBMC or TM-based OBMC, MHP, GPM (e.g., GPM, GPM-TM, GPM-MMVD, GPM-intra), bilateral / template matching AMVP-Merge mode. 2. It is proposed to obtain the prediction of the first component (e.g., chrominance component) of a video unit in the same manner as the second component (e.g., luminance component). a. In one example, whether and how to obtain the prediction of the first component in the same manner as the second component may depend on whether a single tree or a dual tree is used. i. In one example, when a single tree is used, the prediction of the first component is obtained in the same manner as the second component. b. In one example, the prediction of the first component may be obtained using specific prediction methods other than IBC GPM (such as IBC or intra prediction or inter-frame prediction). 3. In one example, the prediction of the region along the geometric segmentation edge may be obtained by mixing the predictions of two sub-segmentations. a. In one example, the weights for mixing may be derived in real time or predefined. b. In one example, the weights for mixing may depend on the distance between the mixing samples and the geometric segmentation edge. c. In one example, the width of the mixing region may be signaled. i. In one example, the determination of the mixing region may depend on codec information such as block size. 4. In one example, IBC-GPM may not be allowed to be used with one or more specific codec tools. a. In one example, specific codec tools may refer to IBC AMVP mode, or IBC Merge mode, or IBC-TM mode, or IBC-MBVD mode, or RR-IBC mode, or IBC-LIC, or CIBCIP (IBC-CIIP) or AMVR for IBC. b. Alternatively, IBC-GPM can be used together with one or more of the above encoding / decoding tools. Store BV and IPM and MV 5. In one example, the encoding / decoding information of a video unit encoded / decoded using IBC-GPM can be stored and used by the following video units in the current picture and / or the video units in the following pictures. a. In one example, the encoding / decoding information can refer to block vectors and / or intra prediction modes and / or motion vectors. b. In one example, one or more BVs can be used to construct a Merge / AMVP candidate list for the following video units, or inserted into the historical BV cache for future reference. i. In one example, BVs in sub-partitions with a larger size than other sub-partitions can be used. ii. In one example, BVs in the first / last sub-partition can be used. iii. Alternatively, BVs in IBC-GPM are not allowed to be used. c. In one example, one or more IPMs can be used to construct an MPM list for the following video units, or stored in the IPM cache, or used for chrominance prediction. i. In one example, IPMs in sub-partitions with a larger size than other sub-partitions can be used. ii. In one example, IPMs in the first / last sub-partition can be used. iii. Alternatively, IPMs in IBC-GPM are not allowed to be used. 1) In one example, default IPMs can be used, such as planar or DC. d. In one example, one or more MVs can be used to construct a Merge / AMVP candidate list for the following video units, or stored in the motion information cache. i. In one example, MVs in sub-partitions with a larger size than other sub-partitions can be used. ii. In one example, MVs in the first / last sub-partition can be used. iii. Alternatively, MVs in IBC-GPM are not allowed to be used. 1) In one example, the motion information for the video unit is set to be equal to the IBC mode. Regarding enabling IBC - Control of GPM 6. The determination of whether a block is allowed to be encoded / decoded using the IBC-GPM mode can depend on the encoding / decoding information. a. In one example, the codec information may indicate whether IBC (Merge and / or AMVP) is allowed. b. In one example, the codec information may indicate the block dimension and / or block size. i. In one example, when the block size (W×H) is less than or equal to a threshold (T), the block is allowed to be coded / decoded using IBC-GPM, where W and H represent the block width and block height respectively. ii. In one example, T = 256, or 512, or 1024, or 2048, or 4096. c. In one example, the codec information may indicate the depth of the block. d. In one example, the codec information may indicate the block position, e.g., whether the current block is the first row / column of the CTU. e. In one example, the codec information may indicate the slice / picture type. i. In one example, IBC-GPM may be applied only to I slices / pictures. f. In one example, the codec information may indicate the information of the temporal layer (e.g., the temporal layer index). g. In one example, the codec information may indicate the information of the color component. 7. In one example, whether and / or how to apply IBC-GPM may depend on the color format and / or color component. a. In one example, IBC-GPM may be applied to all color components. b. In one example, when IBC-GPM is applied to the chrominance component, the derivation of intra prediction may be different from the derivation of intra prediction for the luminance component. i. In one example, intra prediction may be obtained using CCLM, or MMLM, or CCCM, or chroma-DIMD, or chroma-TIMD, or a combination of CCLM / MMLM / CCCM and angular modes. c. In one example, whether and / or how to apply IBC-GPM to the first component may depend on whether IBC-GPM is applied to the second component. i. In one example, the first component may refer to the chrominance component (e.g., Cb and / or Cr), and the second component may refer to the luminance component (e.g., Y). ii. In one example, the way to apply IBC-GPM to the first component may be the same as that of the second component. 1) Alternatively, the way to apply IBC-GPM to the first component may be different from that of the second component. a) In one example, the weighting parameters may be different. d. In one example, IBC-GPM can be applied to the luminance component, but not to the chrominance component. i. In one example, the luminance component can refer to Y in the YCbCr color space or G in the RGB color space. ii. In one example, the chrominance component can refer to Cb and / or Cr in the YCbCr color space or R and / or B in the RGB color space. Signaling for IBC - GPM 8. The indication of IBC-GPM can be conditionally transmitted by a signal, where the conditions can include: a. Whether the IBC Merge / AMVP mode is allowed, b. Whether specific codec tools are allowed, such as IBC-TM, or IBC-MBVD, or RR-IBC, or IBC-LIC, or CIBCIP (IBC-CIIP), c. Block dimension and / or block size, d. The codec information can refer to the depth of the block, e. Slice / picture type and / or partition tree type (single tree or dual tree, or local dual tree), f. Block position, g. Color component. 9. Whether and how to apply IBC-GPM to a video unit can be transmitted by a signal in the bitstream. a. In one example, whether to enable IBC-GPM for a video unit can be transmitted by a signal using one or more syntax elements. i. In one example, a syntax element can be used to indicate whether IBC-GPM is applied to the video unit, and / or one or more syntax elements can be used to indicate whether IBC AMVP / Merge is used to obtain prediction. ii. In one example, a syntax element can be used to indicate whether the IBC Merge mode is used in IBC-GPM. iii. In one example, a syntax element can be used to indicate whether the IBC AMVP mode is used in IBC-GPM. b. In one example, how and whether to divide a video unit into more than one sub-partition can be predefined or derived or transmitted by a signal in the bitstream. i. In one example, a syntax element can be transmitted by a signal to indicate the way to divide the video unit. ii. In one example, one or more syntax elements can be used to indicate how to obtain the prediction for each sub-partition, such as IBC, and / or inter-frame prediction, and / or intra-frame prediction. 1) In one example, one or more syntax elements may be signaled to indicate the IPM used for intra prediction for sub - partitioning. 2) In one example, one or more syntax elements may be signaled to indicate the IBC AMVP index or the IBC Merge index to indicate candidates for IBC for sub - partitioning. c. In one example, a syntax element indicating which IPM of the IPM candidate list is used to obtain intra prediction for one or more sub - partitions may be signaled. d. In one example, how to use multiple sub - partitions for hybrid prediction may be signaled in the bitstream. i. In one example, a syntax element indicating the hybrid width may be signaled. e. In one example, the above - mentioned syntax elements may be binarized using fixed - length coding / decoding or truncated unary coding / decoding or unary coding / decoding or exponential - Golomb coding / decoding or a coded flag. i. In one example, the syntax element may be bypass - coded. ii. Alternatively, the syntax element may be context - coded. 1) The context may depend on coding information such as block dimension and / or block size, and / or slice / picture type, and / or information of neighboring blocks (adjacent or non - adjacent), and / or information of other coding tools used for the current block, and / or information of the temporal layer. f. In one example, the syntax element may be signaled before or after the indication of IBC - TM mode or IBC - MBVD mode or RR - IBC mode or IBC - LIC or CIBCIP (IBC - CIIP). i. In one example, whether to signal and / or how the syntax element may depend on whether the IBC mode, or IBC - TM mode, or IBC - MBVD mode, or RR - IBC mode, or IBC - LIC, or CIBCIP (IBC - CIIP) is enabled for the video unit. g. In one example, one or more syntax elements may be signaled at the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / picture group header. h. In one example, IBC - GPM and GPM may share the signaling method for at least one syntax element (such as binarization, signaling condition, and coding context). General Requirements 10. In the above examples, the video unit may refer to a color component / sub-picture / strip / slice / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transformation unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transformation block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel. 11. Whether and / or how to apply the methods disclosed above may be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as at the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 12. Whether and / or how to apply the methods disclosed above may be signaled at the PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / any other type of region containing more than one sample or pixel. 13. Whether and / or how to apply the methods disclosed above may depend on coding information, such as block size, color format, single / double tree splitting, color component, strip / picture type.
[0623] 5. Embodiments
[0624] 5.1 Embodiment 1
[0625] In this paper, three aspects are proposed to extend the use of IBC:
[0626] Aspect #1: Combined IBC and Intra Prediction (IBC-CIIP);
[0627] Aspect #2: IBC with Geometric Partitioning (IBC-GPM);
[0628] Aspect #3: IBC with Local Illumination Compensation (IBC-LIC).
[0629] Combined IBC and Intra Prediction (IBC-CIIP)
[0630] When IBC-CIIP is applied to a CU, two prediction signals are obtained using IBC and intra prediction. The two prediction signals are weighted and summed to generate the final prediction. IBC-CIIP can be applied to the IBC AMVP mode and the IBC Merge mode. The CU flag is signaled to indicate the use of IBC-CIIP.
[0631] IBC with Geometric Partitioning (IBC-GPM)
[0632] When IBC-GPM is applied to a CU, the CU is geometrically partitioned into two sub-partitions. The predicted signals of the two sub-partitions are generated using IBC and intra prediction. IBC-GPM can be applied to the IBC Merge mode. A CU flag is signaled to indicate the use of IBC-GPM.
[0633] IBC with local illumination compensation (IBC-LIC)
[0634] When IBC-LIC is applied to a CU, the local illumination change between the CU and its predicted block is modeled as a linear equation. The parameters of the linear equation are derived similar to those for LIC used in inter prediction. IBC-LIC can be applied to the IBC AMVP mode and the IBC Merge mode. For the IBC AMVP mode, an IBC-LIC flag is signaled to indicate the use of IBC-LIC. For the IBC Merge mode, the IBC-LIC flag is deduced from the Merge candidates.
[0635] As used herein, the term "video unit" or "video block" may be a sequence, picture, slice, tile, sub-picture, coding tree unit (CTU) / coding tree block (CTB), CTU / CTB row, one or more coding units (CU) / coding blocks (CB), one or more CTUs / CTBs, one or more virtual pipeline data units (VPDU), a sub-region within a picture / slice / tile. The term "reference line" may refer to the reconstructed samples of a row and / or column adjacent or non-adjacent to the current block, which are used to derive the intra prediction of the current video unit via an interpolation filter along a specific direction, and the specific direction is determined by an intra prediction mode (e.g., conventional intra prediction with an intra prediction mode), or to derive the intra prediction of the current video unit by weighting the reference samples of the reference line with a matrix or vector (e.g., MIP).
[0636] Figure 46 A flowchart of a method 4600 for video processing according to an embodiment of the present disclosure is shown. Method 4600 is implemented during the conversion between a video unit of a video and the bitstream of the video.
[0637] At block 4610, for the conversion between a video unit of a video and the bitstream of the video, the video unit is partitioned into a plurality of sub-partitions using a predefined method, wherein the video unit is coded using intra block copy (IBC)-geometric partitioning mode (GPM).
[0638] At block 4620, the prediction of at least one sub-partition of the video unit is obtained using intra block copy.
[0639] At block 4630, the transformation of performing prediction based on at least one sub - partition of a video unit is carried out. In some embodiments, the transformation may include encoding the video unit into a bitstream. Alternatively or additionally, the transformation may include decoding the video unit from the bitstream. This can improve the encoding and decoding efficiency and performance.
[0640] In some embodiments, the predefined method is the same as the method used for geometric partition mode (GPM). Alternatively, the predefined method is the same as the method used for GPM - Intra.
[0641] In some embodiments, the predefined method is different from the method used for GPM. Alternatively, the predefined method is different from the method used for GPM - Intra.
[0642] In some embodiments, the geometric angle associated with the predefined method is different from the geometric angle used for GPM or GPM - Intra. In some embodiments, the geometric offset associated with the predefined method is different from the geometric offset used for GPM or GPM - Intra. In some embodiments, a geometric offset equal to 0 is not allowed to be used in IBC - GPM.
[0643] In some embodiments, the predefined method of partitioning a video unit depends on at least one of the following: the block size or block dimension of the video unit. In one example, how to partition the video unit can be different for video units with different block sizes / dimensions. In some embodiments, at least one geometric partition mode used for GPM or GPM - Intra is not allowed to be used in IBC - GPM.
[0644] In some embodiments, the maximum number of geometric partition modes is predefined. Alternatively, the maximum number of geometric partition modes is indicated.
[0645] In some embodiments, the sub - partitions adjacent to at least one side of the video unit are predicted by intra - prediction. In some embodiments, at least one side includes at least one of the following: the left side of the video unit, the upper side of the video unit, the right side of the video unit, or the bottom side of the video unit.
[0646] In some embodiments, the sub - partitions adjacent to at least one side of the video unit are not predicted by intra - prediction. In some embodiments, at least one side includes at least one of the following: the right side of the video unit or the bottom side of the video unit.
[0647] In some embodiments, the plurality of sub - partitions includes N sub - partitions, where N is an integer greater than 1. In some embodiments, N is equal to 2.
[0648] In some embodiments, N equals 2, the prediction signal of the first sub - partition of the video unit is derived using a first method, and the prediction signal of the second sub - partition of the video unit is derived using a second method. In some embodiments, one of the first method and the second method (e.g., the first or the second) includes one of the following: IBC, palette, or intra - template matching, and one of the first method and the second method (e.g., the second or the first) includes one of the following: IBC, palette, intra - template matching, intra - prediction, or inter - prediction.
[0649] In some embodiments, the first method and the second method are different. In some embodiments, the first method is IBC, and the second method is intra - prediction. Alternatively, the second method is IBC, and the first method is intra - prediction.
[0650] In some embodiments, method 4600 further includes: determining which method will be used to derive the prediction signal for at least one sub - partition. In some embodiments, the determination is indicated, or the determination is predefined, or the determination is derived in real - time. In one example, the determination of which method to use to derive the prediction signal for the sub - partition can be transmitted by a signal, or be predefined, or be derived in real - time.
[0651] In some embodiments, at least one syntax element is used for the determination. In some embodiments, if N equals 2, the syntax element is used to indicate whether intra - prediction is used to derive the prediction signal of one sub - partition among multiple sub - partitions, where the number of syntax elements equal to a predefined number indicates that intra - prediction is used. In one example, when N = 2, the syntax element (ibc_gpm_intra_flag) is used to indicate whether intra - prediction is used to derive the prediction signal of the first (second) sub - partition, where ibc_gpm_intra_flag equal to X indicates that intra - prediction is used. In some embodiments, the syntax element equals a predefined number for one of the multiple sub - partitions, and a specific prediction method is used for the other of the multiple sub - partitions. In one example, when ibc_gpm_intra_flag equals X for the first (second) sub - partition, a specific prediction method, such as IBC, is used for the second (first) sub - partition.
[0652] In some embodiments, the determination depends on the codec information. In some embodiments, the codec information includes at least one of the following: block size, block dimension, or partition mode.
[0653] In some embodiments, whether to construct an intra - prediction mode (IPM) candidate list and / or the method of constructing the IPM candidate list depends on the codec information. In some embodiments, the codec information includes at least one of the following: block size, block dimension, or partition mode, which geometrically divides the video unit into multiple sub - partitions.
[0654] In some embodiments, the size of the IPM candidate list is the same for different block sizes or different block dimensions. Alternatively, the size of the IPM candidate list is the same for all partitioning modes.
[0655] In some embodiments, the size of the IPM candidate list or the maximum number of IPMs used to derive intra prediction for sub - partitioning is predefined. Alternatively, the size of the IPM candidate list or the maximum number of IPMs used to derive intra prediction for sub - partitioning is indicated. In some embodiments, the size of the IPM candidate list or the maximum number of IPMs used to derive intra prediction for sub - partitioning is derived in real - time.
[0656] In some embodiments, IBC - GPM is not allowed to be used with one or more coding tools. Alternatively, IBC - GPM is used with one or more coding tools. In some embodiments, one or more coding tools include at least one of the following: IBC Advanced Motion Vector Prediction (AMVP) mode and IBC Merge mode, IBC Template Matching (IBC - TM) mode, IBC Merge mode with Block Vector Difference (IBC - MBVD) mode, Reconstruction Re - ordering IBC (RR - IBC) Merge mode, IBC with Local Illumination Compensation (IBC - LIC) mode, IBC - CIIP mode, or Adaptive Motion Vector Resolution (AMVR) for IBC.
[0657] In some embodiments, the determination of whether a video unit is allowed to be coded with the IBC - GPM mode depends on coding information. In some embodiments, the coding information includes at least one of the following: block dimension or block size.
[0658] In some embodiments, if the block size represented as W×H is less than or equal to a threshold, the video unit is allowed to be coded with the IBC - GPM mode, where W and H represent the block width and block height respectively. In some embodiments, the threshold is one of the following: 256, 512, 1024, 2048, or 4096.
[0659] In some embodiments, the coding information includes at least one of the following: slice type or picture type. In some embodiments, IBC - GPM is only applied to I - slices or I - pictures.
[0660] In some embodiments, whether to apply IBC-GPM and / or the method of applying IBC-GPM depends on at least one of the following: color format or color component. In some embodiments, IBC-GPM is applied to all color components. In some embodiments, IBC-GPM is applied to the chrominance component, and the derivation of intra prediction is different from the derivation of intra prediction for the luminance component.
[0661] In some embodiments, intra prediction is obtained using at least one of the following: Cross-Component Linear Mode (CCLM) mode, Multi-Mode Linear Mode (MMLM), Convolutional Cross-Component Model (CCCM), Decoder-Side Intra Mode Derivation (DIMD) for chrominance, Template-Based Intra Mode Derivation (TIMD) for chrominance, a combination of CCLM and angular modes, a combination of MMLM and angular modes, or a combination of CCCM and angular modes.
[0662] In some embodiments, whether to apply IBC-GPM to the first component and / or the method of applying IBC-GPM to the first component depends on whether IBC-GPM is applied to the second component. In some embodiments, the first component is the chrominance component and the second component is the luminance component. In some embodiments, the method of applying IBC-GPM to the first component is the same as that for the second component. In some embodiments, the method of applying IBC-GPM to the first component is different from that for the second component. In some embodiments, the weighting parameters are different.
[0663] In some embodiments, IBC-GPM is applied to the luminance component and not to the chrominance component. In some embodiments, the luminance component includes Y in the YCbCr color space. Alternatively, the luminance component includes G in the Red-Green-Blue (RGB) color space.
[0664] In some embodiments, the chrominance component includes at least one of Cb or Cr in the YCbCr color space. Alternatively, the chrominance component includes at least one of the following: R or B in the RGB color space.
[0665] In some embodiments, the indication of IBC-GPM is signaled based on a condition. The condition may include whether a codec tool is allowed. In some embodiments, the codec tool includes at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC mode, or IBC-CIIP mode.
[0666] In some embodiments, a method of whether to apply IBC-GPM to a video unit and / or apply IBC-GPM to a video unit is indicated in a bitstream. In some embodiments, how and whether to divide a video unit into multiple sub-partitions is predefined. Alternatively, how and whether to divide a video unit into multiple sub-partitions is derived. In some embodiments, how and whether to divide a video unit into multiple sub-partitions is signaled in a bitstream.
[0667] In some embodiments, one or more syntax elements are signaled to indicate an intra prediction mode (IPM) used for intra prediction in sub-partitioning. In some embodiments, one or more syntax elements are signaled to indicate an IBC AMVP index or an IBC Merge index to indicate candidate predictions used for IBC in sub-partitioning.
[0668] In some embodiments, syntax elements are signaled before an indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-CIIP. Alternatively, syntax elements are signaled after an indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-CIIP.
[0669] In some embodiments, whether to signal a syntax element and / or a method of signaling a syntax element depends on whether at least one of the following is enabled for a video unit: IBC mode, IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-CIIP.
[0670] In some embodiments, a video unit includes at least one of the following: a color component, a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a codec unit (CU), a codec tree unit (CTU), a CTU row, a group of CTUs, a slice, a picture, a sub-picture, a block, a sub-region within a block, or a region containing more than one sample or pixel.
[0671] In some embodiments, an indication of whether to obtain a prediction of at least one sub-partition of a video unit by using IBC and / or how to obtain a prediction of at least one sub-partition of a video unit by using IBC is indicated at one of the following: sequence level, picture group level, picture level, slice level, or slice group level.
[0672] In some embodiments, an indication of whether to obtain prediction of at least one sub - partition of a video unit by using IBC and / or how to obtain prediction of at least one sub - partition of a video unit by using IBC is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header or slice group header.
[0673] In some embodiments, an indication of whether to obtain prediction of at least one sub - partition of a video unit by using IBC and / or how to obtain prediction of at least one sub - partition of a video unit by using IBC is included in one of the following: prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, picture, sub - picture or a region containing more than one sample or pixel.
[0674] In some embodiments, method 4600 further includes: determining, based on the codec information of the video unit, whether to obtain prediction of at least one sub - partition of the video unit by using IBC and / or how to obtain prediction of at least one sub - partition of the video unit by using IBC, where the codec information includes at least one of the following: block size, color format, single and / or dual - tree partitioning, color component, slice type or picture type.
[0675] According to further embodiments of the present disclosure, a non - transitory computer - readable recording medium is provided. The non - transitory computer - readable recording medium stores a bitstream generated by a method executed by a device for video processing. The method includes: dividing a video unit of a video into a plurality of sub - partitions by using a predefined method, where the video unit is coded and decoded by using intra - block copy (IBC) - geometric partitioning mode (GPM); obtaining prediction of at least one sub - partition of the video unit by using intra - block copy; and generating a bitstream based on the prediction of at least one sub - partition of the video unit.
[0676] According to still further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: dividing a video unit of a video into a plurality of sub - partitions by using a predefined method, where the video unit is coded and decoded by using intra - block copy (IBC) - geometric partitioning mode (GPM); obtaining prediction of at least one sub - partition of the video unit by using intra - block copy; generating a bitstream based on the prediction of at least one sub - partition of the video unit; and storing the bitstream in a non - transitory computer - readable recording medium.
[0677] Embodiments of the present disclosure can be described according to the following items, and the features can be combined in any reasonable manner.
[0678] Item 1. A method for video processing, comprising: for the conversion between a video unit of a video and a bitstream of the video, dividing the video unit into a plurality of sub-divisions using a predefined method, wherein the video unit is encoded and decoded using an Intra Block Copy (IBC)-Geometry Partition Mode (GPM); obtaining a prediction of at least one sub-division of the video unit using Intra Block Copy; and performing the conversion based on the prediction of the at least one sub-division of the video unit.
[0679] Item 2. The method according to Item 1, wherein the predefined method is the same as the method used for the Geometry Partition Mode (GPM), or wherein the predefined method is the same as the method used for GPM-Intra.
[0680] Item 3. The method according to Item 1, wherein the predefined method is different from the method used for GPM, or wherein the predefined method is different from the method used for GPM-Intra.
[0681] Item 4. The method according to Item 3, wherein the geometric angle associated with the predefined method is different from the geometric angle used for the GPM or the GPM-Intra.
[0682] Item 5. The method according to Item 3, wherein the geometric offset associated with the predefined method is different from the geometric offset used for the GPM or the GPM-Intra.
[0683] Item 6. The method according to Item 5, wherein a geometric offset equal to 0 is not allowed to be used in the IBC-GPM.
[0684] Item 7. The method according to Item 3, wherein the predefined method for dividing the video unit depends on at least one of the following: the block size or block dimension of the video unit.
[0685] Item 8. The method according to Item 1, wherein at least one geometry partition mode used for GPM or GPM-Intra is not allowed to be used for IBC-GPM.
[0686] Item 9. The method according to Item 1, wherein the maximum number of geometry partition modes is predefined, or wherein the maximum number of the geometry partition modes is indicated.
[0687] Item 10. The method according to Item 1, wherein sub-divisions adjacent to at least one side of the video unit are predicted by intra prediction.
[0688] Item 11. The method according to Item 10, wherein the at least one side includes at least one of the following: the left side of the video unit, the upper side of the video unit, the right side of the video unit, or the bottom side of the video unit.
[0689] Item 12. The method according to Item 1, wherein sub - partitions adjacent to at least one side of the video unit are not predicted by intra - prediction.
[0690] Item 13. The method according to Item 12, wherein the at least one side includes at least one of the following: the right side of the video unit, or the bottom side of the video unit.
[0691] Item 14. The method according to Item 1, wherein the plurality of sub - partitions includes N sub - partitions, where N is an integer greater than 1.
[0692] Item 15. The method according to Item 14, wherein N is equal to 2.
[0693] Item 16. The method according to Item 14, wherein N is equal to 2, the prediction signal of the first sub - partition of the video unit is derived using a first method, and the prediction signal of the second sub - partition of the video unit is derived using a second method.
[0694] Item 17. The method according to Item 16, wherein one of the first method and the second method includes one of the following: IBC, palette, or intra - template matching, and wherein the other of the first method and the second method includes one of the following: IBC, palette, intra - template matching, intra - prediction, or inter - prediction.
[0695] Item 18. The method according to Item 16, wherein the first method and the second method are different.
[0696] Item 19. The method according to Item 18, wherein the first method is IBC and the second method is intra - prediction, or wherein the second method is IBC and the first method is intra - prediction.
[0697] Item 20. The method according to Item 14, further comprising: determining which method will be used to derive the prediction signal for the at least one sub - partition.
[0698] Item 21. The method according to Item 20, wherein the determination is indicated, or wherein the determination is predefined, or wherein the determination is derived in real - time.
[0699] Item 22. The method according to Item 20, wherein at least one syntax element is used for the determination.
[0700] Item 23. The method according to Item 20, wherein if N is equal to 2, a syntax element is used to indicate whether intra prediction is used to derive the prediction signal for one of the plurality of sub - partitions, and the syntax elements equal to a predefined number indicate that the intra prediction is used.
[0701] Item 24. The method according to Item 23, wherein the syntax element is equal to the predefined number for one of the plurality of sub - partitions, and a specified prediction method is used for another of the plurality of sub - partitions.
[0702] Item 25. The method according to Item 20, wherein the determination depends on codec information.
[0703] Item 26. The method according to Item 25, wherein the codec information includes at least one of the following: block size, block dimension, or partition mode.
[0704] Item 27. The method according to Item 1, wherein whether to construct an intra prediction mode (IPM) candidate list and / or the method of constructing the IPM candidate list depends on codec information.
[0705] Item 28. The method according to Item 27, wherein the codec information includes at least one of the following: block size, block dimension, or partition mode, and the partition mode geometrically divides the video unit into the plurality of sub - partitions.
[0706] Item 29. The method according to Item 27, wherein the size of the IPM candidate list is the same for different block sizes or different block dimensions, or wherein the size of the IPM candidate list is the same for all partition modes.
[0707] Item 30. The method according to Item 27, wherein the size of the IPM candidate list or the maximum number of IPMs used to derive the intra prediction for the sub - partition is predefined, or wherein the size of the IPM candidate list or the maximum number of IPMs used to derive the intra prediction for the sub - partition is indicated, or wherein the size of the IPM candidate list or the maximum number of IPMs used to derive the intra prediction for the sub - partition is derived in real - time.
[0708] Item 31. The method according to Item 1, wherein IBC - GPM is not allowed to be used with one or more codec tools, or wherein the IBC - GPM is used with the one or more codec tools.
[0709] Item 32. The method according to Item 31, wherein the one or more codec tools include at least one of the following: IBC Advanced Motion Vector Prediction (AMVP) mode, and IBC Merge mode, IBC Template Matching (IBC-TM) mode, IBC Merge mode with Block Vector Difference (IBC-MBVD) mode, Reconstruction-Reordering IBC (RR-IBC) Merge mode, IBC with Local Illumination Compensation (IBC-LIC) mode, IBC-CIIP mode, or Adaptive Motion Vector Resolution (AMVR) for IBC.
[0710] Item 33. The method according to Item 1, wherein the determination of whether the video unit is allowed to be coded / decoded using the IBC-GPM mode depends on coding / decoding information.
[0711] Item 34. The method according to Item 33, wherein the coding / decoding information includes at least one of the following: block dimension or block size.
[0712] Item 35. The method according to Item 34, wherein if the block size represented as W×H is less than or equal to a threshold, the video unit is allowed to be coded / decoded using the IBC-GPM mode, where W and H represent the block width and block height, respectively.
[0713] Item 36. The method according to Item 35, wherein the threshold is one of the following: 256, 512, 1024, 2048, or 4096.
[0714] Item 37. The method according to Item 33, wherein the coding / decoding information includes at least one of the following: slice type or picture type.
[0715] Item 38. The method according to Item 37, wherein IBC-GPM is only applied to I slices or I pictures.
[0716] Item 39. The method according to Item 1, wherein whether to apply IBC-GPM and / or the method of applying IBC-GPM depends on at least one of the following: color format or color component.
[0717] Item 40. The method according to Item 39, wherein IBC-GPM is applied to all color components.
[0718] Item 41. The method according to Item 39, wherein IBC-GPM is applied to the chrominance components, and the derivation of intra prediction is different from the derivation of intra prediction for the luminance component.
[0719] Item 42. The method according to Item 41, wherein the intra prediction is obtained using at least one of the following: cross-component linear mode (CCLM) mode, multi-mode linear mode (MMLM), convolutional cross-component model (CCCM), chrominance decoder-side intra mode derivation (DIMD), template-based intra mode derivation for chrominance (TIMD), a combination of CCLM and angular mode, a combination of MMLM and angular mode, or a combination of CCCM and angular mode.
[0720] Item 43. The method according to Item 1, wherein whether to apply IBC-GPM to the first component and / or the method of applying IBC-GPM to the first component depends on whether to apply IBC-GPM to the second component.
[0721] Item 44. The method according to Item 43, wherein the first component is a chrominance component and the second component is a luminance component.
[0722] Item 45. The method according to Item 43, wherein the method of applying IBC-GPM to the first component is the same as that of the second component.
[0723] Item 46. The method according to Item 43, wherein the method of applying IBC-GPM to the first component is different from that of the second component.
[0724] Item 47. The method according to Item 46, wherein the weighting parameters are different.
[0725] Item 48. The method according to Item 39, wherein IBC-GPM is applied to the luminance component and not to the chrominance component.
[0726] Item 49. The method according to Item 48, wherein the luminance component includes Y in the YCbCr color space, or wherein the luminance component includes G in the red-green-blue (RGB) color space.
[0727] Item 50. The method according to Item 48, wherein the chrominance component includes at least one of Cb or Cr in the YCbCr color space, or wherein the chrominance component includes at least one of the following: R or B in the RGB color space.
[0728] Item 51. The method according to Item 1, wherein the indication of IBC-GPM is signaled based on a condition, and the condition includes whether the codec tool is allowed.
[0729] Item 52. The method according to Item 51, wherein the codec tool comprises at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC mode, or IBC-CIIP mode.
[0730] Item 53. The method according to Item 1, wherein whether the IBC-GPM is applied to the video unit and / or the method of applying the IBC-GPM to the video unit is indicated in the bitstream.
[0731] Item 54. The method according to Item 53, wherein how and whether the video unit is divided into the plurality of sub-partitions is predefined, or wherein how and whether the video unit is divided into the plurality of sub-partitions is derived, or wherein how and whether the video unit is divided into the plurality of sub-partitions is signaled in the bitstream.
[0732] Item 55. The method according to Item 54, wherein one or more syntax elements are signaled to indicate the intra prediction mode (IPM) used for intra prediction in the sub-partition.
[0733] Item 56. The method according to Item 54, wherein one or more syntax elements are signaled to indicate the IBC AMVP index or the IBC Merge index to indicate candidate predictions in the IBC for the sub-partition.
[0734] Item 57. The method according to Item 53, wherein the syntax element is signaled before the indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC or IBC-CIIP, or wherein the syntax element is signaled after the indication of at least one of the following: the IBC-TM mode, the IBC-MBVD mode, the RR-IBC mode, the IBC-LIC or the IBC-CIIP.
[0735] Item 58. The method according to Item 57, wherein whether the syntax element is signaled and / or the method of signaling the syntax element depends on whether at least one of the following is enabled for the video unit: the IBC mode, the IBC-TM mode, the IBC-MBVD mode, the RR-IBC mode, the IBC-LIC, or the IBC-CIIP.
[0736] Item 59. The method according to any one of Items 1 to 58, wherein the video unit comprises at least one of the following: a color component, a prediction block (PB), a transform block (TB), a coding / decoding block (CB), a prediction unit (PU), a transform unit (TU), a coding / decoding tree block (CTB), a coding / decoding unit (CU), a coding / decoding tree unit (CTU), a CTU row, a CTU group, a stripe, a slice, a sub-picture, a block, a sub-region within the block, or a region containing more than one sample or pixel.
[0737] Item 60. The method according to any one of Items 1 to 58, wherein an indication of whether to obtain the prediction of at least one sub-division of the video unit by using IBC and / or how to obtain the prediction of at least one sub-division of the video unit by using IBC is indicated at one of the following: sequence level, picture group level, picture level, stripe level, or slice group level.
[0738] Item 61. The method according to any one of Items 1 to 58, wherein an indication of whether to obtain the prediction of at least one sub-division of the video unit by using IBC and / or how to obtain the prediction of at least one sub-division of the video unit by using IBC is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), stripe header, or slice group header.
[0739] Item 62. The method according to any one of Items 1 to 58, wherein an indication of whether to obtain the prediction of at least one sub-division of the video unit by using IBC and / or how to obtain the prediction of at least one sub-division of the video unit by using IBC is included in one of the following: prediction block (PB), transform block (TB), coding / decoding block (CB), prediction unit (PU), transform unit (TU), coding / decoding unit (CU), virtual pipeline data unit (VPDU), coding / decoding tree unit (CTU), CTU row, stripe, slice, sub-picture, or a region containing more than one sample or pixel.
[0740] Item 63. The method according to any one of Items 1 to 58, further comprising: determining whether to obtain the prediction of at least one sub-division of the video unit by using IBC and / or how to obtain the prediction of at least one sub-division of the video unit by using IBC based on the coding / decoding information of the video unit, the coding / decoding information comprising at least one of the following: block size, color format, single and / or dual-tree segmentation, color component, stripe type, or picture type.
[0741] Item 64. The method according to any one of Items 1 to 63, wherein the conversion includes encoding the video unit into the bitstream.
[0742] Item 65. The method according to any one of Items 1 to 63, wherein the conversion includes decoding the video unit from the bitstream.
[0743] Item 66. An apparatus for video processing, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of Items 1 - 65.
[0744] Item 67. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of Items 1 - 65.
[0745] Item 68. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing, wherein the method includes: dividing video units of the video into a plurality of sub-divisions using a predefined method, wherein the video units are encoded and decoded using an Intra Block Copy (IBC)-Geometry Partitioning Mode (GPM); obtaining a prediction of at least one sub-division of the video unit using Intra Block Copy; and generating the bitstream based on the prediction of the at least one sub-division of the video unit.
[0746] Item 69. A method for storing a bitstream of a video, including: dividing video units of the video into a plurality of sub-divisions using a predefined method, wherein the video units are encoded and decoded using an Intra Block Copy (IBC)-Geometry Partitioning Mode (GPM); obtaining a prediction of at least one sub-division of the video unit using Intra Block Copy; generating the bitstream based on the prediction of the at least one sub-division of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0747] Example device
[0748] Figure 47 A block diagram of a computing device 4700 in which various embodiments of the present disclosure can be implemented is shown. The computing device 4700 can be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or can be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).
[0749] It should be understood that Figure 47The computing device 4700 shown is for illustrative purposes only and does not imply any limitation, in any way, to the functionality and scope of the embodiments of the present disclosure.
[0750] As Figure 47 shown, the computing device 4700 includes a general-purpose computing device 4700. The computing device 4700 may include at least one or more processors or processing units 4710, a memory 4720, a storage unit 4730, one or more communication units 4740, one or more input devices 4750, and one or more output devices 4760.
[0751] In some embodiments, the computing device 4700 may be implemented as any user terminal or server terminal with computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is conceivable that the computing device 4700 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0752] The processing unit 4710 may be a physical processor or a virtual processor and may implement various processes based on the programs stored in the memory 4720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 4700. The processing unit 4710 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0753] The computing device 4700 generally includes various computer storage media. Such media can be any media accessible by the computing device 4700, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4720 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 4730 can be any removable or non-removable media and can include machine-readable media such as memory, flash drives, magnetic disks, or other media that can be used to store information and / or data and can be accessed in the computing device 4700.
[0754] The computing device 4700 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 47 , a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk can be provided. In such a case, each drive can be connected to a bus (not shown) via one or more data media interfaces.
[0755] The communication unit 4740 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 4700 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Thus, the computing device 4700 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0756] The input device 4750 can be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 4760 can be one or more of various output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 4740, the computing device 4700 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 4700 can also communicate with one or more devices that enable a user to interact with the computing device 4700, or if needed, the computing device 4700 can also communicate with any device (such as a network card, modem, etc.) that enables the computing device 4700 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0757] In some embodiments, some or all components of computing device 4700 may also be arranged in a cloud computing architecture rather than integrated in a single device. In a cloud computing architecture, components may be provided remotely and work together to implement the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require an end user to be aware of the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses suitable protocols to provide services via a wide area network, such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. Software or components of the cloud computing architecture and corresponding data may be stored on a server at a remote location. Computing resources in a cloud computing environment may be consolidated or distributed at locations of remote data centers. Cloud computing infrastructure may provide services through shared data centers, although to a user, they appear as a single access point. Thus, a cloud computing architecture may be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein may be provided by a conventional server or installed directly or otherwise on a client device.
[0758] In an embodiment of the present disclosure, computing device 4700 may be used to implement video encoding / decoding. Memory 4720 may include one or more video codec modules 4725 having one or more program instructions. These modules are accessible and executable by processing unit 4710 to perform the functions of the various embodiments described herein.
[0759] In an example embodiment of performing video encoding, input device 4750 may receive video data as input 4770 to be encoded. The video data may be processed, for example, by video codec module 4725 to generate an encoded bitstream. The encoded bitstream may be provided as output 4780 via output device 4760.
[0760] In an example embodiment of performing video decoding, input device 4750 may receive the encoded bitstream as input 4770. The encoded bitstream may be processed, for example, by video codec module 4725 to generate decoded video data. The decoded video data may be provided as output 4780 via output device 4760.
[0761] Although the present disclosure has been specifically shown and described with reference to preferred embodiments of the present disclosure, those skilled in the art will understand that various changes may be made in form and detail without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be restrictive.
Claims
1. A method for video processing, comprising: For the conversion between a video unit of a video and the bitstream of the video, using a predefined method to divide the video unit into a plurality of sub-divisions, wherein the video unit is encoded and decoded using an Intra Block Copy (IBC)-Geometry Partition Mode (GPM); Using Intra Block Copy to obtain a prediction of at least one sub-division of the video unit; and Performing the conversion based on the prediction of the at least one sub-division of the video unit.
2. The method according to claim 1, wherein the predefined method is the same as the method used for the Geometry Partition Mode (GPM), or wherein the predefined method is the same as the method used for GPM-Intra.
3. The method according to claim 1, wherein the predefined method is different from the method used for GPM, or wherein the predefined method is different from the method used for GPM-Intra.
4. The method according to claim 3, wherein the geometric angle associated with the predefined method is different from the geometric angle used for the GPM or the GPM-Intra.
5. The method according to claim 3, wherein the geometric offset associated with the predefined method is different from the geometric offset used for the GPM or the GPM-Intra.
6. The method according to claim 5, wherein a geometric offset equal to 0 is not allowed to be used in the IBC-GPM.
7. The method according to claim 3, wherein the predefined method for dividing the video unit depends on at least one of the following: the block size or block dimension of the video unit.
8. The method according to claim 1, wherein at least one geometry partition mode used for GPM or GPM-Intra is not allowed to be used for IBC-GPM.
9. The method according to claim 1, wherein the maximum number of geometry partition modes is predefined, or wherein the maximum number of the geometry partition modes is indicated.
10. The method according to claim 1, wherein sub-divisions adjacent to at least one side of the video unit are predicted by intra prediction.
11. The method according to claim 10, wherein the at least one side includes at least one of the following: The left side of the video unit, The upper side of the video unit, The right side of the video unit, or The bottom side of the video unit.
12. The method according to claim 1, wherein sub-divisions adjacent to at least one side of the video unit are not predicted by intra prediction.
13. The method according to claim 12, wherein the at least one side includes at least one of the following: The right side of the video unit, or The bottom side of the video unit.
14. The method according to claim 1, wherein the plurality of sub-divisions includes N sub-divisions, where N is an integer greater than 1.
15. The method according to claim 14, wherein N is equal to 2.
16. The method according to claim 14, wherein N is equal to 2, the prediction signal of the first sub - division of the video unit is derived using a first method, and the prediction signal of the second sub - division of the video unit is derived using a second method.
17. The method according to claim 16, wherein one of the first method and the second method comprises one of the following: IBC, palette, or intra - template matching, and wherein the other of the first method and the second method comprises one of the following: IBC, palette, intra - template matching, intra - prediction, or inter - prediction.
18. The method according to claim 16, wherein the first method and the second method are different.
19. The method according to claim 18, wherein the first method is IBC and the second method is intra - prediction, or wherein the second method is IBC and the first method is intra - prediction.
20. The method according to claim 14, further comprising: determining which method will be used to derive the prediction signal for the at least one sub - division.
21. The method according to claim 20, wherein the determination is indicated, or wherein the determination is predefined, or wherein the determination is derived in real - time.
22. The method according to claim 20, wherein at least one syntax element is used for the determination.
23. The method according to claim 20, wherein if N is equal to 2, a syntax element is used to indicate whether intra - prediction is used to derive the prediction signal of one sub - division of the plurality of sub - divisions, and the syntax elements equal to a predefined number indicate that the intra - prediction is used.
24. The method according to claim 23, wherein the syntax element is equal to the predefined number for one of the plurality of sub - divisions, and a specified prediction method is used for the other of the plurality of sub - divisions.
25. The method according to claim 20, wherein the determination depends on codec information.
26. The method according to claim 25, wherein the codec information comprises at least one of the following: block size, block dimension, or split mode.
27. The method according to claim 1, wherein whether to construct an intra - prediction mode (IPM) candidate list and / or the method of constructing the IPM candidate list depends on codec information.
28. The method according to claim 27, wherein the codec information comprises at least one of the following: block size, block dimension, or split mode, which geometrically divides the video unit into the plurality of sub - divisions.
29. The method according to claim 27, wherein the size of the IPM candidate list is the same for different block sizes or different block dimensions, or wherein the size of the IPM candidate list is the same for all split modes.
30. The method according to claim 27, wherein the size of the IPM candidate list or the maximum number of IPMs used to derive intra - prediction for a sub - division is predefined, or wherein the size of the IPM candidate list or the maximum number of IPMs used to derive the intra prediction for the sub - partition is indicated, or wherein the size of the IPM candidate list or the maximum number of IPMs used to derive the intra prediction for the sub - partition is derived in real - time.
31. The method according to claim 1, wherein IBC - GPM is not allowed to be used with one or more codec tools, or wherein the IBC - GPM is used with the one or more codec tools.
32. The method according to claim 31, wherein the one or more codec tools include at least one of the following: IBC Advanced Motion Vector Prediction (AMVP) mode, and IBC Merge mode, IBC Template Matching (IBC - TM) mode, IBC Merge mode with Block Vector Difference (IBC - MBVD) mode, Reconstruction - Reordering IBC (RR - IBC) Merge mode, IBC with Local Illumination Compensation (IBC - LIC) mode, IBC - CIIP mode, or Adaptive Motion Vector Resolution (AMVR) for IBC.
33. The method according to claim 1, wherein the determination of whether the video unit is allowed to be coded / decoded in IBC - GPM mode depends on the codec information.
34. The method according to claim 33, wherein the codec information includes at least one of the following: block dimension or block size.
35. The method according to claim 34, wherein if the block size represented as W×H is less than or equal to a threshold, the video unit is allowed to be coded / decoded in IBC - GPM mode, where W and H represent the block width and block height respectively.
36. The method according to claim 35, wherein the threshold is one of the following: 256, 512, 1024, 2048 or 4096.
37. The method according to claim 33, wherein the codec information includes at least one of the following: slice type or picture type.
38. The method according to claim 37, wherein IBC - GPM is only applied to I - slices or I - pictures.
39. The method according to claim 1, wherein whether to apply IBC - GPM and / or the method of applying IBC - GPM depends on at least one of the following: color format or color component.
40. The method according to claim 39, wherein IBC - GPM is applied to all color components.
41. The method according to claim 39, wherein IBC - GPM is applied to the chrominance component, and the derivation of intra prediction is different from the derivation of intra prediction for the luminance component.
42. The method according to claim 41, wherein the intra prediction is obtained using at least one of the following: Cross - Component Linear Mode (CCLM) mode, Multi - Mode Linear Mode (MMLM), Convolutional Cross - Component Model (CCCM), Decoder - side Intra Mode Derivation for Chrominance (DIMD), Template - based Intra Mode Derivation for Chrominance (TIMD), Combination of CCLM and angular mode, Combination of MMLM and angular mode, or Combination of CCCM and angular mode.
43. The method according to claim 1, wherein whether to apply IBC-GPM to the first component and / or the method of applying IBC-GPM to the first component depends on whether to apply IBC-GPM to the second component.
44. The method according to claim 43, wherein the first component is a chrominance component and the second component is a luminance component.
45. The method according to claim 43, wherein the method of applying IBC-GPM to the first component is the same as that of the second component.
46. The method according to claim 43, wherein the method of applying IBC-GPM to the first component is different from that of the second component.
47. The method according to claim 46, wherein the weighting parameters are different.
48. The method according to claim 39, wherein IBC-GPM is applied to the luminance component and not to the chrominance component.
49. The method according to claim 48, wherein the luminance component includes Y in the YCbCr color space, or wherein the luminance component includes G in the Red-Green-Blue (RGB) color space.
50. The method according to claim 48, wherein the chrominance component includes at least one of Cb or Cr in the YCbCr color space, or wherein the chrominance component includes at least one of the following: R or B in the RGB color space.
51. The method according to claim 1, wherein the indication of IBC-GPM is transmitted by a signal based on a condition, where the condition includes: Whether the codec tool is allowed.
52. The method according to claim 51, wherein the codec tool includes at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC mode, or IBC-CIIP mode.
53. The method according to claim 1, wherein whether to apply IBC-GPM to the video unit and / or the method of applying IBC-GPM to the video unit is indicated in the bitstream.
54. The method according to claim 53, wherein how and whether to divide the video unit into the plurality of sub-partitions is predefined, or wherein how and whether to divide the video unit into the plurality of sub-partitions is derived, or wherein how and whether to divide the video unit into the plurality of sub-partitions is transmitted by a signal in the bitstream.
55. The method according to claim 54, wherein one or more syntax elements are transmitted by a signal to indicate the intra prediction mode (IPM) used for intra prediction in the sub-partition.
56. The method according to claim 54, wherein one or more syntax elements are transmitted by a signal to indicate the IBCAMVP index or the IBC Merge index to indicate the candidate prediction in IBC for the sub-partition.
57. The method according to claim 53, wherein the syntax element is signaled before the indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC or IBC-CIIP, or wherein the syntax element is signaled after the indication of at least one of the following: the IBC-TM mode, the IBC-MBVD mode, the RR-IBC mode, the IBC-LIC or the IBC-CIIP.
58. The method according to claim 57, wherein whether the syntax element is signaled and / or the way the syntax element is signaled depends on whether at least one of the following is enabled for the video unit: the IBC mode, the IBC-TM mode, the IBC-MBVD mode, the RR-IBC mode, the IBC-LIC, or the IBC-CIIP.
59. The method according to any one of claims 1 to 58, wherein the video unit comprises at least one of the following: color component, prediction block (PB), transformation block (TB), coding block (CB), prediction unit (PU), transformation unit (TU), coding tree block (CTB), coding unit (CU), coding tree unit (CTU), CTU row, CTU group, slice, picture, sub-picture, block, sub-region within a block, or region containing more than one sample or pixel.
60. The method according to any one of claims 1 to 58, wherein an indication of whether the prediction of at least one sub-division of the video unit is obtained by using IBC and / or how the prediction of at least one sub-division of the video unit is obtained by using IBC is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or group of slices level.
61. The method according to any one of claims 1 to 58, wherein an indication of whether the prediction of at least one sub-division of the video unit is obtained by using IBC and / or how the prediction of at least one sub-division of the video unit is obtained by using IBC is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or group of slices header.
62. The method according to any one of claims 1 to 58, wherein an indication of whether the prediction of at least one sub-division of the video unit is obtained by using IBC and / or how the prediction of at least one sub-division of the video unit is obtained by using IBC is included in one of the following: prediction block (PB), transformation block (TB), coding block (CB), prediction unit (PU), transformation unit (TU), coding unit (CU), virtual pipeline data unit (VPDU), coding tree unit (CTU), CTU row, slice, picture, sub-picture, or A region containing more than one sample or pixel.
63. The method according to any one of claims 1 to 58, further comprising: Based on the coding and decoding information of the video unit, determining whether to obtain the prediction of at least one sub - partition of the video unit by using IBC and / or how to obtain the prediction of at least one sub - partition of the video unit by using IBC, where the coding and decoding information includes at least one of the following: Block size, Color format, Single and / or dual - tree segmentation, Color component, Slice type, or Picture type.
64. The method according to any one of claims 1 to 63, wherein the transformation includes encoding the video unit into the bitstream.
65. The method according to any one of claims 1 to 63, wherein the transformation includes decoding the video unit from the bitstream.
66. An apparatus for video processing, comprising a processor and a non - transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of claims 1 - 65.
67. A non - transitory computer - readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1 - 65.
68. A non - transitory computer - readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method comprises: Dividing video units of the video into multiple sub - partitions using a predefined method, where the video units are coded and decoded using intra - block copy (IBC) - geometric partitioning mode (GPM); Obtaining the prediction of at least one sub - partition of the video unit using intra - block copy; and Generating the bitstream based on the prediction of the at least one sub - partition of the video unit.
69. A method for storing a bitstream of a video, comprising: Dividing video units of the video into multiple sub - partitions using a predefined method, where the video units are coded and decoded using intra - block copy (IBC) - geometric partitioning mode (GPM); Obtaining the prediction of at least one sub - partition of the video unit using intra - block copy; Generating the bitstream based on the prediction of the at least one sub - partition of the video unit; and Storing the bitstream in a non - transitory computer - readable recording medium.