Method and device for video processing and medium

By using an affine model to process multiple sub-blocks of a video unit in video encoding and decoding, the problem of insufficient efficiency in texture processing encoding and decoding in existing technologies is solved, and more efficient encoding and decoding performance is achieved.

CN121264040APending Publication Date: 2026-01-02DOUYIN VISION CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480038116.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-06-07
Filing Date
2024-06-05
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies are inefficient when dealing with textures, especially when considering translational motion within an image, and fail to fully utilize the features of video data.

Method used

The block vector of the video unit is derived using an affine model. Multiple sub-blocks of the video unit are processed using the same block size and affine mode parameters or position information, and transformation is performed based on the block vector to improve encoding and decoding performance.

Benefits of technology

The application of affine models enhances the texture processing capabilities of video encoding and decoding, thereby improving encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121264040A_ABST
    Figure CN121264040A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is presented. The method comprises, for a conversion between a video unit of the video and a bitstream of the video, deriving a block vector (BV) for at least one of a plurality of sub-blocks of the video unit using an affine model, where a same block size is used for all of the plurality of sub-blocks, and / or at least one of a parameter of the affine mode or a position of the video unit is applied to a subsequent video unit of the video unit; and performing the conversion based on the BV.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present disclosure generally relate to video processing technology, and more particularly, to intra block copy (IBC) with affine model. BACKGROUND

[0002] Nowadays, digital video capability is being applied to various aspects of people's life. For video coding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, in the current design of IBC in video coding, only translational motion within a picture is considered. Therefore, it is generally desirable to further improve the coding efficiency of video coding technology. SUMMARY

[0003] Embodiments of the present disclosure provide a solution for video processing.

[0004] In a first aspect, a method for video processing is proposed. The method comprises deriving, for a conversion between a video unit of a video and a bitstream of the video, a block vector (BV) for at least one sub-block of a plurality of sub-blocks of the video unit using an affine model, wherein a same block size is used for all sub-blocks in the plurality of sub-blocks, and / or at least one of parameters of the affine model or a position of the video unit is applied to a subsequent video unit of the video unit; and performing the conversion based on the BV. Compared with the conventional solution, the method according to the first aspect of the present disclosure can improve the coding performance to handle texture by the affine model.

[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.

[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions to cause a processor to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing. The method includes deriving a block vector (BV) for at least one sub-block of a plurality of sub-blocks of a video unit of the video using an affine model, wherein a same block size is used for all of the plurality of sub-blocks, and / or at least one of parameters of the affine model or a position of the video unit is applied to a subsequent video unit of the video unit, and generating the bitstream based on the BV.

[0008] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes deriving a block vector (BV) for at least one sub-block of a plurality of sub-blocks of a video unit of the video using an affine model, wherein a same block size is used for all of the plurality of sub-blocks, and / or at least one of parameters of the affine model or a position of the video unit is applied to a subsequent video unit of the video unit, generating the bitstream based on the BV, and storing the bitstream in a non-transitory computer-readable recording medium.

[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to the like elements throughout. In the example embodiments of the present disclosure, like reference numerals refer to like elements throughout.

[0011] Figure 1 A block diagram showing an example video coding system is shown in accordance with some embodiments of the present disclosure; Figure 2 A block diagram showing a first example video encoder is shown in accordance with some embodiments of the present disclosure; Figure 3 A block diagram showing an example video decoder is shown in accordance with some embodiments of the present disclosure; Figure 4 An example of an encoder block diagram is shown; Figure 5 Sixty-seven intra prediction modes are shown; Figure 6 Reference samples for wide-angle intra prediction are shown; Figure 7 Discontinuity issues in cases where the direction exceeds 45° are shown; Figure 8MMVD search points are shown; Figure 9 is a diagram for symmetric MVD mode; Figure 10 Extended CU regions used in BDOF are shown; Figure 11a and Figure 11b Control point based affine motion model is shown; Figure 12 Affine MVFs per subblock are shown; Figure 13 Location of inherited affine motion predictor is shown; Figure 14 Control point motion vector inheritance is shown; Figure 15 Location of candidate positions for constructed affine Merge mode is shown; Figure 16 is a diagram for motion vector usage for proposed combination method; Figure 17 Subblock MV VSB and pixel ; Figure 18a and Figure 18b SbTMVP process in VVC is shown; Figure 19 Local illumination compensation is shown; Figure 20 No downsampling for short side is shown; Figure 21 Decoder-side motion vector refinement is shown; Figure 22 Diamond-shaped region in search area is shown; Figure 23 Location of spatial Merge candidate is shown; Figure 24 Candidate pairs considered for redundancy check for spatial Merge candidate are shown; Figure 25 is a diagram for motion vector scaling for temporal Merge candidate; Figure 26 Candidate positions for temporal Merge candidate, C0 and C1 are shown; Figure 27 VVC spatial neighboring blocks of current block are shown; Figure 28 is a diagram for virtual blocks in i-th round search; Figure 29 Example of GPM partitioning grouped by same angle is shown; Figure 30An example of a unidirectional prediction MV selection for a geometric partition mode is shown; Figure 31 An example of a bending weight using a geometric partition mode is shown Figure 32 An example of a spatial neighboring block used to derive a spatial Merge candidate is shown; Figure 33 An example of a template matching performed on a search region around the initial MV is shown; Figure 34 An example of a subblock where OBMC is applied is shown; Figure 35 An example of SBT position, type and transform type is shown; Figure 36 An example of neighboring samples used to compute SAD is shown; Figure 37 An example of neighboring samples used to compute SAD for sub-CU level motion information is shown; Figure 38 An example of a sorting process is shown; Figure 39 An example of a reordering process in the encoder is shown; Figure 40 An example of a reordering process in the decoder is shown; Figure 41 An example of an extended reference region is shown; Figure 42 An example of an IBC reference region depending on the current CU position is shown; Figure 43 An example of a symmetry in a screen content picture is shown; Figure 44a An example of a BV adjustment for horizontal flipping is shown, and Figure 44b An example of a BV adjustment for vertical flipping is shown; Figure 45 An example of an intra template matching search region used is shown; Figure 46 An example of a template region is shown; Figure 47 An example of a spatial part of a convolution filter is shown; Figure 48 An example of a reference region (and its padding) used to derive filter coefficients is shown; Figure 49 An example of four Sobel-based gradient modes for GLM is shown; Figure 50 An example of non-reconstructed samples (shaded) in a reference block being padded by copying their predicted samples is shown; Figure 51 ​Five positions in the reconstructed luma samples are shown; Figure 52 A prediction process for DBV mode is shown; Figure 53 A spatial domain portion of the filter is shown; Figure 54 A flowchart of a method for video processing according to an embodiment of the disclosure is shown; Figure 1 A block diagram of a computing device in which various embodiments of the disclosure can be implemented is shown.

[0012] Throughout the drawings, identical or similar reference numerals are generally used to refer to identical or similar elements throughout the drawing figures. DETAILED DESCRIPTION

[0013] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the embodiments are described for illustrative purposes only and to help the person skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways in addition to the ways described below.

[0014] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.

[0015] References in the present disclosure to “one embodiment”, “an embodiment”, “example embodiments”, etc. indicate that the embodiment described can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that the embodiment is not limited to those particular features, structures, or characteristics, whether or not the particular feature, structure, or characteristic is recited in the claims, unless otherwise expressly so limited by the claims.

[0016] It should be understood that although the terms “first” and “second” etc. can be used herein to describe various elements, the elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element can be called a second element, and similarly, a second element can be called a first element, without departing from the scope of the example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the listed terms.

[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises," "comprising," "includes" and / or "including," when used herein, specify the presence of stated features, elements and / or components, but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.

[0018] Example Environment Figure 2 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As illustrated, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 can also be referred to as a video encoding device, and destination device 120 can also be referred to as a video decoding device. In operation, source device 110 can be configured to generate encoded video data, and destination device 120 can be configured to decode the encoded video data generated by source device 110. Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0019] Video source 112 can include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system to generate video data, and / or a combination thereof.

[0020] Video data can include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that form an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be transmitted directly to destination device 120 by way of I / O interface 116, through network 130A. The encoded video data can also be stored onto a storage medium / server 130B for access by destination device 120.

[0021] Destination device 120 can include I / O interface 126, video decoder 124, and display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120, which is configured to interface with an external display device.

[0022] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other existing and / or further standards.

[0023] Figure 1 is a block diagram illustrating an example of a video encoder 200 that can be Figure 2 an example of video encoder 114 in system 100 shown.

[0024] Video encoder 200 can be configured to implement any or all of the techniques of this disclosure. In Figure 2 example, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared between the components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0025] In some embodiments, video encoder 200 can include partitioning unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.

[0026] In other examples, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra-block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.

[0027] Furthermore, although some components, such as motion estimation unit 204 and motion compensation unit 205, can be integrated, for purposes of explanation, these components are shown as separate components in Figure 3are shown separately in the example.

[0028] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.

[0029] Mode selection unit 203 can select one of a plurality of coding modes (intra- or inter-coding), e.g., based on the error results, and provide the resulting intra- or inter-coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, mode selection unit 203 can select a combined intra-inter prediction (CIIP) mode in which prediction is based on both an inter prediction signal and an intra prediction signal. In the case of inter prediction, mode selection unit 203 can also select a resolution for a motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).

[0030] To perform inter prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from cache 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from cache 213 other than the picture associated with the current video block.

[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations for a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" can refer to a portion of a picture composed of macroblocks all of which are based on macroblocks within the same picture. Further, as used herein, a "P slice" and a "B slice" can refer, in some aspects, to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture.

[0032] In some examples, motion estimation unit 204 can perform uni-directional prediction for a current video block, and motion estimation unit 204 can search a reference picture of list 0 or list 1 for a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference picture of list 0 or list 1 containing the reference video block and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.

[0033] Alternatively, in other examples, the motion estimation unit 204 can perform bi-prediction for the current video block. The motion estimation unit 204 can search the reference pictures in List 0 for one reference video block for the current video block and also search the reference pictures in List 1 for another reference video block for the current video block. The motion estimation unit 204 can then generate a plurality of reference indices that indicate the plurality of reference pictures in List 0 and List 1 that contain the plurality of reference video blocks and a plurality of motion vectors that indicate a plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 can output the plurality of reference indices and the plurality of motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information for the current video block.

[0034] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoding process at the decoder. Alternatively, in some embodiments, the motion estimation unit 204 can signal the motion information for the current video block with reference to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information for the current video block is sufficiently similar to the motion information of a neighboring video block.

[0035] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0036] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0037] As discussed above, the video encoder 200 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.

[0038] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.

[0039] Residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by the minus sign) the prediction video block(s) from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0040] In other examples, such as in a skip mode, there can be no residual data for the current video block for the current video block, and residual generation unit 207 can not perform the subtraction operation.

[0041] Transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0042] Quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block after transform processing unit 208 generates the transform coefficient video blocks associated with the current video block.

[0043] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video blocks to reconstruct the residual video blocks from the transform coefficient video blocks. Reconstruction unit 212 can add the reconstructed residual video blocks to corresponding samples from the prediction video block(s) generated by prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in buffer 213.

[0044] Loop filtering operations can be performed to reduce video block artifacts in the video blocks after reconstruction unit 212 reconstructs the video blocks.

[0045] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.

[0046] Figure 1 FIG. 3 is a block diagram illustrating an example of a video decoder 300 that can be Figure 3 of the video decoder 124 in the system 100 shown.

[0047] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3In examples of the video decoder 300, the video decoder 300 includes a number of functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0048] In Figure 4 In examples of the video decoder 300, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transformation unit 305, and a reconstruction unit 306 and buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200.

[0049] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy encoded video data and the motion compensation unit 302 can determine motion information from the entropy decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes. AMVP is used including deriving a number of most probable candidates based on data from neighboring PBs and reference pictures. The motion information typically includes a horizontal motion vector displacement value and a vertical motion vector displacement value, one or two reference picture indices, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, “Merge mode” can refer to deriving motion information from a spatially or temporally neighboring block.

[0050] The motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax elements.

[0051] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 from the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate the prediction block.

[0052] Motion compensation unit 302 can use at least some of the syntax information to determine the size of the blocks used to encode the frame(s) and / or slice(s) of the coded video sequence, partitioning information describing how each macroblock of a picture of the coded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and lists of reference frames) for each inter-coded block, and other information used to decode the coded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture, or can also be a region of a picture.

[0053] Intra prediction unit 303 can use, for example, intra prediction modes received in the bitstream to form a prediction block from spatial neighboring blocks. Dequantization unit 304 dequantizes quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0054] Reconstruction unit 306 can obtain a decoded block, for example, by adding the residual block to the corresponding prediction block generated by motion compensation unit 302 or intra prediction unit 303. If desired, a deblocking filter can also be applied to filter the decoded block in order to remove blockiness artifacts. The decoded video blocks are then stored in buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and which also produces decoded video for presentation on a display device.

[0055] Some example embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for convenience only and not to be construed as limiting the embodiments disclosed in that section to that section only. Furthermore, although some embodiments are described with reference to a multi-functional video codec or other specific video codec, the disclosed techniques are applicable to other video codec technologies. Moreover, although some embodiments are described in detail with reference to video encoding steps, it should be understood that corresponding decoding steps to undo the encoding would be implemented by a decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding in which video pixels are represented from one compressed format to another compressed format or at a different compressed bit rate.

[0056] 1. BRIEF OVERVIEW The present disclosure relates to video coding techniques. In particular, it relates to Intra Block Copy (IBC) with affine model, and how to apply IBC with affine model to IBC coded blocks, and other coding tools in image / video coding. It can be applied to existing video coding standards like HEVC or Versatile Video Coding (VVC). It can also be applicable to future video coding standards or video codecs.

[0057] 2. Introduction Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263 standards, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. From H.262 onwards, the video coding standards are based on the hybrid video coding structure where temporal prediction plus transform coding is utilized. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly formed the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has taken many new methods and incorporated them into a reference software named Joint Exploration Model (JEM). In April 2018, the Joint Video Team (JVT) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was formed to work on the VVC standard, targeting 50% bitrate reduction compared to HEVC. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5) are studying the potential need for standardization of future video coding technologies with significantly improved compression capability over the current VVC standard. Such future standardization activities can take the form of additional extensions of VVC or an entirely new standard. These groups are working together in a joint collaborative effort called the Joint Video Exploration Team (JVET) to evaluate compression technology designs proposed by experts in the field. New coding features and coding methods implemented in the Enhanced Compression Mode (ECM) software, as potential enhanced video coding technologies beyond VVC capability, are being explored by the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG.

[0058] 2.1. Coding process of a typical video codec Figure 5An example of an encoder block diagram of VVC is shown, which contains three in-loop filtering blocks: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF which uses a pre-defined filter, SAO and ALF utilize the original samples of the current picture, by adding an offset and by applying a Finite Impulse Response (FIR) filter, respectively, and reduce the mean square error between the original and the reconstructed samples by signaling the offset and the filter coefficients as side information. ALF is located at the last processing stage of each picture and can be considered as a tool that tries to capture and fix artifacts caused by previous stages.

[0059] 2.2. Intra mode coding with 67 intra prediction modes Figure 5 The 67 intra prediction modes are shown. To capture arbitrary edge directions that are present in natural videos, as Figure 6 shown, the number of directional intra modes is extended from 33 as used in HEVC to 65, and the planar and DC modes remain unchanged. These denser directional intra prediction modes are applied to all block sizes and to both luma and chroma intra prediction.

[0060] In HEVC, each intra coded block has a square shape and the length of each side is a power of 2. Therefore, no division operation is needed to generate the intra prediction value using the DC mode. In VVC, the block can have a rectangular shape, which in general case requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value for non-square blocks.

[0061] 2.2.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. The regular angular intra prediction directions are defined as clockwise directions from 45 degrees to -135 degrees. In VVC, for non-square blocks, several regular angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes is unchanged, i.e., 67, and the intra mode coding method is unchanged.

[0062] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 6 shown. Figure 7 The reference samples for wide-angle intra prediction are shown.

[0063] The number of modes replaced in the wide-angle directional modes depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2-1.

[0064] Table 2-1 - Intra prediction modes replaced by wide-angle modes

[0065] Figure 7 The problem of discontinuity in the case of directions beyond 45° is shown. As shown in Figure 9 , in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap p α If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that satisfy this condition, i.e., [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted by these modes, the samples in the reference buffer are copied directly without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of non-fraction modes in regular prediction modes with the wide-angle modes.

[0066] In VVC, 4:2:2 and 4:4:4 chroma formats are supported in addition to 4:2:0. The chroma derivation mode (DM) derivation table for 4:2:2 chroma format is originally ported from HEVC, extending the number of entries from 35 to 67 to align with the extension of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes with range from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for 4:2:2 chroma format is updated by replacing some values of the entries of the mapping table to convert the prediction angles for chroma blocks more accurately.

[0067] 2.3. Inter prediction For each inter prediction CU, the motion parameters include the motion vector, the reference picture index and the reference picture list usage index, and additional information needed for inter prediction sample generation for VVC's new coding features. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no coded motion vector delta or reference picture index. Merge mode is specified, in which the motion parameters for the current CU are derived from neighboring CUs, including spatial and temporal candidates and additional scheduling introduced in VVC. Merge mode can be applied to any inter prediction CU, not only for skip mode. An alternative to Merge mode is the explicit signaling of motion parameters, in which the motion vector, the corresponding reference picture index for each reference picture list and the reference picture list usage flag and other needed information are explicitly signaled for each CU.

[0068] 2.4. Intra block copy (IBC) Intra block copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed inside the current picture. The luma block vector of an IBC-coded CU is in integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. An IBC-coded CU is considered as a third prediction mode different from intra or inter prediction modes. IBC mode is applicable to CUs with width and height less than or equal to 64 luma samples.

[0069] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check for blocks with width or height no larger than 16 luma samples. For non-Merge mode, the block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a block matching based local search will be performed.

[0070] In hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key computation for each position in the current picture is based on 4x4 sub-blocks. For a current block of larger size, the hash key is determined to match the hash key of a reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference positions. If multiple reference blocks' hash keys are found to match the hash key of the current block, the block vector cost of each matched reference is computed, and the one with the smallest cost is selected.

[0071] In block matching search, the search range is set to cover both the previous CTU and the current CTU.

[0072] At CU level, the IBC mode is signaled with a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows: - IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidates IBC-coded blocks is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates and paired candidates.

[0073] - IBC AMVP mode: The block vector difference is coded in the same way as the motion vector difference. The block vector prediction method uses two candidates as the predictor, one from the left neighboring block and one from the above neighboring block (if IBC-coded). When either neighboring block is unavailable, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.

[0074] 2.5. IBC motion candidates The term "block" can refer to a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB, or a unit of video processing including a plurality of samples / pixels. The block can be rectangular or non-rectangular.

[0075] For IBC-coded blocks, a block vector (BV) is used to indicate the displacement from the current block to a reference block that has been reconstructed inside the current picture.

[0076] W and H are the width and height of the current block (e.g. luma block).

[0077] The non-adjacent spatial candidate of the current coding block is the adjacent spatial candidate of the virtual block in the i-th round search (as in Figure 8The width and height of the virtual block for the i-th search round are calculated by the following equations: newWidth = i x 2 x gridX + W, newHeight = i x 2 x gridY + H. Obviously, if the search round i is 0, the virtual block is the current block.

[0078] In the following, a BV predictor is also a BV candidate. The skip mode is also a Merge mode.

[0079] BV candidates can be divided into several groups according to some criteria. Each group is called a sub-group. For example, we can put the neighboring spatial and temporal BV candidates as the first sub-group, and the remaining BV candidates as the second sub-group; in another example, we can also put the first N (N > 2) BV candidates as the first sub-group, the following M (M > 2) BV candidates as the second sub-group, and the remaining BV candidates as the third sub-group.

[0080] 2.6. Merge mode with MVD (MMVD) In addition to the Merge mode (where the implicitly derived motion information is directly used for the prediction sample generation of the current CU), the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag to specify whether the MMVD mode is used for the CU.

[0081] In MMVD, after the Merge candidate is selected, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, the index to specify the motion magnitude, and the index to indicate the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected as the MV base. The MMVD candidate flag is signaled to specify which one is used between the first Merge candidate and the second Merge candidate.

[0082] The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 8 MMVD search points are shown. As Figure 23 shown, the offset is added to the horizontal component or the vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2-2.

[0083] Table 2-2 - Relationship between distance index and predefined offset

[0084] The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions as shown in Table 2-3. It should be noted that the meaning of the MVD sign can change depending on the information of the starting MV. When the starting MV is a uni-predicted MV or a bi-predicted MV and both lists point to the same side of the current picture (i.e., both reference POCs are greater than or less than the POC of the current picture), the sign in Table 2-3 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bi-predicted MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture and the other is less than the POC of the current picture), and the difference of POCs in list 0 is greater than the difference of POCs in list 1, the sign in Table 2-3 specifies the sign of the MV offset added to the listO MV component of the starting MV and the sign for the listl MV has the opposite value. Otherwise, if the difference of POCs in list 1 is greater than the difference of POCs in list 0, the sign in Table 2-3 specifies the sign of the MV offset added to the listl MV component of the starting MV and the sign for the listO MV has the opposite value.

[0085] The MVD is scaled according to the difference of POCs in each direction. If the difference of POCs in the two lists is the same, no scaling is needed. Otherwise, if the difference of POCs in list 0 is greater than the difference of POCs in list 1, the MVD for list 1 is scaled as shown in Figure 10

[0086] Table 2-3 - Sign of MV offset specified by direction index

[0087] 2.7. Symmetric MVD coding In VVC, in addition to the normal uni-predicted and bi-predicted mode MVD signaling, a symmetric MVD mode for bi-predicted MVD signaling is also applied. In the symmetric MVD mode, the motion information (including the reference picture index for both list 0 and list 1 and the MVD for list 1) is not signaled but derived.

[0088] The decoding process for the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0 and RefIdxSymL1 are derived as follows: ​- If mvd_l1_zero_flag is equal to 1, BiDirPredFlag is set equal to 0.

[0089] - Otherwise, if the most recent reference picture in list 0 and the most recent reference picture in list 1 form a pair of forward and backward reference pictures or a pair of backward and forward reference pictures, BiDirPredFlag is set equal to 1 and both list 0 and list 1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set equal to 0.

[0090] 2) At the CU level, if the CU is bi-directionally prediction coded and BiDirPredFlag is equal to 1, a symmetric mode flag indicating whether the symmetric mode is used or not is explicitly signaled.

[0091] When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. The reference indices of list 0 and list 1 are set equal to the pair of reference pictures, respectively. MVD1 is set equal to (-MVD0). The final motion vector is given by the following equation.

[0092] (2-1) In the encoder, symmetric MVD motion estimation starts from the initial MV estimation. A set of initial MV candidates, including the MVs obtained from the uni-prediction search, the MVs obtained from the bi-prediction search, and the MVs from the AMVP list. One with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search.

[0093] 2.8. Bi-directional optical flow (BDOF) The bi-directional optical flow (BDOF) tool is included in VVC. BDOF (previously called BIO) is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version which requires much less computation, especially in terms of the number of multiplications and the size of the multiplier.

[0094] BDOF is used to refine the bi-prediction signal of a CU at the 4x4 subblock level. BDOF is applied to a CU if it satisfies all the following conditions: - The CU is coded using the "true" bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order; - The distance (i.e., POC difference) from the two reference pictures to the current picture is the same; - Both reference pictures are short-term reference pictures; - CU is not encoded or decoded using affine mode or SbTMVP Merge mode; - The CU has more than 64 luminance samples; - Both the CU height and CU width are greater than or equal to 8 luminance samples; - The BCW weight index indicates equal weights; - WP is not enabled for the current CU; - CIIP mode is not used in the current CU.

[0095] BDOF is applied only to the luma component. As the name suggests, the BDOF mode is based on the concept of optical flow, which assumes that the motion of an object is smooth. Motion refinement is performed for each 4×4 sub-block. The difference between the L0 and L1 predicted samples is calculated. Motion refinement is then used to adjust the bidirectional predicted sample values ​​in the 4x4 sub-blocks. The following steps are applied during the BDOF process.

[0096] First, the horizontal and vertical gradients of the two predicted signals. and , It is calculated by directly calculating the difference between two neighboring sample points, i.e.

[0097] in It is a list , Coordinates of the predicted signal in The sample value at the location, and shift1 is calculated based on the luminance bit depth bitDepth as shift1 = max(6, bitDepth-6).

[0098] Then, the autocorrelation and cross-correlation of the gradients. , , , and Calculated as

[0099] in

[0100] in It is a 6x6 window surrounding a 4x4 sub-block, and and The values ​​are respectively set to min(1, bitDepth) 11) and min(4, bitDepth) 8).

[0101] Motion refinement Then, the following formula is derived using cross-correlation and autocorrelation terms:

[0102] in , , . It is a floor function, and .

[0103] Based on motion refinement and gradients, the following adjustments are calculated for each sample point in the 4×4 sub-block:

[0104] Finally, the BDOF samples of CU are calculated by adjusting the bidirectional prediction samples as follows:

[0105] These values ​​were chosen so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.

[0106] To derive the gradient values, the list outside the current CU boundary is used. ( Some predicted samples in ) It needs to be generated. Figure 10 The extended CU region used in BDOF is shown. For example... Figure 11a As shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region (white area) are generated by directly taking reference samples at nearby integer positions (using floor() operations on coordinates) without interpolation, and a normal 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values ​​are only used in gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values ​​outside the CU boundary are needed, they are filled from their nearest neighbors (i.e., repeated).

[0107] When the width and / or height of a CU is greater than 16 luminance samples, it will be divided into sub-blocks with a width and / or height equal to 16 luminance samples, and the sub-block boundaries will be considered as CU boundaries in the BDOF process. The maximum cell size for the BDOF process is limited to 16×16. The BDOF process can be skipped for each sub-block. The BDOF process is not applied to the sub-block when the SAD between the initial L0 predicted samples and the L1 predicted samples is less than a threshold. The threshold is set to equal to (8... W (H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 prediction samples and L1 prediction samples calculated during the DVMR process is reused here.

[0108] Bidirectional optical flow (BDOF) is disabled if BCW is enabled for the current block, meaning the BCW weight index indicates unequal weights. Similarly, BDOF is disabled if WP is enabled for the current block, meaning either luma_weight_lx_flag is 1 for either of the two reference images. BDOF is also disabled when the CU is encoded / decoded in symmetric MVD or CIIP mode.

[0109] 2.9. Inter-frame and Intra-frame Joint Prediction (CIIP) 2.10. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many types of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 11b and Figure 11a An affine motion model based on control points is shown, where Figure 11b A 4-parameter affine model is shown, and Figure 11a A 6-parameter affine model is shown. (Example) Figure 11b and x, y As shown, the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).

[0110] For a 4-parameter affine motion model, the location of the sample point in the block ( x, y The motion vector at point () is derived as: (2-8) For a 6-parameter affine motion model, the location of the sample points in the block ( mv The motion vector at point () is derived as: (2-9) in( mv 0x , mv 0y ) is the motion vector of the upper left control point, ( mv 1x , mv 1y ) is the motion vector of the upper right control point, and ( mv 2x , Figure 122y ) is the motion vector of the bottom-left control point.

[0111] Figure 12 The affine MVFs for each sub-block are shown. To simplify the motion compensated prediction, block-based affine transform prediction is applied. To derive the motion vector for each 4x4 luma sub-block, the motion vector of the center sample of each sub-block is calculated according to the equation above (as shown in Figure 13 ), and is rounded to 1 / 16 fractional precision. Then a motion compensated interpolation filter is applied to generate the prediction for each sub-block with the derived motion vector. The sub-block size for chroma components is also set to 4x4. The MVs for 4x4 chroma sub-blocks are calculated as the average of the MVs of the four corresponding 4x4 luma sub-blocks.

[0112] As with translational inter prediction, there are also two affine inter prediction modes: affine Merge mode and affine AMVP mode.

[0113] 2.10.1. Affine Merge Prediction The AF M ERGE mode can be applied to a CU whose width and height are both greater than or equal to 8. In this mode, the CPMVs for the current CU are generated based on the motion information of spatial neighboring CUs. There can be up to five CPMV candidates, and one of them is indicated by signaling an index to be used for the current CU. The following three types of CPMV candidates are used to form the affine Merge candidate list: - Inherited affine Merge candidates inferred from the CPMVs of neighboring CUs.

[0114] - Constructed affine Merge candidate CPMVPs derived using the translational MVs of neighboring CUs.

[0115] - Zero MV.

[0116] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks, one from the left neighboring CU and one from the top neighboring CU. Figure 13 The positions of the inherited affine motion predictors are shown. The candidate blocks are shown as Figure 14 . For the left-side predictors, the scan order is A0->A1, and for the top-side predictors, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No de-duplication check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMVP candidates in the affine Merge list of the current CU. As shown, if the neighboring bottom-left block A is coded in affine mode, the motion vectors of the top-left, top-right, and bottom-left corners of the CU containing block A , and are obtained. When block A is coded with 4-parameter affine model, two CPMVs of the current CU are derived from and are calculated. When block A is coded with 6-parameter affine model, three CPMVs of the current CU are derived from , and are calculated. Figure 15 Control point motion vector inheritance is shown.

[0117] Constructed affine candidates mean that the candidate is constructed by combining the neighboring translational motion information of each control point. Figure 15 Position of the constructed affine Merge mode candidate position is shown. The motion information of the control points is derived from Figure 15 the specified spatial and temporal neighbors as shown. CPMVs k (k = 1, 2, 3, 4) denote the k-th control point. For CPMV1, the B2->B3->A2 blocks are checked and the MV of the first available block is used. For CPMV2, the Bl->B0 blocks are checked and for CPMV3, the Al->A0 blocks are checked. For TMVP, if available, it is used as CPMV4.

[0118] After obtaining the MVs of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, { CPMV1, CPMV3}.

[0119] The combinations of 3 CPMVs construct the 6-parameter affine Merge candidate and the combinations of 2 CPMVs construct the 4-parameter affine Merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the related combination of control point MVs is discarded.

[0120] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list.

[0121] 2.10.2. Affine AMVP prediction Affine AMVP mode can be applied to a CU whose width and height are both greater than or equal to 16. An affine flag at CU level is signaled in the bitstream to indicate whether affine AMVP mode is used or not, and then another flag is signaled to indicate whether 4-parameter affine or 6-parameter affine is used. In this mode, the difference between the CPMVs of the current CU and its prediction CPMVPs are signaled in the bitstream. The affine AMVP candidate list size is 2, and is generated by using the following four types of CPMV candidates in order: - Inherited affine AMVP candidate inferred from CPMVs of neighboring CUs.

[0122] - Constructed affine AMVP candidate CPMVP derived using translation MVs of neighboring CUs.

[0123] - Translation MVs from neighboring CUs.

[0124] - Zero MV.

[0125] The checking order of inherited affine AMVP candidates is the same as that of inherited affine merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting inherited affine motion predictors into the candidate list, no de-duplication process is applied.

[0126] Constructed AMVP candidates are derived from the specified spatial neighbors as shown in Table 1. mv The same checking order as done in affine merge candidate construction is used. In addition, the reference picture index of the neighboring block is also checked. The first block in the checking order that is inter coded and has the same reference picture as in the current CU is used. There is only one. When the current CU is coded with 4-parameter affine mode, and mv 0 and mv 1 are available, they are added as one candidate in the affine AMVP list. When the current CU is coded with 6-parameter affine mode, and all three CPMVs are available, they are added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable.

[0127] If the affine AMVP list candidate is still less than 2 after the inherited affine AMVP candidates and constructed AMVP candidate are checked, mv 0, mv 1 and Figure 16 2 are added in order as translation MVs that predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, a zero MV is used to fill the affine AMVP list.

[0128] 2.10.3. Affine motion information storage In VVC, the CPMVs of an affine CU are stored in a separate buffer. The stored CPMVs are only used to generate inherited CPMVs in affine Merge mode and affine AMVP mode for the recently coded CUs. The subblock MVs derived from the CPMVs are used for motion compensation, MV derivation for the Merge / AMVP list of translational MVs, and deblocking.

[0129] To avoid picture line buffer for additional CPMVs, the affine motion data inheritance from a CU of the above CTU is handled differently from the inheritance from regular neighboring CUs. If the candidate CU for affine motion data inheritance is in the above CTU line, the bottom-left and bottom-right subblock MVs in the line buffer, not the CPMVs, are used for affine MVP derivation. In this way, the CPMVs are only stored in the local buffer. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to the 4-parameter model. Figure 16 The motion vector usage for the proposed combined method is shown. As shown in Figure 17 the bottom-left and bottom-right subblock motion vectors of a CU are used for affine inheritance of the CUs in the bottom CTU along the top CTU boundary.

[0130] 2.10.4. Prediction refinement with optical flow for affine mode Compared with pixel-based motion compensation, subblock-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of prediction accuracy loss. To achieve more fine-grained motion compensation, prediction refinement with optical flow (PROF) is used to refine the subblock-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after the subblock-based affine motion compensation is performed, the luma prediction samples are refined by adding the difference derived by the optical flow equation. PROF is described as the following four steps: Step 1) Subblock-based affine motion compensation is performed to generate subblock prediction .

[0131] Step 2) The spatial gradient of the subblock prediction is calculated at each sample position using a 3-tap filter [-1, 0, 1] and . The gradient calculation is exactly the same as in BDOF.

[0132] (2-10) (2-11) used to control the precision of the gradient. The subblock (i.e. 4x4) prediction is extended by one sample on each side for the gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extended boundaries are copied from the nearest integer pixel positions in the reference picture.

[0133] Step 3) The luminance prediction refinement is calculated by the following optical flow equation.

[0134] (2-12) where is the sample MV (denoted by ) calculated for the sample position is the difference between the sample MV and the subblock MV of the subblock the sample belongs to, as shown in Figure 17 is quantized in 1 / 32 luminance sample precision. Figure 19 The subblock MV VSB and the pixel Δv(i,j) are shown.

[0135] Since the affine model parameters and the sample position relative to the subblock center do not change from one subblock to another, the can be calculated for the first subblock and reused for other subblocks in the same CU. Let and be the horizontal and vertical offsets from the sample position to the center of the subblock , the can be derived by the following equations: (2-13) (2-14) To maintain accuracy, the center of the subblock is calculated as ((W SB 1) / 2, (H SB 1) / 2), where W SB and H SB are the width and height of the subblock.

[0136] For the 4-parameter affine model, (2-15) For the 6-parameter affine model, (2-16) where , , are the motion vectors of the top-left, top-right and bottom-left control points,​ and W and H are the width and height of the CU.

[0137] Step 4) Finally, the luminance prediction refinement is added to the sub-block prediction The final prediction I’ is generated as the following equation.

[0138] (2-17) PROF is not applied to an affine coded CU in the following two cases: 1) all control point MVs are the same, which indicates the CU has only translational motion; 2) the affine motion parameters are larger than a specified limit, because the sub-block based affine MC is downgraded to CU based MC to avoid large memory access bandwidth requirement.

[0139] A fast coding method is applied to reduce the coding complexity of affine motion estimation with PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) if the CU is not the root block and its parent block does not select affine mode as its best mode, PROF is not applied because the probability that the current CU selects affine mode as the best mode is low; b) if the amplitudes of the four affine parameters (C, D, E, F) are all smaller than a predefined threshold, and the current picture is not a low-delay picture, PROF is not applied because the improvement introduced by PROF is small in this case. In this way, the affine motion estimation with PROF can be accelerated.

[0140] 2.11. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports a sub-block based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in a co-located picture to improve the motion vector prediction and the Merge mode for a CU in the current picture. The same co-located picture used by TMVP is used for SbTVMP. SbTMVP is distinguished from TMVP in the following two main aspects: - TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; - While TMVP obtains the temporal motion vector from a co-located block in the co-located picture (the co-located block is the right-bottom or center block relative to the current CU), SbTMVP applies a motion shift before obtaining the temporal motion information from the co-located picture, where the motion shift is obtained from a motion vector from one of the spatial neighboring blocks of the current CU.

[0141] The SbTMVP process is illustrated in FIG. 18. SbTMVP predicts the motion vector of a sub-CU within the current CU in two steps. In the first step, the spatial neighbor Al in (a) is examined. If Al has a motion vector using a collocated picture as its reference picture, then that motion vector is selected as the motion shift to be applied. If no such motion is identified, then the motion shift is set to (0, 0).

[0142] In the second step, the motion shift identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the collocated picture as shown in (b). The example in (b) assumes that the motion shift is set to the motion of block Al. Then, for each sub-CU, the motion information of its corresponding block (covering the minimum motion grid of the center sample) in the collocated picture is used to derive the motion information for the sub-CU. After the motion information of the collocated sub-CU is identified, it is converted to the motion vector and reference index of the current sub-CU in a similar way as the TMVP process of HEVC, where the temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.

[0143] In VVC, the combination of the subblock-based Merge list containing both SbTMVP candidates and affine Merge candidates is used for the signaling of the subblock-based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry of the list of subblock-based Merge candidates, followed by the affine Merge candidates. The size of the subblock-based Merge list is signaled in the SPS, and the maximum allowed size of the subblock-based Merge list is 5 in VVC.

[0144] The sub-CU size used in SbTMVP is fixed to 8x8, and as with the affine Merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8.

[0145] The coding logic of the additional SbTMVP Merge candidate is the same as that of other Merge candidates, i.e., for each CU in a P slice or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.

[0146] 2.12. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in slice header is equal to 0, the motion vector difference (MVD) between the motion vector of a CU and the predicted motion vector is signaled in quarter luma sample units. In VVC, the CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be coded with different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows: - Normal AMVP mode: quarter luma sample, half luma sample, integer luma sample, or four luma sample.

[0147] - Affine AMVP mode: quarter luma sample, integer luma sample, or 1 / 16 luma sample.

[0148] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both horizontal and vertical MVDs for reference list L0 and reference list L1) are zero, quarter luma sample MVD resolution is inferred.

[0149] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luma sample MVD precision is used for the CU. If the first flag is 0, no further signaling is needed and quarter luma sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half luma sample or other MVD precision (integer or four luma sample) is used for normal AMVP CU. In the case of half luma sample, a 6-tap interpolation filter instead of the default 8-tap interpolation filter is used for half luma sample positions. Otherwise, a third flag is signaled to indicate whether integer luma sample or four luma sample MVD precision is used for normal AMVP CU. In the case of affine AMVP CU, the second flag is used to indicate whether integer luma sample or 1 / 16 luma sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter luma sample, half luma sample, integer luma sample, or four luma sample), the motion vector predictor of the CU will be rounded to the same precision as the MVD before being added with the MVD. The motion vector predictor is rounded to zero (i.e., negative motion vector predictor is rounded to positive infinity and positive motion vector predictor is rounded to negative infinity).

[0150] The encoder uses RD check to determine the motion vector resolution for the current CU. To avoid performing four times of CU-level RD check for each MVD resolution, in VTM11, the RD check for MVD precision other than quarter luma sample is only conditionally invoked. For normal AMVP mode, the RD cost of quarter luma sample MVD precision and integer luma sample MV precision are first calculated. Then, the RD cost of integer luma sample MVD precision is compared with the RD cost of quarter luma sample MVD precision to decide whether it is necessary to further check the RD cost of quarter luma sample MVD precision. When the RD cost of quarter luma sample MVD precision is much smaller than the RD cost of integer luma sample MVD precision, the RD check of quarter luma sample MVD precision is skipped. Then, the check of half luma sample MVD precision is skipped if the RD cost of integer luma sample MVD precision is significantly larger than the best RD cost of previously tested MVD precisions. For affine AMVP mode, if the affine inter mode is not selected after checking the rate-distortion cost of affine Merge / skip mode, Merge / skip mode, quarter luma sample MVD precision normal AMVP mode and quarter luma sample MVD precision affine AMVP mode, the 1 / 16 luma sample MV precision and 1-pixel MV precision affine inter mode are not checked. In addition, in 1 / 16 luma sample and quarter luma sample MV precision affine inter mode, the affine parameters obtained in quarter luma sample MV precision affine inter mode are used as the starting search points.

[0151] 2.13. Bi-prediction with CU-level weights (BCW) In HEVC, bi-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, bi-prediction mode is extended beyond simple averaging to allow weighted averaging of two prediction signals.

[0152]

[0153] Five weights are allowed in weighted average bi-prediction, For each bi-predicted CU, the weight w is determined in one of two ways: 1) for non-Merge CU, the weight index is signaled after the motion vector difference; 2) for Merge CU, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w e {3, 4, 5}) are used.

[0154] - At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. For more details, the reader can refer to the VTM software. When combined with AMVR, if the current picture is a low-delay picture, unequal weights are only conditionally checked for 1-pel and 4-pel motion vector precision.

[0155] - When combined with affine, affine ME for unequal weights is performed if and only if the affine mode is selected as the current best mode.

[0156] - When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked.

[0157] - When certain conditions are met, unequal weights are not searched, depending on the POC distance between the current picture and its reference pictures, the coding QP, and the temporal level.

[0158] The BCW weight index is coded using one context-coded bin followed by bypass-coded bins. The first context-coded bin indicates whether equal weights are used; and if unequal weights are used, additional bins are signaled using bypass coding to indicate which unequal weight is used.

[0159] Weighted prediction (WP) is a coding tool supported by H.264 / AVC and HEVC standards for efficiently coding video content with gradual transitions. Support of WP is also added to the VVC standard. WP allows the signaling of a weighting parameter (weight and offset) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weight(s) and offset(s) of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid the interaction between WP and BCW, which would complicate the VVC decoder design, if a CU uses WP, the BCW weight index is not signaled and w is assumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is assumed from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using constructed affine Merge mode is simply set equal to the BCW index of the first control point MV.

[0160] In VVC, CIIP and BCW cannot be jointly applied for a CU. When a CU is coded in CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights.

[0161] 2.14. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is a coding tool to address the local illumination change between the current picture and its temporal reference picture. LIC is based on a linear model, where scaling factors and offsets are applied to the reference samples to obtain the prediction samples of the current block. Specifically, LIC can be mathematically modeled by the following equation:

[0162] where is the prediction signal of the current block at coordinates ; is the reference block pointed by the motion vector ; and are the corresponding scaling factors and offsets applied to the reference block. Figure 19 The LIC process is shown. In Figure 19 , when LIC is applied to a block, a least mean square error (LMSE) method is employed to derive the values of the LIC parameters (i.e., Figure 19 and T ) by minimizing the difference between the neighboring samples of the current block (i.e., Figure 19 in T0 ) and their corresponding reference samples (i.e., T1 in ) in the temporal reference picture. In addition, to reduce the computational complexity, both the template samples and the reference template samples are down-sampled (adaptive down-sampling) to derive the LIC parameters, i.e., only the shaded samples in and Figure 20 are used to derive and .

[0163] To improve the coding performance, as shown in Figure 21 , no down-sampling is performed for the short side.

[0164] 2.15. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the MVs of the Merge mode, a decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bi-prediction operation, around the initial MV in the reference picture list L0 and the reference picture list L1, a refined MV is searched. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and list L1. As shown in Figure 22As shown, the SAD between two blocks based on each MV candidate (e.g., MV0' and MV1') around the initial MV is computed. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predicted signal.

[0165] In VVC, the application of DMVR is restricted, and is only applied to CUs coded with the following modes and features: - CU-level Merge mode with bi-predicted MV.

[0166] - One reference picture is past and the other reference picture is future relative to the current picture.

[0167] - The distance (i.e., POC difference) from the two reference pictures to the current picture is the same.

[0168] - Both reference pictures are short-term reference pictures.

[0169] - The CU has more than 64 luma samples.

[0170] - The CU height and CU width are both greater than or equal to 8 luma samples.

[0171] - The BCW weight index indicates equal weights.

[0172] - WP is not enabled for the current block.

[0173] - CIIP mode is not used for the current block.

[0174] The refined MV derived by the DMVR process is used to generate the inter- predicted samples, and is also used in temporal motion vector prediction for future picture coding. While the original MV is used in the deblocking process, and is also used in spatial motion vector prediction for future CU coding.

[0175] Additional features of DMVR are mentioned in the following sub-items.

[0176] 2.15.1. Search scheme In DVMR, the search points are around the initial MV, and the MV offsets follow the MV difference mirroring rule. In other words, any point (denoted by the candidate MV pair (MV0, MV1)) examined by DMVR follows the following two equations: (2-19) (2-20) where denotes the refinement offset between the initial MV and the refined MV in one of the two reference pictures. The refinement search range is two integer luma samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase.

[0177] A 25-point full search is applied to the integer sample offset search. The SAD of the initial MV pair is first computed. If the SAD of the initial MV pair is less than a threshold, the integer sample phase of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is computed and checked in a raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search phase. To reduce the loss of uncertainty of DMVR refinement, it is proposed to bias the original MV forward during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidate is reduced by 1 / 4 of the SAD value.

[0178] The integer sample search is followed by a fractional sample refinement. To save computational complexity, the fractional sample refinement is derived by using a parametric error surface equation, instead of by an additional search with SAD comparison. The fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. The fractional sample refinement is further applied when the integer sample search phase is terminated with the center having the smallest SAD in the first or second iteration search.

[0179] In the parametric error surface based sub-pixel offset estimation, the center position cost and the costs at the four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form (2-21) where (x, y) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of the five search points, (x, y) is computed as: is computed as: (2-22) (2-23) and The values of (x, y) are automatically constrained between -8 and 8, because all cost values are positive, and the minimum value is This corresponds to the half-pixel offset in VVC with 1 / 16-pixel MV precision. The computed fractional (x, y) is added to the integer distance refined MV to get the sub-pixel accurate refinement delta MV.

[0180] 2.15.2. Bilinear Interpolation and Sample Padding ​​In VVC, the resolution of MV is 1 / 16 luma samples. The samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are around the initial fractional pixel MV with integer sample offset, so the samples at those fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, with 2-sample search range, DMVR does not access more reference samples compared to the normal motion compensation process. After the refined MV is obtained by the DMVR search process, the normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples than the normal MC process, the samples needed by the interpolation process based on the original MV do not need to be filled from those available samples but the interpolation process based on the refined MV does.

[0181] 2.15.3. Maximum DMVR processing unit When the width and / or height of a CU is larger than 16 luma samples, it will be further divided into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16.

[0182] 2.16. Multi-pass decoder-side motion vector refinement In this contribution, a multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16x16 sub-block within the coded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bi-directional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction.

[0183] 2.16.1. First pass - block-based bilateral matching MV refinement In the first pass, the refined MVs are derived by applying BM to the coded block. Similar to decoder-side motion vector refinement (DMVR), the refined MVs are searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.

[0184] The BM performs a local search to derive the integer sample precision intDeltaMV and the half-pel sample precision halfDeltaMv. The local search applies a 3x3 square search pattern to loop over a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum values of sHor and sVer are 8.

[0185] The bilateral matching cost is computed as: bilCost = mvDistanceCost + sadCost. When the block size cbW When cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between the reference blocks. The intDeltaMV or halfDeltaMV local search is terminated when the bilCost at the center point of the 3x3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern and the search for the minimum cost continues until it reaches the end of the search range.

[0186] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: • MV0_pass1 = MV0 + deltaMV • MV1_pass1 = MV1 - deltaMV 2.16.2. Second pass - sub-block based bilateral matching MV refinement In the second pass, the refined MV is derived by applying the BM to 16x16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) for the reference picture lists L0 and L1 obtained by the first pass. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.

[0187] For each sub-block, the BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum values of sHor and sVer are 8.

[0188] The bilateral matching cost is computed by applying a cost factor to the SATD cost between the two reference sub-blocks as follows: bilCost = satdCost costFactor. The search area (2 sHor + 1) (2 sVer + 1) is divided into Figure 22 up to 5 diamond shaped search areas as shown. Each search area is assigned a costFactor which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in a raster scan order from the top-left corner to the bottom-right corner of the area. If the minimum bilCost within the current search area is less than a threshold (which is equal to sbW sbH), the integer-pel full search is terminated, otherwise, the integer-pel full search continues to the next search area until all search points are checked. Figure 23 The diamond areas in the search area are shown.

[0189] The BM performs a local search to derive the half-pel precision halfDeltaMv. The search pattern and cost function are the same as defined in 2.9.1.

[0190] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV(sbIdx2). The refined MV at the second pass is then derived as: • MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2) • MV1_pass2(sbIdx2) = MV1_pass1 – deltaMV(sbIdx2) 2.16.3. Third pass - subblock-based bilateral optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8x8 grid subblocks. For each 8x8 subblock, the BDOF refinement is applied to derive the scaled Vx and Vy without clipping from the refined MV of the parent subblock at the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.

[0191] The refined MV at the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as: • MV0_pass3(sbldx3) = MV0_pass2(sbldx2) + bioMv • MV1_pass3(sbldx3) = MV0_pass2(sbldx2) - bioMv 2.17. Sample-based BDOF In sample-based BDOF, motion refinements (Vx, Vy) are not derived based on blocks, but are performed for each sample.

[0192] A coded block is divided into 8x8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to a sub-block, for each sample in the sub-block, a 5x5 sliding window is used, and for each sliding window, the existing BDOF process is applied to derive Vx and Vy. The derived motion refinements (Vx, Vy) are applied to adjust the bi-predicted sample value for the center sample of the window.

[0193] 2.18. Extended Merge prediction In VVC, the Merge candidate list is constructed by including the following 5 types of candidates in order: (1) Spatial MVP from spatial neighboring CUs.

[0194] (2) Temporal MVP from collocated CUs.

[0195] (3) History-based MVP from a FIFO table.

[0196] (4) Pairwise average MVP.

[0197] (5) Zero MV.

[0198] The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coded in Merge mode, the index of the best Merge candidate is coded using truncated unary binarization (TU). The first bin of the Merge index is coded with a context, and bypass coding is used for the other bins.

[0199] This section provides the derivation process of each type of Merge candidate. As done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain size region.

[0200] 2.18.1. Spatial candidate derivation The derivation of spatial Merge candidates in VVC is the same as in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates located at Figure 24 positions shown in FIG. 6, at most four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more than one CU at positions B0, A0, B1, A1 is not available (e.g., because it belongs to another slice or tile) or is intra coded. After the addition of the candidate at position A1, a redundancy check is performed for the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thus improving coding efficiency. To reduce the computational complexity, all possible pairs of candidates are not considered in the mentioned redundancy check. Instead, only pairs linked by an arrow in FIG. 6 are considered, and only if the corresponding candidates for the redundancy check do not have the same motion information, the candidate is added to the list. Figure 25

[0201] 2.18.2. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on a collocated CU belonging to a collocated reference picture. The reference picture list to be used for the derivation of the collocated CU is explicitly signaled in the slice header. As shown by the dashed line in FIG. 7, the scaled motion vector of the temporal Merge candidate is obtained by scaling the motion vector of the collocated CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set equal to 0. Figure 26

[0202] As shown in FIG. 8, the position for the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is not available, intra coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal Merge candidate. Figure 27

[0203] 2.18.3. History-based Merge candidate derivation ​​​A history-based MVP (HMVP) Merge candidate is added to the Merge list, after the spatial MVP and the TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as MVPs for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (emptied). As long as there is a non-subblock inter coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.

[0204] The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table and all the HMVP candidates after it are moved forward and the identical HMVP is inserted to the last entry of the table.

[0205] The HMVP candidates can be used in the Merge candidate list construction process. The latest few HMVP candidates in the table are checked in order and inserted into the candidate list, after the TMVP candidates. A redundancy check is applied to the HMVP candidates against the spatial or temporal Merge candidates.

[0206] To reduce the number of redundancy check operations, the following simplification is introduced: The number of HMPV candidates used for Merge list generation is set to (N <= 4)? M : (8 N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table.

[0207] The Merge candidate list construction process from HMVP is terminated once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1.

[0208] 2.18.4. Pairwise average Merge candidate derivation Pairwise average candidates are generated by averaging predefined pairs of candidates in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vectors are calculated separately for each reference list. If both motion vectors are available in one list, they are averaged even if they point to different reference pictures; if only one motion vector is available, it is used directly; if no motion vector is available, the list is kept invalid.

[0209] When the Merge list is not full after adding the pairwise average Merge candidates, a zero MVP is inserted at the end until the maximum number of Merge candidates is reached.

[0210] 2.18.5. Merge estimation region The Merge estimation region (MER) allows the Merge candidate list to be derived independently for CUs in the same Merge estimation region (MER). A candidate block that is within the same MER as the current CU is not included for the generation of the Merge candidate list of the current CU. In addition, the update process of the history-based motion vector predictor candidate list is only updated if ( xCb + cbWidth ) » Log2ParMrgLevel is greater than xCb » Log2ParMrgLevel and ( yCb + cbHeight ) » Log2ParMrgLevel is greater than ( yCb » Log2ParMrgLevel ), where ( xCb, yCb ) is the top-left luma sample position of the current CU in the picture and ( cbWidth, cbHeight ) is the CU size. The MER size is selected at the encoder side and is signaled in the sequence parameter set as log2_parallel_merge_level_minus2.

[0211] 2.19. New Merge candidate 2.19.1. Non-adjacent Merge candidate derivation In VVC, Figure 28 The 5 spatial neighboring blocks and 1 temporal neighbor shown are used to derive the Merge candidate.

[0212] It is proposed to derive additional Merge candidates from positions that are non-adjacent to the current block using the same modes as in VVC. To achieve this, for each round of search i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated by the following equation: Offsetx = -i x gridX, Offsety = -i x gridY where Offsetx and Offsety represent the offset of the top-left corner of the virtual block to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.

[0213] Second, the width and height of the virtual block are calculated by the following equation: newWidth = i x 2 x gridX + currWidth, newHeight = i x 2 x gridY + currHeight.

[0214] where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.

[0215] gridX and gridY are currently set to currWidth and currHeight, respectively.

[0216] Figure 29 is an illustration of the virtual block in the i-th search round. After the virtual block is generated, blocks A i , B i , C i , D i , and E i can be regarded as the VVC spatial neighboring blocks of the virtual block, and their positions are obtained with the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, blocks A i , B i , C i , D i , and E i are the spatial neighboring blocks used in the VVC Merge mode.

[0217] When constructing the Merge candidate list, de-duplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that 5 non-adjacent spatial neighboring blocks are utilized.

[0218] The non-adjacent spatial Merge candidates are inserted into the Merge list after the temporal Merge candidates in the order of B1->A1->C1->D1->E1.

[0219] 2.19.2. STMVP It is proposed to use three spatial Merge candidates and one temporal Merge candidate to derive the average candidate as the STMVP candidate.

[0220] STMVP is inserted before the top-left spatial Merge candidate.

[0221] The STMVP candidate is de-duplicated together with all previous Merge candidates in the Merge list.

[0222] For spatial candidates, the first three candidates in the current Merge candidate list are used.

[0223] For temporal candidates, the same position as the VTM / HEVC collocated position is used.

[0224] For spatial candidates, the first, second and third candidates inserted in the current Merge candidate list before STMVP are denoted as F, S and T.

[0225] The temporal candidate with the same position as the VTM / HEVC collocated position used in TMVP is denoted as Col.

[0226] The motion vector of the STMVP candidate in prediction direction X (denoted as mvLX) is derived as follows: 1) if the reference indices of the four Merge candidates are all valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_S + mvLX_T + mvLX_Col)>>2 2) if the reference indices of three of the four Merge candidates are valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F x 3 + mvLX_S x 3 + mvLX_Col x 2)>>3 or mvLX = (mvLX_F x 3 + mvLX_T x 3 + mvLX_Col x 2)>>3 or mvLX = (mvLX_S x 3 + mvLX_T x 3 + mvLX_Col x 2)>>3 3) if the reference indices of two of the four Merge candidates are valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_Col)>>1 or mvLX = (mvLX_S + mvLX_Col)>>1 or mvLX = (mvLX_T + mvLX_Col) » 1 NOTE: If temporal candidates are not available, the STMVP mode is turned off.

[0227] 2.19.3. Merge list size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header and the maximum allowed size of the Merge list is 8.

[0228] 2.20. Geometric partition mode (GPM) In VVC, a geometric partition mode is supported for inter prediction. The geometric partition mode is signaled as a kind of Merge mode using a CU-level flag, where other Merge modes include regular Merge mode, MMVD mode, CIIP mode and subblock Merge mode. For each possible CU size (where excluding 8x64 and 64x8), the geometric partition mode supports 64 partitions in total.

[0229] Figure 29 An example of GPM partitioning grouped with the same angle is shown. When this mode is used, the CU is partitioned into two parts by a geometrically positioned straight line Figure 30 ). The position of the partition line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter predicted using its own motion; only uni-prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure the same as regular bi-prediction, i.e., only two motion- compensated predictions are needed for each CU. The uni-prediction motion of each partition is derived using the process described in 2.20.1.

[0230] If the geometric partition mode is used for the current CU, further two Merge indices (one for each partition) and the geometric partition index of the partition mode (angle and offset) of the geometric partition are signaled. The number of maximum GPM candidate size is explicitly signaled in the SPS and specifies the syntax binarization for the GPM Merge indices. After predicting each part of the geometric partition, a hybrid process with adaptive weights as in 2.20.2 is used to adjust the sample values along the geometric partition edge. This is the prediction signal for the whole CU and the transform and quantization processes will be applied to the whole CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partition mode is stored as shown in 2.20.3.

[0231] 2.20.1. Unidirectional prediction candidate list construction The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 2.18. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate, where X equals the parity of n, is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are marked with "x" in Figure 31 . If the corresponding LX motion vector of the nth extended Merge candidate does not exist, the L(1 X) motion vector of the same candidate is used as the unidirectional prediction motion vector for the geometric partition mode.

[0232] 2.20.2. Blending along the geometric partition edge After each part of the geometric partition is predicted using its own motion, blending is applied to the two prediction signals to derive the samples around the geometric partition edge. The blending weight for each location of the CU is derived based on the distance between the single location and the partition edge.

[0233] Location The distance to the partition edge is derived as: (2-24) (2-25) (2-26) (2-27) where is the index for the angle and offset of the geometric partition, which depends on the geometric partition index signaled through. and the sign of .

[0234] The weight of each part of the geometric partition is derived as follows: (2-28) (2-29) (2-30) partldx depends on the angle index . An example of the weight is shown in Figure 31 . Otherwise, if Mvl and Mv2 are from the same list, only the uni-predicted motion Mv2 is stored. An exemplary generation of the bending weight using the geometric partition mode is shown.

[0235] 2.20.3. Motion field storage for geometric partition mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and a combined Mv of Mv1 and Mv2 are stored in the motion field of the CU coded with the geometric partition mode.

[0236] The motion vector type stored for each individual position in the motion field is determined as: (2-31) where motionldx is equal to which is re-computed from equation (2-18). partldx depends on the angle index .

[0237] If sType is equal to 0 or 1, Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, a combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following process: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), Mv1 and Mv2 are simply combined to form a bi-predictive motion vector.

[0238] Figure 32

[0239] 2.21. Multi-hypothesis prediction In multi-hypothesis prediction (MHP), up to two additional prediction values are signaled in addition to the inter AMVP mode, regular Merge mode, affine Merge and MMVD modes. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.

[0240]

[0241] Weighting factors α are specified according to the following Table 2-4: Table 2-4 - Weighting factors for MHP

[0242] For the inter AMVP mode, MHP is applied only when non-equal weights in BCW are selected in bi-predictive mode.

[0243] The additional assumption can be either Merge mode or AMVP mode. In the case of Merge mode, the motion information is indicated by the Merge index and the Merge candidate list is the same as in the geometric partition mode. In the case of AMVP mode, the reference index, the MVP index and the MVD are signaled.

[0244] 2.22. Non-adjacent spatial candidates Figure 32 The spatial neighboring blocks used to derive the spatial Merge candidates are shown. The non-adjacent spatial Merge candidates are inserted after the TMVP in the regular Merge candidate list. The pattern of the spatial Merge candidates is shown in Figure 33 . The distance between the non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block.

[0245] 2.23. Template matching (TM) Template matching (TM) is a decoder-side MV derivation method to refine the motion information of a current CU by finding the closest match between a template (i.e., the top and / or left neighboring blocks of the current CU) in the current picture and a block (i.e., of the same size as the template) in the reference picture. As shown in Figure 34 , within a search range of [-8, +8] pixels around the initial motion of the current CU, a better MV will be searched. Template matching proposes two modifications: determining the search step size based on the AMVR mode, and in Merge mode TM can be cascaded with the bilateral matching process.

[0246] In AMVP mode, the MVP candidate is determined based on the template matching error to pick one that achieves the minimum difference between the current block template and the reference block template, and then TM only performs MV refinement for that specific MVP candidate. TM refines this MVP candidate by using an iterative diamond search, starting from the full-pel MVD precision (or 4-pixel for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined by using a cross search with full-pel MVD precision (or 4-pixel for 4-pixel AMVR mode), and then quarter-pel and half-pel in turn according to the AMVR mode specified in Table 2-5. This search process ensures that the MVP candidate still maintains the same MV precision as indicated by the AMVR mode after the TM process.

[0247] Table 2-5 - Search pattern for AMVR and search pattern for Merge mode with AMVR

[0248] In Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 2-5, TM can proceed up to 1 / 8 pixel MVD precision, or skip those precisions beyond half-pixel MVD precision, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in half-pixel mode). Furthermore, when TM mode is enabled, template matching can operate as a standalone process or as an additional MV refinement process between the block-based bilateral matching (BM) method and the sub-block-based bilateral matching (BM) method, depending on whether BM can be enabled according to its enable condition check.

[0249] 2.24. Overlapping Block Motion Compensation (OBMC) Overlapping Block Motion Compensation (OBMC) has been previously used in H.263. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, it is performed on all motion compensation (MC) block boundaries except for the right and bottom boundaries of the CU. Furthermore, it is applied to both the luma and chroma components. In JEM, MC blocks correspond to codec blocks. When a CU is encoded and decoded in sub-CU modes (including sub-CU merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To handle CU boundaries uniformly, OBMC is performed at the sub-block level for all MC block boundaries, where the sub-block size is set to equal to 4×4, such as... Figure 34 As shown. Above This is a schematic diagram of a sub-block in the OBMC application.

[0250] When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of the four connected neighboring sub-blocks (if available and different from the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block.

[0251] The predicted block based on the motion vectors of neighboring sub-blocks is represented as P N ,in N Indicates targeting neighboring Below , Above Left , Right and Figure 35 The index of the sub-block, and the predicted block based on the motion vector of the current sub-block is represented as: P C .when P N When based on the motion information of neighboring sub-blocks that contain the same motion information as the current sub-block, the OBMC is not removed from...P N is added to P N . Otherwise, P C the same sample in P N is added to P C . Weight factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N and weight factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . The exception is for small MC blocks (i.e., when the height or width of the coded block is equal to 4 or the CU is coded in sub-CU mode), for which only P N rows / columns of P C are added to P N . In this case, weight factors {1 / 4, 1 / 8} are used for P C and weight factors {3 / 4, 7 / 8} are used for P N . For P N generated based on the motion vector of the vertically (horizontally) neighboring sub-block, the samples in the same row (column) of P C are added to

[0252] In JEM, for CUs with size smaller than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU or not. For CUs with size larger than 256 luma samples or not coded in AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top and left neighboring blocks is used to compensate the top and left boundaries of the original signal of the current CU, then the normal motion estimation process is applied.

[0253] 2.25. Multiple Transform Selection (MTS) for core transform In addition to DCT-II already used in HEVC, a multiple transform selection (MTS) scheme is used for residual coding of both inter-coded blocks and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the transform basis functions of the selected DST / DCT.

[0254] Table 2-6 - Transform basis functions of DCT-II / VIII and DST VII for N-point input

[0255] To preserve the orthogonality of the transform matrices, the transform matrices are quantized more precisely than the transform matrices in HEVC. To keep the mid value of the transform coefficients within the 16-bit range, all coefficients must have 10 bits after the horizontal transform and after the vertical transform.

[0256] To control the MTS scheme, separate enabling flags are specified at the SPS level for intra coding and inter coding, respectively. When MTS is enabled at the SPS, CU level flags are signaled to indicate whether MTS is applied or not. Here, MTS is only applied to luma. The MTS signaling is skipped when one of the following conditions is applied.

[0257] - The position of the last significant coefficient for luma TB is less than 1 (i.e., DC only).

[0258] - The last significant coefficient of luma TB is located in the MTS zero-out region.

[0259] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are additionally signaled to indicate the transform type for the horizontal direction and the vertical direction, respectively. The transform and signaling mapping table is shown in Table 2-7. A unified transform selection for ISP and implicit MTS is used by removing the intra mode and block shape dependency. If the current block is in ISP mode, or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are on, only DST7 is used for both the horizontal transform kernel and the vertical transform kernel. When it comes to transform matrix precision, 8-bit primary transform kernels are used. Therefore, all transform kernels used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8) use 8-bit primary transform kernels.

[0260] Table 2-7 –Transform and signaling mapping table

[0261] To reduce the complexity of large size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with size (width or height, or both width and height) equal to 32, high frequency transform coefficients are zeroed. Only the coefficients within the 16x16 low frequency region are kept.

[0262] As in HEVC, the residual of a block can be coded in transform skip mode. To avoid the redundancy of syntax coding, the transform skip flag is not signaled when CU level MTS CU flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for inter coded blocks, implicit MTS can still be enabled.

[0263] 2.26. Sub-block transform (SBT) In VTM, sub-block transform is introduced for inter predicted CUs. In this transform mode, only a sub-part of the residual block is coded for a CU. When a CU has cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the whole residual block or a sub-part of the residual block is coded. In the former case, inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is coded with a deterministic adaptive transform and another part of the residual block is zeroed.

[0264] When SBT is used for an inter coded CU, SBT type and SBT position information are signaled in the bitstream. As shown in Figure 35 there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half or ¼ of the CU width (or height), resulting in 2:2 partition or 1:3 / 3:1 partition. 2:2 partition is like binary tree (BT) partition, while 1:3 / 3:1 partition is like asymmetric binary tree (ABT) partition. In ABT partition, only the small region contains non-zero residual. If one dimension of the CU is 8 in luma samples, 1:3 / 3:1 partition along that dimension is forbidden. There are at most 8 SBT modes for a CU.

[0265] Position dependent transform kernel selection is applied to luma transform blocks in SBT-V and SBT-H (chroma TB always uses DCT-2). The two positions of SBT-H and SBT-V are associated with different core transforms. More specifically, the horizontal transform and the vertical transform of each SBT position are in Figure 35The SBT position, type and transform type are shown in Table 2.1.1-1. Figure 36 The SBT position, type and transform type are shown in Table 2.1.1-1.

[0266] SBT is not applied to CUs coded in inter-intra joint mode.

[0267] 2.27. Template matching based adaptive Merge candidate reordering To improve coding efficiency, after constructing the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list according to ascending order of the template matching cost. It is operated in the form of subgroups.

[0268] The template matching cost is measured by the SAD (sum of absolute difference) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes motion information of bi-prediction, the corresponding reference samples are the average of the corresponding reference samples in reference list 0 and the corresponding reference samples in reference list 1, as shown in Figure 37 If the Merge candidate includes motion information at sub-CU level, the corresponding reference samples consist of the neighboring samples of the corresponding reference sub-blocks, as shown in Figure 38

[0269] The ordering process is operated in the form of subgroups, as shown in Figure 39 The first three Merge candidates are ordered together. The last three Merge candidates are ordered together.

[0270] The template size (width of the left template or height of the top template) is 1. The subgroup size is 3.

[0271] 2.28. Adaptive Merge candidate list It can be assumed that the number of Merge candidates is 8. We take the first 5 Merge candidates as the first subgroup and the last 3 Merge candidates as the second subgroup (i.e. the last subgroup).

[0272] For the encoder, after constructing the Merge candidate list, some Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, as shown in Figure 40

[0273] ​​More specifically, template matching cost is calculated for Merge candidates in all subgroups except the last one; then, Merge candidates in the own subgroup are reordered except the last one; finally, the final Merge candidate list is obtained.

[0274] For the decoder, after constructing the Merge candidate list, some / no Merge candidates are adaptively reordered in ascending order of Merge candidate cost, as shown in Figure 40 Figure 41 In the selected subgroup, the subgroup in which the selected (signaled) Merge candidate is located.

[0275] More specifically, if the selected Merge candidate is located in the last subgroup, after deriving the selected Merge candidate, the Merge candidate list construction process is terminated, no reordering is performed, and the Merge candidate list is not changed; otherwise, the process is performed as follows: After deriving all Merge candidates in the selected subgroup, the Merge candidate list construction process is terminated; template matching cost is calculated for Merge candidates in the selected subgroup; Merge candidates in the selected subgroup are reordered; finally, the new Merge candidate list is obtained.

[0276] For both the encoder and the decoder, the template matching cost is derived as a function of T and RT, where T is the set of samples in the template, and RT is the set of reference samples for the template.

[0277] When deriving the reference samples of the template of a Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.

[0278] The reference samples of the template (RT) for bi-prediction are derived by weighted average of the reference samples of the template in reference list 0 (L0) ) and the reference samples of the template in reference list 1 (L1) ).

[0279] (2-32) where the weight of the reference template in reference list 0 (8-w) and the weight of the reference template in reference list 1 (w) are determined by the BCW index of the Merge candidate. The BCW index equal to {0, 1, 2, 3, 4} corresponds to w equal to {-2, 3, 4, 5, 10}, respectively.

[0280] If the local illumination compensation (LIC) flag of the Merge candidate is true, the reference samples of the template are derived using the LIC method.

[0281] ​The template matching cost is computed based on the sum of absolute difference (SAD) of T and RT.

[0282] The template size is 1. This means that the width of the left template and / or the height of the top template is 1.

[0283] If the coding mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered.

[0284] If the coding mode is GPM, the Merge candidates used to derive the uni-prediction candidate list are not reordered.

[0285] 2.29. IBC with extended reference region An IBC reference region design is proposed that does not increase the current memory region required by ECM-3 and tests the performance.

[0286] Figure 42 The extended reference region is shown. In the figure, the blue square represents the current CTU, and the green square represents the CTU that can be used by IBC reference. Specifically, assuming W represents the maximum horizontal CTU index, and the current CTU index is (m, n), for the coding unit in the current CTU, the CTUs with indices (0, n)…(m, n) and (m-1, n)…(W, n) define the reference region that can be used by IBC.

[0287] One reason for having such a design is that in the current ECM, the left, top, and top-left CTUs are used, so they need to be saved. To achieve this, all the CTUs to the right of the top CTU in the top CTU row (for the CTU to be coded in the current CTU row) and all the CTUs to the left of the current CTU in the current CTU row (for the CTU to be coded in the next CTU row) must be kept. This means that such a design does not increase the buffer size required by the current ECM.

[0288] 2.30. IBC with template matching It is proposed to also use IBC with template matching for both IBC Merge mode and IBC AMVP mode.

[0289] The IBC-TM Merge list has been modified compared to the list used by the regular IBC Merge mode so that the candidates are selected according to a de-weighting method with the motion distance between the candidates in the regular TM Merge mode. The end zero motion satisfaction (which is meaningless with respect to intra coding) has been replaced by the motion vector to the left (-W, 0), above (0, -H) and top-left (-W, -H) CUs, then, if necessary, the list is satisfied with the one on the left without de-weighting.

[0290] In the IBC-TM Merge mode, the selected candidate is refined with a template matching method before the RDO or the decoding process. The IBC-TM Merge mode has been competing with the regular IBC Merge mode and the TM-Merge flag is signaled.

[0291] In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC Merge list. Each of the 3 selected candidates is refined using a template matching method and ordered according to the resulting template matching cost. Then, only the first 2 are usually considered in the motion estimation process.

[0292] The template matching refinement for both the IBC-TM Merge mode and the AMVP mode is quite simple because the IBC motion vector is constrained to be integer and within the reference region as shown in Figure 43 Therefore, in the IBC-TM Merge mode, all refinements are performed with integer precision while in the IBC-TM AMVP mode, they are performed with integer precision or 4-pixel precision. In both cases, the refined motion vector in each refinement step must respect the constraint of the reference region.

[0293] 2.31. Reconstructed Re-ordered IBC (RR-IBC) Figure 43 An example of symmetry in screen content pictures is shown. Screen content coding tools such as Intra Block Copy (IBC) generate a prediction block by directly copying a previously coded reference region in the same picture. Symmetry is often observed in video content, especially in text character regions and computer generated graphics in screen content sequences as shown in Figure 44a Therefore, a specific screen content coding tool that takes into account the symmetry would efficiently compress such video content.

[0294] For video encoding and decoding of screen content, a Reconstruction Reordering IBC (RR-IBC) mode is proposed. When applied, samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the predicted block is derived without flipping. On the decoder side, the reconstructed block is flipped back to recover the original block.

[0295] For blocks encoded with RR-IBC, two flipping methods are supported: horizontal flipping and vertical flipping. Syntax flags are first signaled for blocks encoded with IBC AMVP, indicating whether the reconstruction is flipped. If flipped, another flag is further signaled to specify the flipping type. For IBC Merge, the flipping type is inherited from neighboring blocks, with no syntax signaling. Considering horizontal or vertical symmetry, the current block and reference block are typically aligned horizontally or vertically. Therefore, when a horizontal flip is applied, the vertical component of the BV is not signaled and is presumed to be 0. Similarly, when a vertical flip is applied, the horizontal component of the BV is not signaled and is presumed to be 0.

[0296] To better utilize the symmetry property, a flip-aware BV adjustment method was applied to refine the block vector candidates. Figure 44b This is a diagram illustrating the adjustment of BV for horizontal flipping, and Figure 44a This is a diagram illustrating BV adjustments for vertical flipping. For example, as... Figure 44b and BV As shown, ( x nbr , y nbr )and( x cur , y cur () represents the coordinates of the center sample points of the neighboring blocks and the current block, respectively. BV nbr and BV cur These represent the BV of the neighboring block and the current block, respectively. This is relevant when the neighboring block is encoded and decoded using a horizontal flipping method. BV cur The horizontal component does not inherit BV directly from neighboring blocks, but rather through... BV nbr The horizontal component (represented as) BV nbr h The motion shift is added and calculated, i.e. BV cur h =2 ( x nbr - x cur )+BV nbr h Similarly, in case the neighboring block is coded with vertical flipping, BV cur The vertical component of BV nbr is calculated by adding the motion shift to BV nbr v , i.e. BV cur v = 2 ( y nbr - y cur ) + 2 ( Figure 45 nbr v .

[0297] 2.32. Intra Template Matching Prediction Intra Template Matching Prediction (Intra-TMP) is a special Intra prediction mode, which copies the best prediction block from the reconstructed part of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches for the template in the reconstructed part of the current frame that is most similar to the current template and uses the corresponding block as the prediction block. Then, the encoder signals the use of this mode and the same prediction operation is performed at the decoder side.

[0298] The prediction signal is generated by matching the L-shaped causal neighbors of the current block with another block in a predefined search region in Figure 45 , which consists of the following parts: R1 : the current CTU R2: the top-left CTU R3: the top CTU R4: the left CTU SAD is used as the cost function.

[0299] Within each region, the decoder searches for the template that has the smallest SAD with respect to the current template and uses its corresponding block as the prediction block.

[0300] The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a BlkW SearchRange_h = a BlkH where "K" is a constant controlling the gain / complexity trade-off. In practice, "K" is equal to 5.

[0301] Figure 46 The used intra template matching search region is shown. The intra template matching tool is enabled for CUs with width and height size smaller or equal to 64. This maximum CU size for intra template matching is configurable.

[0302] When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level by a dedicated flag.

[0303] 2.33. Intra prediction blending The intra prediction blending method uses multiple prediction values generated from different modes / reference lines.

[0304] In sub-test a, multiple intra prediction values are generated and then blended by weighted average. The process to derive the prediction values to be used in the blending process is described as follows: 1) For angular intra prediction modes including single mode cases of TIMD and DIMD, the proposed method derives the intra prediction by weighting the intra predictions obtained from multiple reference lines denoted as where is the intra prediction from the default reference line, and is the prediction from the line above the default reference line. The weights are set to and .

[0305] 2) For TIMD mode with mixing, is used for the first mode ( ), and is used for the second mode ( ).

[0306] 3) For DIMD mode with mixing, the number of prediction values selected for weighted average is increased from 3 to 6.

[0307] In sub-test b, the intra prediction blending is performed on the reference lines instead of the prediction block. Two reference lines (denoted as and ) are used for the intra prediction blending. The corresponding to the intra prediction angle is considered in the blending process. Each value in the blended reference line ( ) is derived from the following equation:

[0308] ​​When the intra-frame mode has a non-integer slope (required reference sample interpolation) and the block size is greater than 16, the proposed intra-frame prediction fusion is applied to the luma block, used together with MRL, but not to the ISP-encoded block. In the method studied in subtest a, PDPC is applied to the intra-frame prediction mode using the reference line closest to the current block.

[0309] 2.34. Template-based Multi-Reference Line Intra-Frame Prediction (TMRL) The proposed TMRL model includes the following aspects: a) Expanded reference line candidate list and intra-prediction mode candidate list The extended reference row candidate list used in this proposal is {1, 3, 5, 7, 12}. The constraint on the top CTU row remains unchanged. The size of the intra-prediction mode candidate list is 10. The construction of the intra-prediction mode candidate list is similar to that of MPM. The differences are: The PLANAR mode was excluded from the proposed list of intra-prediction mode candidates.

[0310] If not already included, the DC mode is added after the modes of the five adjacent PUs and the DIMD mode.

[0311] Added (compared to existing angle modes in the intra-prediction mode candidate list) with features from arrive The incremental angle mode.

[0312] b) Construction of the TMRL candidate list There are 5 × 10 = 50 combinations of extended reference lines and allowed intra-frame prediction modes for the block. Since the extended reference lines start from reference line 1, the region covered by reference line 0 is used for template matching. Between prediction (generated from the 50 combinations) and reconstruction, for the template region (see... Figure 47 The SAD cost of each combination is calculated. The 20 combinations with the lowest SAD cost are selected in ascending order to form the TMRL candidate list.

[0313] c) Signaling of TMRL Instead of directly encoding and decoding the reference line and intra-frame mode, the index of the TMRL candidate list is encoded and decoded to indicate which combination of the reference line and prediction mode is used to encode and decode the current block. In the proposed TMRL mode, rounding Golomb-Rice encoding and decoding with a divisor of 4 is used to encode and decode the combinations selected from the combination list. The binarization process and codewords are shown in Table 2-8.

[0314] Table 2-8 – Binarization process of TMRL index

[0315] d) Encoder-side modifications Encoder-side modifications were tested to further improve coding efficiency. For intra blocks larger than 8x8, an additional TMRL RDO was added if no TMRL mode was selected by the SATD comparison.

[0316] 2.35. Convolutional Cross-Component Model (CCCM) for Intra Prediction It was proposed to apply a Convolutional Cross-Component Model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit as done for the current CCLM mode. As for CCLM, when chroma downsampling is used, reconstructed luma samples are downsampled to match the lower resolution chroma grid.

[0317] In addition, similarly to CCLM, there is an option to have a single model or a multi-model variant with CCCM. The multi-model variant uses two models, one model is derived for samples above the average luma reference value and the other model is derived for the rest of the samples (in the spirit of the CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples.

[0318] 2.35.1. Convolutional filter The proposed convolutional 7-tap filter consists of a 5-tap plus sign-shaped spatial component, a non-linear term and a bias term. The input of the spatial 5-tap component of the filter consists of the center (C) luma sample co-located with the chroma sample to be predicted and its above / north (N), below / south (S), left / west (W) and right / east (E) neighbors as shown in Figure 48

[0319] The non-linear term P is expressed as the 2nd power of the center luma sample C and is scaled to the sample value range of the content: P = ( C C + midVal )>>bitDepth i.e. for 10-bit content, it is computed as: P = ( C C + 512 )>>10 The bias term B represents a scalar offset between the input and the output (similar to the offset term in CCLM) and is set to the mid chroma value (512 for 10-bit content).

[0320] The output of the filter is computed as the convolution of the filter coefficients c i and the input values, and is clipped to the range of valid chroma samples: ​predChromaVal = c0C + c1N + c2S + c3E + c4W + c5P + c6B 2.35.2. Calculation of filter coefficients Filter coefficients c i The filter coefficients are computed by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. Figure 49 A reference region consisting of 6 lines of chroma samples above and to the left of the PU is shown. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to only include available samples. The extension of the region shown in blue is needed to support the "side samples" of the shape-adaptive spatial filter and is padded when in unavailable regions.

[0321] The MSE minimization is performed by computing the auto-correlation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The auto-correlation matrix is LDL-decomposed and the final filter coefficients are computed using back-substitution. The process roughly follows the calculation of the ALF filter coefficients in the ECM, however, LDL-decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method only uses integer operations.

[0322] 2.35.3. Bitstream signaling The use of the mode is signaled by a PU-level flag that is CABAC coded. A new CABAC context is included to support this. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, the CCCM flag is only signaled when the intra prediction mode is either LM_CHROMA IDX (to enable single mode CCCM) or MM_LM_CHROMA IDX (to enable multi-model CCCM).

[0323] 2.36. Gradient Linear Model (GLM) In contrast to CCLM, GLM utilizes luma sample gradients to derive the linear model instead of down-sampled luma values. Specifically, when GLM is applied, the input of the CCLM process (i.e., the down-sampled luma samples ) are replaced by luma sample gradients . The other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged.

[0324]

[0325] For signaling, when CCLM mode is enabled for the current CU, two flags are signaled separately for Cb and Cr components to indicate whether GLM is enabled for each component; if GLM is enabled for one component, one syntax element is further signaled to select one of the 4 gradient filters for gradient calculation. Four gradient filters are enabled for GLM as shown in Figure 50

[0326] 2.37. Temporal block vector prediction Temporal BV prediction (TBVP) is proposed to mimic TMVP. Temporal BV candidates are introduced to the IBC Merge / AMVP candidate list. The IBC Merge candidate list includes regular IBC Merge candidate list, IBC-TM Merge candidate list and IBC-MBVD base candidate list.

[0327] TBVP candidates are derived from the same temporal location as TMVP with full de -repetition. They are placed before the HMVP candidates.

[0328] 2.38. Copy padding for IBC Copy padding can be applied to the overlapping region. Figure 50 It is shown that the unreconstructed samples in the reference block (shaded) are padded by copying their predicted samples. With copy padding, the unreconstructed samples in the overlapping region can be padded by copying its predicted samples as shown in Figure 51 P'(x,y) = P(x+BVx, y+BVy) where P'(x,y) is the padded sample at position (x,y), P(x+BVx, y+BVy) is the predicted sample, and (BVx, BVy) is the BV of the current block.

[0329] Copy padding is performed only when the horizontal BV component is less than or equal to 0 and the vertical BV component is less than or equal to 0.

[0330] 2.39. Direct block vector (DBV) mode for chroma prediction In ECM-6.0, the intra prediction modes for chroma components include 6 cross- component linear model (LM) modes, convolution cross-component model (CCCM) mode, gradient linear model (GLM) mode, DIMD mode, direct mode (DM) and four default intra prediction modes.

[0331] In the signaling of chroma intra modes, intra_chroma_pred_mode is signaled to indicate the specific coding mode as shown in Table 2-9. ​​

[0332] Table 2-9 - binarization process for intra_chroma_pred_mode in ECM6.0

[0333] ECM6.0 includes a method for chroma prediction using block vector in MODE IB C. In single tree partitioning, the prediction process of IBC is applied to both luma and chroma components, and the chroma block vector is derived from the corresponding luma block vector according to the chroma format sampling structure. In dual tree partitioning, the prediction process of IBC is applied to luma component only.

[0334] This contribution proposes a method for improving the coding efficiency of chroma intra prediction for screen content, i.e. the direct block vector (DBV) mode.

[0335] Figure 51 Five positions in the reconstructed luma samples are shown. For chroma components, when chroma dual tree is activated in intra slice, if one of the following five positions is coded with MODE IB C, its block vector bvL is used and scaled to derive the chroma block vector bvC. The scaling factor depends on the chroma format sampling structure. Figure 52

[0336] Then, by using the position of the current chroma block (xCb, yCb) and its bvC, the corresponding offset position (xCb+ bvC[0], yCb + bvC[1]) is determined, and the block copy prediction is performed as shown in Figure 52 . Figure 53 The prediction process of DBV mode is shown.

[0337] A CU level flag is signaled to indicate whether the proposed DBV mode is applied or not, as shown in Table 2-10.

[0338] Table 2-10 - binarization process for intra_chroma_pred_mode in the proposed method

[0339] 2.40. Filtering for IBC predicted blocks The proposed 6-tap filter consists of a 5-tap plus sign-shaped spatial component and a bias term. The input of the spatial 5-tap component of the filter consists of the center (C) samples in the reference block, which have corresponding positions with the samples in the current block to be predicted, and the above / north (N), below / south (S), left / west (W) and right / east (E) neighbors, as shown below. Intra block copy with affine model The spatial part of the filter is shown. ​

[0340] The bias term B represents a scalar offset between input and output, and is set to the middle luma value (512 for 10-bit content).

[0341] The output of the filter is computed as follows: predLumaVal = c0C + c1N + c2S + c3E + c4W + c5B.

[0342] The filter coefficients ci are computed by minimizing the MSE between the reference template and the current template, which is very similar to the CCCM process.

[0343] This filtered mode is used as an additional mode for non-Merge IBC blocks. For non-Merge blocks, this mode cannot be applied together with IBC-LIC, IBC-CIIP or RR-IBC. For IBC Merge mode, the filtering mode is inherited when the Merge mode list is constructed, so there is no extra signaling.

[0344] 3. Problem In the current design of Intra Block Copy (IBC) in video coding standards or video codecs (e.g. VVC, ECM), only translational motion within a picture is considered. Similar to inter prediction, an affine model can be used to improve coding performance to handle textures with zooming-in / out, rotation, perspective motion and other irregular motion.

[0345] 4. Detailed Solution The following detailed solutions should be considered as examples to explain the general concept. The embodiments should not be interpreted in a narrow way. Furthermore, the solutions can be combined in any way.

[0346] In this disclosure, Intra Block Copy (IBC) can not be limited to the current IBC technique, but can be interpreted as a technique that utilizes samples in the current slice / tile / subpicture / picture / other video unit (e.g. CTU row) to obtain a reference (or prediction) block, excluding the regular intra prediction method.

[0347] In the following discussion, IBC can be replaced by other coding tools that rely on coded / decoded, decoded or reconstructed information within the same region, e.g. palette, intra template matching.

[0348] In the following discussion, the block vector (BV) used in the affine model is denoted as control point block vector (CPBV).

[0349] bv (Presenting the proposed method as IBC-Affine mode) 1. It is proposed that a video unit can be divided into more than one sub-block, where the block vector (BV) for the sub-blocks can be derived using an affine model. The block size of the video unit is denoted as WxH.

[0350] a. In one example, one affine model can be applied to all sub-blocks.

[0351] b. In one example, at least two affine models can be applied to two sub-blocks.

[0352] i. Alternatively, additionally, the control point BVs for the two sub-blocks associated with the affine model can be different.

[0353] ii. Alternatively, additionally, each sub-block can be associated with its own affine model.

[0354] c. In one example, the video unit can be coded with IBC mode or Intra TMP mode.

[0355] d. In one example, the same size can be used for all sub-blocks.

[0356] i. In one example, the size of the sub-blocks can be W1xH1, where W1 is less than or equal to W and H1 is less than or equal to H, except that W1 = W and H1 = H.

[0357] 1) In one example, W1 can be equal to H1.

[0358] a) For example, W1 = H1 = 2.

[0359] b) For example, W1 = H1 = 4.

[0360] c) For example, W1 = H1 = 8.

[0361] d) For example, W1 = H1 = 16.

[0362] 2) In one example, a filtering method can be applied to the boundary of the sub-blocks, such as OBMC.

[0363] 3) In one example, a pixel-level refinement process can be applied when the number of samples of the sub-blocks is greater than 1.

[0364] ii. In one example, the size of the sub-blocks can depend on the coding information, such as color component and / or chroma format and / or video resolution.

[0365] iii. Alternatively, the size of at least one sub-block can be different from the other sub-blocks.

[0366] iv. In one example, one sub-block can be one pixel, i.e. W1=H1=1.

[0367] v. In one example, the sub-block partitioning can depend on the color format / color component.

[0368] 1) In one example, Wc = max(1, W1 / xS) and Hc = max(1, H1 / yS), where W1 and H1 are the width and height of the luma sub-block, Wc and Hc are the width and height of the chroma sub-block, and xS and yS are scaling factors depending on the color format.

[0369] a) For example, for YUV 4:2:0, xS = yS = 2.

[0370] b) For example, for YUV 4:4:4, xS = yS = 1.

[0371] c) For example, for YUV 4:2:2, xS = 2, yS = 1.

[0372] vi. In one example, the BVs of the chroma sub-blocks can be derived from at least one BV of at least one co-located luma sub-block.

[0373] 1) In one example, a co-located luma sub-block can refer to a luma sub-block composed of at least one co-located luma sample of a chroma sample in the chroma sub-block.

[0374] 2) In one example, a co-located luma sub-block can refer to a luma sub-block composed of a co-located luma sample of the center of the chroma sub-block, or a co-located luma sample of the top-left of the chroma sub-block, or a co-located luma sample of the top-right of the chroma sub-block, or a co-located luma sample of the bottom-left of the chroma sub-block, or a co-located luma sample of the bottom-right of the chroma sub-block.

[0375] 3) In one example, the BVs from the co-located luma sub-blocks can be weighted average.

[0376] 4) In one example, the BV of the co-located luma sub-block composed of the majority of the co-located luma samples can be used.

[0377] 5) In one example, the BV of the co-located luma sub-block at a specific location can be used, such as the center.

[0378] e. In one example, a T1 parameter affine model can be used.

[0379] i. For example, a 4 parameter affine model is used (i.e. T1 = 4).

[0380] 1) In one example, 2 CPBVs can be used in the 4 parameter affine model.

[0381] 2) In one example, the BVs for a sub-block can be derived using a 4-parameter affine model as follows: a) where (X0, y0) is the block vector of the top-left control point, (X1, y1) is the block vector of the top-right control point, (X2, y2) is the block vector of the bottom-left control point, and (X3, y3) is the block vector of the bottom-right control point. bv 0x , bv 0y ) is the block vector of the top-left control point, (X1, y1) is the block vector of the top-right control point, and (X2, y2) is the block vector of the bottom-left control point. bv 1x , bv 1y ) is the block vector of the top-right control point.

[0382] 3) Alternatively, the top-left control point and the bottom-left control point can be used.

[0383] 4) Alternatively, the top-left control point and the bottom-right control point can be used.

[0384] 5) Alternatively, the top-right control point and the bottom-left control point can be used.

[0385] 6) Alternatively, the top-right control point and the bottom-right control point can be used.

[0386] 7) Alternatively, the bottom-left control point and the bottom-right control point can be used.

[0387] ii. For example, a 6-parameter affine model is used (i.e., T1= 6).

[0388] 1) In one example, 3 CPBVs can be used in a 6-parameter affine model.

[0389] 2) In one example, the BVs for a sub-block can be derived using a 6-parameter affine model as follows: a) where (X0, y0) is the block vector of the top-left control point, (X1, y1) is the block vector of the top-right control point, (X2, y2) is the block vector of the bottom-left control point, and (X3, y3) is the block vector of the bottom-right control point. bv 0x , bv 0y ) is the block vector of the top-left control point, (X1, y1) is the block vector of the top-right control point, and (X2, y2) is the block vector of the bottom-left control point. bv 1x , bv 1y ) is the block vector of the top-right control point, and (X2, y2) is the block vector of the bottom-left control point. bv 2x , General aspects 2y ) is the block vector of the bottom-right control point.

[0390] 3) Alternatively, the top-left control point, the top-right control point, and the bottom-right control point can be used.

[0391] 4) Alternatively, the top-left control point, the bottom-right control point, and the bottom-left control point can be used.

[0392] 5) Alternatively, the top-right control point, the bottom-right control point and the bottom-left control point can be used.

[0393] iii. In one example, which affine model is used and / or how the affine model is used can be predefined or signaled or derived.

[0394] iv. In one example, the derived BV can be rounded to a precision.

[0395] 1) For example, the derived BV can be rounded to integer-pixel precision.

[0396] 2) For example, the derived BV can be rounded to half-pixel precision.

[0397] 3) For example, the derived BV can be rounded to quarter-pixel precision.

[0398] v. In one example, the derived BV can be clipped to a region.

[0399] 1) In one example, the region can only include samples that have already been reconstructed before the current block.

[0400] 2) In one example, the region can only include samples in a certain CTU or samples in a certain CTU row.

[0401] 3) In one example, the region can refer to IBC cache.

[0402] 2. It is proposed that only uni-prediction affine models are applied to IBC-affine coded blocks.

[0403] a. Alternatively, bi-prediction affine models can be applied to IBC-affine coded blocks.

[0404] i. In one example, each prediction direction can be associated with an affine model.

[0405] 3. It is proposed that one or more CPBV candidate lists can be constructed for a video unit coded with IBC-affine mode.

[0406] a. In one example, IBC-affine mode can be applied to IBC AMVP mode (IBC affine AMVP mode).

[0407] i. In one example, CPBV candidate list can refer to CPBV AMVP candidate list.

[0408] b. In one example, IBC-Affine mode can be applied to IBC Merge mode (IBC-Affine Merge mode).

[0409] i. In one example, CPBV candidate list can refer to CPBV Merge candidate list.

[0410] c. In one example, one or more types of CPBV candidates can be used to build the CPBV candidate list.

[0411] i. In one example, inherited CPBV candidates can be used.

[0412] 1) In one example, inherited CPBV candidates can be inferred from CPBVs of neighboring video units.

[0413] a) In one example, neighboring video units can refer to spatially neighboring and / or non-neighboring and / or temporally neighboring video units.

[0414] b) In one example, the inference method can be the same as used for affine inter prediction.

[0415] ii. In one example, built CPBV candidates can be used.

[0416] 1) In one example, built CPBV candidates can refer to CPBV candidates built by combining normal block vectors of neighboring video units for each control point.

[0417] a) In one example, neighboring video units can refer to spatially neighboring and / or non-neighboring and / or temporally neighboring video units.

[0418] b) In one example, BVs from the HBVP table can be used to build CPBV candidates.

[0419] c) In one example, temporal BVs can be used to build CPBV candidates.

[0420] iii. In one example, normal BVs from spatially neighboring and / or non-neighboring neighboring video units can be used to build the candidate list.

[0421] iv. In one example, temporal normal BVs can be used to build the candidate list.

[0422] v. In one example, normal BVs from the HBVP table can be used to build the candidate list.

[0423] vi. In one example, normal BVs from spatially neighboring and / or non-neighboring neighboring video units and / or temporal normal BVs and / or normal BVs from the HBVP table can be directly used as candidates in the CPBV candidate list.

[0424] vii. In one example, pair-wise candidates can be used.

[0425] viii. In one example, predefined candidates can be used, e.g., zero CPBVs.

[0426] ix. In one example, the way to construct the CPBV candidate list for IBC AMVP mode can be the same as for IBC Merge mode.

[0427] 1) Alternatively, the way to construct the CPBV candidate list for IBC AMVP mode can be different from IBC Merge mode.

[0428] d. In one example, de-duplication can be used during the construction of the CPBV candidate list.

[0429] i. In one example, de-duplication can be performed for some or all CPBV candidates.

[0430] ii. In one example, de-duplication can be performed when the motion information of two CPBV candidates are the same.

[0431] 1) In one example, the BVs of all control points are the same.

[0432] 2) In one example, the motion information can include other coding information, such as whether a specific coding tool is used.

[0433] a) In one example, the specific coding tool can refer to IBC-LIC, RR-IBC, filter-based IBC.

[0434] iii. In one example, de-duplication can be performed when the motion information of two CPBV candidates are similar.

[0435] 1) In one example, two CPBV candidates can be considered similar when the difference between the BVs of one or more control points of the two CPBV candidates is less than or equal to a threshold (T2).

[0436] e. In one example, reordering can be used during the construction of the CPBV candidate list.

[0437] i. In one example, reordering can depend on neighboring predicted / reconstructed samples.

[0438] ii. In one example, template matching cost can be used for reordering.

[0439] iii. In one example, one or more reordering operations can be used.

[0440] iv. In one example, reordering can be performed according to coding information, e.g. block size / dimension or candidate type.

[0441] f. In one example, a method of grouping CPBV candidates can be used.

[0442] i. In one example, a clustering method can be used.

[0443] ii. In one example, CPBV candidates can be divided into different groups / classes.

[0444] iii. In one example, a method of grouping CPBV candidates can be used together with a reordering method.

[0445] g. In one example, one or more CPBV candidates can be refined.

[0446] i. In one example, template matching can be used to refine CPBV candidates.

[0447] ii. In one example, some or all CPBV candidates can be refined.

[0448] h. In one example, one or more BV offsets / differences can be added to CPBV candidates.

[0449] i. In one example, BV offsets / differences can be added to all CPBV candidates.

[0450] ii. Alternatively, no BV offset / difference for at least one CPBV candidate.

[0451] iii. In one example, BV offsets / differences can be signaled or derived.

[0452] 1) In one example, the sign and / or amplitude of BV offsets / differences can be signaled or derived.

[0453] 4. In one example, CPBVs of a current video unit can be used for subsequent video units.

[0454] a. In one example, CPBVs of a current video unit can be stored in a HBVP table.

[0455] i. In one example, CPBVs can be stored in the same way as normal BVs.

[0456] ii. In one example, CPBVs can be stored in a separate HBVP table than the one used for normal BVs.

[0457] iii. In one example, CPBVs can be conditionally stored in the HBVP table.

[0458] 1) In one example, the condition can depend on the coding information, e.g., block dimension / size.

[0459] b. In one example, the parameters of the affine mode and / or position can be stored and used by subsequent video units.

[0460] i. In one example, the parameters of the IBC affine mode can be used for subsequent video units coded with IBC affine.

[0461] ii. In one example, the parameters of the IBC affine mode can be used for subsequent video units coded with inter affine.

[0462] iii. The parameters of the inter affine mode can be used for subsequent video units coded with IBC affine.

[0463] c. In one example, CPBVs can be used to build BV candidate lists for subsequent video units.

[0464] i. In one example, the subsequent video unit can refer to a video unit in the current picture / slice / tile or a video unit in a subsequent picture / slice / tile.

[0465] ii. In one example, CPBVs can only be used to build CPBV candidate lists for subsequent units.

[0466] iii. In one example, CPBVs from neighboring (adjacent or non-adjacent) video units can be used.

[0467] d. In one example, CPBVs of luma video units can be used for chroma video units.

[0468] i. In one example, CPBVs can be used to derive chroma BVs for chroma video units.

[0469] ii. In one example, CPBVs can be used to indicate the region used to train the model for CCCM.

[0470] e. In one example, CPBVs can be used for video units coded with IntraTMP.

[0471] f. In one example, CPBVs can be used only for neighboring video units.

[0472] g. Alternatively, CPBVs are not allowed for subsequent video units.

[0473] h. In one example, whether and / or how CPBVs are stored / used can be different for video units at CTU boundaries and video units not at CTU boundaries.

[0474] 5. In one example, IBC-Affine can be used with a particular coding tool.

[0475] a. In one example, the coding tool can refer to an intra coding tool.

[0476] i. In one example, the intra coding tool can refer to a particular coding tool such as regular intra prediction, DIMD, TIMD, MRL, ISP, MIP, IntraTMP, Intra prediction merge, SGPM, TMRL, PDPC / Gradient PDPC, CCLM, MMLM, CCCM, GLM, chroma merge, or variants thereof.

[0477] b. In one example, the coding tool can refer to an inter coding tool.

[0478] i. In one example, the inter coding tool can refer to a particular coding tool such as CIIP (e.g., CIIP-Planar, CIIP-TIMD, CIIP-TM), BCW (e.g., BCW index derived by TM), MMVD (e.g., MMVD or TM-based MMVD reordering), template matching (TM), affine (e.g., affine-MMVD, TM-based affine MMVD reordering), DMVR / multi-pass DMVR, PROF, BDOF or sample-based BDOF, adaptive decoder-side motion vector refinement (ADMVR), OBMC or TM-based OBMC, MHP, GPM (e.g., GPM, GPM-TM, GPM-MMVD, GPM-Intra), bilateral / template matching AMVP-Merge mode, or variants thereof.

[0479] c. In one example, the coding tool can refer to BDPCM or palette.

[0480] d. In one example, the coding tool can refer to an IBC coding tool.

[0481] i. In one example, the IBC coding tool refers to IBC AMVP mode, which can refer to normal IBC AMVP, or TM-based IBC AMVP, or RR-IBC AMVP mode, or other IBC AMVP mode in which BV predictor is derived and BVD is signaled / derived, or IBC Merge mode which can refer to normal IBC Merge mode, or IBC-TM Merge mode, or IBC-MBVD mode, or IBC-CIIP, or IBC-GPM, or IBC-LIC, or sub-pixel based IBC, or filter based IBC, or variants thereof.

[0482] e. In one example, the coding tool can refer to loop filter method.

[0483] i. In one example, the boundary between different sub-blocks can be handled by deblocking filter.

[0484] f. In one example, the coding tool can refer to transform method, such as DCT / DST / MTS / LFNST / NSPT.

[0485] g. In one example, the coding tool can refer to sign prediction and / or dependent quantization.

[0486] h. In one example, when IBC-afine is used with TM-based IBC AMVP, one or more CPBVs of bi-prediction can be refined using TM.

[0487] i. In one example, the fractional precision of BVs can be derived using TM.

[0488] i. In one example, one or more CPBVs can be refined using template matching and / or bilateral matching.

[0489] j. In one example, when IBC-afine is used with IBC-TM Merge mode, one or more BV offsets can be added to CPBVs.

[0490] k. In one example, when IBC-afine is used with RR-IBC AMVP mode, all CPBVs are flipped.

[0491] i. Alternatively, one of CPBVs can not be flipped.

[0492] l. In one example, when IBC-afine is used with IBC-LIC / filter-based IBC, some or all sub-blocks can be refined using IBC-LIC / filter-based IBC parameters.

[0493] m. In one example, when IBC-affine is used with IBC-CIIP, the prediction signal generated by IBC can be obtained through CPBVs.

[0494] n. In one example, when IBC-affine is used with IBC-GPM, one or more sub-partitioned prediction signals can be obtained through IBC-affine.

[0495] o. In one example, when IBC-affine is used with IBC-MBVD, BV offsets can be derived for one or more CPBVs.

[0496] p. In one example, IBC-affine can be used with sub-pixel / fraction based IBC.

[0497] i. In one example, one or more CPBVs can be sub-pixel / fraction precision.

[0498] ii. In one example, all CPBVs can be sub-pixel / fraction.

[0499] iii. In one example, at most one CPBV is allowed to be sub-pixel / fraction.

[0500] iv. In one example, the derived BVs for each sub-block can be fraction precision or integer.

[0501] v. In one example, sub-pixel / fraction precision can refer to 1 / 2, 1 / 4, 1 / 8, 1 / 16, 1 / 32, 1 / 64.

[0502] vi. In one example, the interpolation for fraction BVs in IBC-affine can be the same as non-IBC-affine.

[0503] 1) Alternatively, the interpolation for fraction BVs in IBC-affine can be different from non-IBC-affine.

[0504] q. Alternatively, IBC-affine is not allowed to be used with specific coding tools.

[0505] i. In one example, the coding tool can refer to IBC AMVP mode, which can refer to normal IBC AMVP, or TM-based IBC AMVP, or RR-IBC AMVP mode, or other IBC AMVP mode in which BV predictor is derived and BVD is signaled / derived, or IBC Merge mode which can refer to normal IBC Merge mode, or IBC-TM Merge mode, or IBC-MBVD mode, or IBC-CIIP, or IBC-GPM, or IBC-LIC, or sub-pixel based IBC, or filter based IBC, or variants thereof.

[0506] 6. Whether and / or how to apply IBC-Affine can depend on coding information, which can refer to: a. whether a particular coding method is allowed, b. block dimension and / or block size, c. block depth, d. slice / picture type and / or partition tree type (single tree, or dual tree, or local dual tree), e. temporal layer identification, f. block position, g. video content (e.g., camera captured content and / or screen content and / or mixed content), h. color format, i. color component.

[0507] i. In one example, IBC-Affine can be applied to all color components.

[0508] ii. In one example, when IBC-Affine is applied to chroma components, it can be different from that for luma component.

[0509] iii. In one example, whether and / or how to apply IBC-Affine to a first component can depend on whether IBC-Affine is applied to a second component.

[0510] 1) In one example, the first component can refer to chroma component (e.g., Cb and / or Cr), and the second component can refer to luma component (e.g., Y).

[0511] 2) In one example, the way IBC-Affine is applied to the first component can be the same as that for the second component.

[0512] a) Alternatively, the way IBC-Affine is applied to the first component can be different from that for the second component.

[0513] iv. In one example, IBC-Affine can be applied to the luma component, but not to the chroma components.

[0514] 1) In one example, the luma component can refer to Y in YCbCr color space or G in RGB color space.

[0515] 2) In one example, the chroma component can refer to Cb and / or Cr in YCbCr color space or R and / or B in RGB color space.

[0516] 7. The indication of IBC-Affine can be conditionally signaled, where the condition can include: a. block dimension and / or block size, b. block depth, c. slice / picture type and / or partition tree type (single tree, or dual tree, or local dual tree), d. temporal layer identification, e. block position, f. color format, g. color component.

[0517] 8. Whether the current block is coded in IBC-Affine mode can be signaled using one or more syntax elements.

[0518] a. In one example, the syntax elements can be binarized or coded as flags with fixed length coding or truncated unary coding or unary coding or EG coding.

[0519] b. In one example, the syntax elements can be bypass coded or context coded.

[0520] i. The context can depend on coded information, such as block dimension and / or block size and / or slice / picture type and / or information of neighboring blocks (adjacent or non-adjacent) and / or information of other coding tools used for the current block and / or information of temporal layer.

[0521] c. In one example, the one or more syntax elements can be signaled at sequence header / picture header / S PS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.

[0522] d. In one example, the syntax elements can be coded in a predictive manner.

[0523] e. For example, the syntax elements of the current block can be predicted by the syntax elements of neighboring blocks.

[0524] 9. Whether the disclosed method is used can depend on content characteristics, e.g. screen content or natural content.

[0525] 10. The disclosed method can be applied to other coding tools, e.g. Intra TMP.

[0526] Figure 54 11. In the above examples, video unit can refer to a video unit, which can refer to a color component / sub-picture / slice / tile / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transform unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transform block (TB) / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.

[0527] 12. Whether and / or how the disclosed method is applied can be signaled at sequence level / group of pictures level / picture level / slice level / tile group level, e.g. in sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.

[0528] 13. Whether and / or how the disclosed method is applied can depend on the following information: a. a message signaled in DPS / SPS / VPS / PPS / APS / picture header / slice header / tile group header / coding tree unit (CTU) / coding unit (CU) / CTU row / CTU group / TU / PU block / video coding unit; b. the location of the CU / PU / TU / block / video coding unit; c. the block dimension of the current block and / or the block dimension of a neighboring block of the current block; d. the block shape of the current block and / or the block shape of a neighboring block of the current block; e. the coding mode of the block, e.g. IBC or non-IBC inter mode or non-IBC sub-block mode; f. an indication of the color format (such as 4:2:0, 4:4:4); g. the coding tree structure; h. the slice / tile group type and / or the picture type; i. the color component (e.g. can be applied only to chroma component or luma component); j. the temporal layer ID; k. the profile / tier / layer of the standard.

[0529] 14. The syntax element disclosed above can be binarized as a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, a truncated binary code, etc. It can be signed or unsigned.

[0530] 15. The syntax element disclosed above can be coded / decoded with at least one context model. Or it can be bypass coded.

[0531] 16. The syntax element disclosed above can be signaled in a conditional way.

[0532] a. The SE is signaled only if the corresponding function is applicable.

[0533] b. The SE is signaled only if the dimension (width and / or height) of the block satisfies a condition.

[0534] 17. The syntax element disclosed above can be signaled at block level / sequence level / group of pictures level / picture level / slice level / tile group level, e.g., in the coding structure of CTU / CU / TU / PU / CTB / CB / TB / PB, or in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.

[0535] 18. The proposed method(s) can be combined with another coding tool, e.g., affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.

[0536] 19. The proposed method(s) can be exclusive with another coding tool, e.g., affine / MTS / LFNST / MMVD / MIP / ISP / CCLM / CCCM / SMVD / BDOF / DMVR / HMVP / template matching / IBC / palette / etc.

[0537] a. In one example, if the proposed method(s) is used, the exclusive coding tool is implicitly disabled without signaling.

[0538] b. In one example, if the exclusive coding tool is used, the proposed method(s) is implicitly disabled without signaling.

[0539] As used herein, the term “video unit” or “video block” can be a sequence, a picture, a slice, a tile, a brick, a subpicture, a coding tree unit (CTU) / coding tree block (CTB), a CTU / CTB row, one or more coding units (CU) / coding blocks (CB), one or more CTU / CTB, one or more virtual pipeline data units (VPDU), a subregion within a picture / slice / tile / brick.

[0540] Figure 55 A flowchart of a method 5400 for video processing according to an embodiment of the disclosure is shown. The method 5400 is implemented during conversion between a video unit of a video and a bitstream of the video.

[0541] At block 5410, for conversion between a video unit of a video and a bitstream of the video, a block vector (BV) for at least one subblock of a plurality of subblocks of the video unit is derived using an affine model. In this case, in some embodiments, a same block size is used for all subblocks in the plurality of subblocks. Alternatively or additionally, at least one of a parameter of the affine model or a location of the video unit is applied to a subsequent video unit of the video unit.

[0542] At block 5420, the conversion is performed based on the BV. In some embodiments, the conversion can include encoding the video unit into the bitstream. Alternatively, the conversion can include decoding the video unit from the bitstream.

[0543] The method 5400 enables the affine model to be used to improve coding performance to handle textures with zooming, rotation, perspective motion, and other irregular motion.

[0544] In some embodiments, a filtering method can be applied to boundaries of the plurality of subblocks. For example, the filtering method can be an overlapped block motion compensation (OBMC). In some other embodiments, if a number of samples of a subblock is greater than a threshold, a pixel-level refinement process can be applied to the subblock. In this case, the threshold is equal to 1.

[0545] In some embodiments, the block size of a sub-block of the plurality of sub-blocks can be one pixel. In this case, the block size of a sub-block of the plurality of sub-blocks is W x H, and W = H = 1, W represents the width of the sub-block, and H represents the height of the sub-block. In some other embodiments, the sub-block partitioning of the plurality of sub-blocks can depend on at least one of: the color format or the color component. In some examples, Wc can be equal to 1 and the maximum of Wl / xS, and Hc can be equal to 1 and the maximum of Hl / yS. In this case, Wl represents the width of a luma sub-block, Hl represents the height of a luma sub-block, Wc represents the width of a chroma sub-block, Hc represents the height of a chroma sub-block, and xS and yS represent scaling factors depending on the color format. For example, for the color format YUV 4:2:0, xS = yS = 2. In some examples, for the color format YUV 4:4:4, xS = yS = 1. In some examples, for the color format YUV 4:2:2, xS = 2 and yS = 1.

[0546] In some embodiments, the BV of a chroma sub-block can be derived from at least one BV of at least one collocated luma sub-block. In some examples, the at least one collocated luma sub-block can comprise a luma sub-block. In this case, the luma sub-block can comprise at least one collocated luma sample of a chroma sample in the chroma sub-block. In some embodiments, the at least one collocated luma sub-block can comprise a luma sub-block. In this case, the luma sub-block can comprise at least one of: a collocated luma sample of the center of the chroma sub-block, a collocated luma sample of the top-left of the chroma sub-block, a collocated luma sample of the top-right of the chroma sub-block, a collocated luma sample of the bottom-left of the chroma sub-block, or a collocated luma sample of the bottom-right of the chroma sub-block. In some other embodiments, the at least one BV of the at least one collocated luma sub-block can be weighted averaged. In some embodiments, the at least one BV of the at least one collocated luma sub-block comprising the majority of collocated luma samples can be used. Alternatively, the at least one BV of the at least one collocated luma sub-block at a position can be used. For example, the position can be the center.

[0547] In some embodiments, the derived BVs can be rounded to a precision. For example, the precision can be at least one of: integer pixel precision, half-pixel precision, or quarter-pixel precision. In some other embodiments, the derived BVs can be clipped to a region. For example, the region can comprise samples that have been reconstructed before the current block. In some examples, the region can comprise at least one of: samples in a coding tree unit (CTU) or samples in a CTU row. For example, the region can comprise an IBC buffer.

[0548] In some embodiments, the parameters of the IBC affine mode can be applied by a subsequent video unit coded with IBC affine. In some other embodiments, the parameters of the IBC affine mode can be applied by a subsequent video unit coded with inter affine. Alternatively, the parameters of the inter affine mode can be applied by a subsequent video unit coded with IBC affine.

[0549] In some embodiments, the IBC affine can be used with a coding tool, and the coding tool can include an in-loop filter method. In some examples, the boundary between different sub-blocks can be processed by a deblocking filter. In some other embodiments, the IBC affine can be used with a coding tool, and the coding tool can include a transform method. For example, the transform method can include at least one of a discrete cosine transform (DCT), a discrete sine transform (DST), a multi-transform selection (MTS), a low frequency non-separable transform (LFNST), or a non-separable primary transform (NSPT).

[0550] In some embodiments, the IBC affine can be used with a coding tool, and the coding tool can include at least one of sign prediction or dependent quantization. In some other embodiments, one or more control point block vectors (CPBVs) can be refined by at least one of template matching or bilateral matching.

[0551] In some embodiments, the indication of whether and / or how to apply the IBC affine can depend on coding information. In this case, the coding information can include video content. For example, the video content can include at least one of camera captured content, screen content, or mixed content.

[0552] In some embodiments, the video unit can include at least one of a color component, a sub-picture, a slice, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), a block, a sub-block of a block, a sub-region within a block, or a region containing more than one sample or pixel.

[0553] In some embodiments, the indication of whether and / or how to use the affine model to derive the BVs of at least one sub-block can be indicated at one of a sequence level, a group of pictures level, a picture level, a slice level, or a tile group level.

[0554] In some embodiments, the indication of whether and / or how to use the affine model to derive the BV of the at least one sub-block can be indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.

[0555] In some embodiments, whether and / or how to use the affine model to derive the BV of the at least one sub-block can depend on at least one of the following: a message indicated in one of the following: a DPS, a SPS, a VPS, a PPS, an APS, a picture header, a slice header, a tile group header, a coding tree unit (CTU), a coding unit (CU), a CTU row, a CTU group, a TU, a PU block, a video coding unit, a location of a CU block, a location of a PU block, a location of a TU block, a location of a video coding unit, a block dimension of a current block, a block dimension of a neighboring block of the current block, a coding mode of a block, an indication of a color format, a coding tree structure, a slice group type, a tile group type, a slice picture type, a tile picture type, a color component, an ID of a temporal layer, a profile of a standard, a level of a standard, or a tier of a standard. In some embodiments, the coding mode of a block can be at least one of the following: an IBC inter mode, a non-IBC inter mode, or a non-IBC sub-block mode. In some embodiments, the indication of a color format can be 4:2:0 or 4:4:4. In some other embodiments, the color component can be applied to a chroma component or a luma component.

[0556] In some embodiments, a syntax element can be binarized as a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code. In such a case, the syntax element can be signed or unsigned. In some other embodiments, the syntax element can be coded with at least one context model, or bypass coded. In some embodiments, the syntax element can be signaled if a corresponding function is applicable, or if a dimension of a block satisfies a condition. In some embodiments, the dimension can be a width and / or a height.

[0557] In some embodiments, a syntax element can be signaled at one of the following: a block level, a sequence level, a group of pictures level, a picture level, a slice level, or a tile group level.

[0558] In some embodiments, the syntax element can be signaled in one of: a coding structure of a CTU, a coding structure of a CU, a coding structure of a TU, a coding structure of a PU, a coding structure of a CTB, a coding structure of a CB, a coding structure of a TB, a coding structure of a PB, a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.

[0559] In some embodiments, IBC with affine model can be combined with at least one of the following coding tools: affine, multi-transform selection (MTS), low-frequency non-separable transform (LFNST), merge mode with motion vector difference (MMVD), MIP, ISP, cross-component linear model (CCLM), convolution cross-component model (CCCM), symmetric motion vector difference (SMVD), bi-directional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based MVP (HMVP), template matching, intra block copy (IBC), or palette. In some other embodiments, IBC with affine model can be exclusive with at least one of the following coding tools: affine, multi-transform selection (MTS), LFNST, MMVD, MIP, ISP, CCLM, CCCM, SMVD, BDOF, DMVR, HMVP, template matching, IBC, or palette. In some embodiments, if IBC with affine model is used, the exclusive coding tools can be disabled without signaling. Alternatively, if the exclusive coding tools are used, IBC with affine model can be disabled without signaling.

[0560] According to a further embodiment of the disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing. The method includes deriving a block vector (BV) for at least one subblock of a plurality of subblocks of a video unit of the video using an affine model, wherein a same block size is used for all of the plurality of subblocks, and / or at least one of parameters of the affine model or a position of the video unit is applied to a subsequent video unit of the video unit, and generating the bitstream based on the BV.

[0561] According to yet some embodiments of the disclosure, a method for storing a bitstream of a video is provided. The method includes deriving, using an affine model, a block vector (BV) for at least one sub-block of a plurality of sub-blocks of a video unit of the video, wherein a same block size is used for all of the plurality of sub-blocks, and / or at least one of parameters of the affine mode or a position of the video unit is applied to a subsequent video unit of the video unit; generating the bitstream based on the BV; and storing the bitstream in a non-transitory computer-readable recording medium.

[0562] Implementations of the disclosure can be described according to the following clauses, which features can be combined in any reasonable manner.

[0563] Clause 1. A method for video processing, comprising: for conversion between a video unit of a video and a bitstream of the video, deriving, using an affine model, a block vector (BV) for at least one sub-block of a plurality of sub-blocks of the video unit, wherein a same block size is used for all of the plurality of sub-blocks, and / or at least one of parameters of the affine mode or a position of the video unit is applied to a subsequent video unit of the video unit; and performing the conversion based on the BV.

[0564] Clause 2. The method of claim 1, wherein a filtering mode is applied to a boundary of the plurality of sub-blocks.

[0565] Clause 3. The method of claim 2, wherein the filtering mode is an overlapped block motion compensation (OBMC).

[0566] Clause 4. The method of claim 1, wherein if a number of samples of a sub-block is greater than a threshold, a pixel-level refinement process is applied to the sub-block, wherein the threshold is equal to 1.

[0567] Clause 5. The method of claim 1, wherein a block size of one of the plurality of sub-blocks is one pixel, wherein the block size of one of the plurality of sub-blocks is WH, and W=H=1, W representing a width of the sub-block and H representing a height of the sub-block.

[0568] Clause 6. The method of claim 1, wherein a sub-block partitioning of the plurality of sub-blocks depends on at least one of: a color format or a color component.

[0569] Item 7. The method of claim 6, wherein Wc is equal to 1 and the maximum of Wl / xS and Hc is equal to 1 and the maximum of Hl / yS, wherein Wl denotes a width of a luma subblock, Hl denotes a height of the luma subblock, Wc denotes a width of a chroma subblock, Hc denotes a height of the chroma subblock, and xS and yS denote scaling factors dependent on the color format.

[0570] Item 8. The method of claim 7, wherein for YUV 4:2:0 of the color format, xS = yS = 2.

[0571] Item 9. The method of claim 7, wherein for YUV 4:4:4 of the color format, xS = yS = 1.

[0572] Item 10. The method of claim 7, wherein for YUV 4:2:2 of the color format, xS = 2 and yS = 1.

[0573] Item 11. The method of claim 1, wherein a BV of a chroma subblock is derived from at least one BV of at least one collocated luma subblock.

[0574] Item 12. The method of claim 11, wherein the at least one collocated luma subblock comprises a luma subblock, wherein the luma subblock comprises at least one collocated luma sample of a chroma sample in the chroma subblock.

[0575] Item 13. The method of claim 11, wherein the at least one collocated luma subblock comprises a luma subblock, wherein the luma subblock comprises at least one of: a collocated luma sample of a center of the chroma subblock, a collocated luma sample of a top-left of the chroma subblock, a collocated luma sample of a top-right of the chroma subblock, a collocated luma sample of a bottom-left of the chroma subblock, or a collocated luma sample of a bottom-right of the chroma subblock.

[0576] Item 14. The method of claim 11, wherein the at least one BV of the at least one collocated luma subblock is weighted averaged.

[0577] Item 15. The method of claim 11, wherein the at least one BV of the at least one collocated luma subblock comprising a majority of collocated luma samples is used.

[0578] Item 16. The method of claim 11, wherein the at least one BV of the at least one collocated luma subblock at a location is used.

[0579] Item 17. The method of claim 11, wherein the location is a center.

[0580] Clause 18. The method of claim 1, wherein the derived BV is rounded to a precision.

[0581] Clause 19. The method of claim 18, wherein the precision is at least one of: integer pixel precision, half-pixel precision, or quarter-pixel precision.

[0582] Clause 20. The method of claim 1, wherein the derived BV is clipped to a region.

[0583] Clause 21. The method of claim 20, wherein the region comprises samples that have been previously reconstructed before a current block.

[0584] Clause 22. The method of claim 20, wherein the region comprises at least one of: samples in a coding tree unit (CTU) or samples in a CTU row.

[0585] Clause 23. The method of claim 20, wherein the region comprises an IBC cache.

[0586] Clause 24. The method of claim 1, wherein the parameters of IBC affine mode are applied by the subsequent video unit coded with IBC affine.

[0587] Clause 25. The method of claim 1, wherein the parameters of IBC affine mode are applied by the subsequent video unit coded with inter affine.

[0588] Clause 26. The method of claim 1, wherein parameters of inter affine mode are applied by the subsequent video unit coded with IBC affine.

[0589] Clause 27. The method of claim 1, wherein IBC affine is used with a coding tool, and the coding tool comprises an in-loop filter mode.

[0590] Clause 28. The method of claim 27, wherein a boundary between different sub-blocks is processed by a deblocking filter.

[0591] Clause 29. The method of claim 1, wherein IBC affine is used with a coding tool, and the coding tool comprises a transform mode.

[0592] Item 30. The method of claim 29, wherein the transform mode comprises at least one of: a discrete cosine transform (DCT), a discrete sine transform (DST), a multiple transform selection (MTS), a low-frequency non-separable transform (LFNST), or a non-separable primary transform (NSPT).

[0593] Item 31. The method of claim 1, wherein IBC affine is used with a coding tool, and the coding tool comprises at least one of: sign prediction or dependent quantization.

[0594] Item 32. The method of claim 1, wherein one or more control point block vectors (CPBVs) are refined by at least one of: template matching or bilateral matching.

[0595] Item 33. The method of claim 1, wherein an indication of whether and / or how to apply IBC affine depends on coding information, wherein the coding information comprises video content.

[0596] Item 34. The method of any of claims 1-33, wherein the video unit comprises at least one of: a color component, a subpicture, a slice, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), a block, a subblock of a block, a subregion within a block, a region containing more than one sample or pixel.

[0597] Item 35. The method of any of claims 1-33, wherein an indication of whether and / or how to use the affine model to derive the BVs of the at least one subblock is indicated at one of: a sequence level, a picture group level, a picture level, a slice level, or a tile group level.

[0598] Item 36. The method of any of claims 1-33, wherein an indication of whether and / or how to use the affine model to derive the BVs of the at least one subblock is indicated in one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.

[0599] Item 37. The method of any of claims 1-33, wherein whether and / or how the affine model is used to derive the BV of the at least one sub-block depends on at least one of a message indicated in one of: a DPS, a SPS, a VPS, a PPS, an APS, a picture header, a slice header, a tile group header, a coding tree unit (CTU), a coding unit (CU), a CTU row, a CTU group, a TU, a PU block, a video coding unit, a location of a CU block, a location of a PU block, a location of a TU block, a location of a video coding unit, a block dimension of a current block, a block dimension of a neighboring block of the current block, a coding mode of a block, an indication of a color format, a coding tree structure, a slice group type, a tile group type, a slice picture type, a tile picture type, a color component, an ID of a temporal layer, a profile of a standard, a level of a standard, or a tier of a standard.

[0600] Item 38. The method of claim 37, wherein the coding mode of the block is at least one of: an IBC inter mode, a non-IBC inter mode, or a non-IBC sub-block mode.

[0601] Item 39. The method of claim 37, wherein the indication of a color format is 4:2:0 or 4:4:4.

[0602] Item 40. The method of claim 37, wherein the color component is applied to a chroma component or a luma component.

[0603] Item 41. The method of any of claims 1-33, wherein the syntax element is binarized as a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code.

[0604] Item 42. The method of any of claims 1-33, wherein the syntax element is signed or unsigned.

[0605] Item 43. The method of any of claims 1-33, wherein the syntax element is coded with at least one context model, or bypass coded.

[0606] Item 44. The method of any of claims 1-33, wherein the syntax element is signaled under at least one condition: if a corresponding function is applicable, or if a dimension of the block satisfies a condition.

[0607] Item 45. The method of claim 44, wherein the dimension is a width and / or a height.

[0608] Item 46. The method of any of claims 1-33, wherein the syntax element is signaled at one of: a block level, a sequence level, a group of pictures level, a picture level, a slice level, or a tile group level.

[0609] Item 47. The method of any of claims 1-33, wherein the syntax element is signaled in one of: a coding structure of a CTU, a coding structure of a CU, a coding structure of a TU, a coding structure of a PU, a coding structure of a CTB, a coding structure of a CB, a coding structure of a TB, a coding structure of a PB, a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.

[0610] Item 48. The method of any of claims 1-33, wherein IBC with affine model is combined with at least one of the following coding tools: affine, multi-transform selection (MTS), low-frequency non-separable transform (LFNST), merge mode with motion vector difference (MMVD), MIP, ISP, cross-component linear model (CCLM), convolution cross-component model (CCCM), symmetric motion vector difference (SMVD), bi-directional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based MVP (HMVP), template matching, intra block copy (IBC), or palette.

[0611] Item 49. The method of any of claims 1-33, wherein IBC with affine model is exclusive with at least one of the following coding tools: affine, multi-transform selection (MTS), LFNST, merge mode with motion vector difference (MMVD), MIP, ISP, cross-component linear model (CCLM), convolution cross-component model (CCCM), SMVD, bi-directional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based MVP (HMVP), template matching, intra block copy (IBC), or palette.

[0612] Item 50. The method of claim 49, wherein if the IBC with affine model is used, the exclusive coding tool is disabled without signaling.

[0613] Item 51. The method of claim 49, wherein if the exclusive coding tool is used, the IBC with affine model is disabled without signaling.

[0614] Clause 52. The method of any of clauses 1-51, wherein the converting comprises encoding the video unit into the bitstream.

[0615] Clause 53. The method of any of clauses 1-51, wherein the converting comprises decoding the video unit from the bitstream.

[0616] Clause 54. An apparatus for video processing comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of clauses 1-53.

[0617] Clause 55. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any of clauses 1-53.

[0618] Clause 56. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing, wherein the method comprises deriving a block vector (BV) for at least one subblock of a plurality of subblocks of a video unit of the video using an affine model, wherein a same block size is used for all subblocks of the plurality of subblocks, and / or at least one of parameters of an affine mode or a position of the video unit is applied to a subsequent video unit of the video unit, and generating the bitstream based on the BV.

[0619] Clause 57. A method for storing a bitstream of a video, comprising deriving a block vector (BV) for at least one subblock of a plurality of subblocks of a video unit of the video using an affine model, wherein a same block size is used for all subblocks of the plurality of subblocks, and / or at least one of parameters of an affine mode or a position of the video unit is applied to a subsequent video unit of the video unit, generating the bitstream based on the BV, and storing the bitstream in a non-transitory computer-readable recording medium.

[0620] Example Device Figure 55 A block diagram of a computing device 5500 in which various embodiments of the present disclosure can be implemented is shown. The computing device 5500 can be implemented as the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300), or can be included in the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).

[0621] It should be understood that, Figure 55The computing device 5500 shown is for illustrative purposes only and is not intended to imply any limitation on the functionality and scope of the embodiments of this disclosure.

[0622] like Figure 55 As shown, computing device 5500 includes general-purpose computing device 5500. Computing device 5500 may include at least one or more processors or processing units 5510, memory 5520, storage unit 5530, one or more communication units 5540, one or more input devices 5550, and one or more output devices 5560.

[0623] In some embodiments, the computing device 5500 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, large computing device, etc., provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 5500 can support any type of interface to the user (such as "wearable" circuitry devices, etc.).

[0624] Processing unit 5510 can be a physical processor or a virtual processor, and can perform various processes based on programs stored in memory 5520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capabilities of computing device 5500. Processing unit 5510 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.

[0625] Computing device 5500 typically includes a variety of computer storage media. Such media can be any media that is accessible by computing device 5500 and can include, without limitation, both volatile and non-volatile media, or removable and non-removable media. Memory 5520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (such as read only memory (ROM), Electrically Programmable Read Only Memory (EPROM), or Electrically Erasable Programmable Read Only Memory (EEPROM)), or any combination thereof. Storage 5530 can be any available storage media that is detachable or non- detachable and that can be used to store information and / or data and that can be accessed by computing device 5500.

[0626] Computing device 5500 can also include additional removable / non-removable, volatile / non-volatile storage media. For example, computing device 5500 can include a mass disk that is removably coupled to computing device 5500 via an interface (e.g., an EIDE, SATA, or IEEE 1394 interface). Like storage 5530, mass disk can comprise a computer storage medium. Examples of mass disk include categories of computer storage media such as one or more of hard disk drives, optical disk drives, and tape drives. ​ Although not shown in FIG. 5, a mass disk drive and / or an optical disk drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile mass disk. In such instances, each can be connected to the bus (not shown) by one or more data media interfaces.

[0627] Communication unit 5540 enables communications with other computing devices via a communication medium. Additionally, the functionality of the components of computing device 5500 can be implemented by a single computing cluster or a plurality of computing machines that can communicate over a communication connection. Thus, computing device 5500 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general network nodes.

[0628] Input device 5550 can be one or more of various input devices such as a mouse, a keyboard, a trackball, a voice input device, and / or the like. Output device 5560 can be one or more of various output devices such as a display, a speaker, a printer, and / or the like. Computing device 5500 can also include communication unit 5540, which can enable computing device 5500 to communicate with one or more external devices such as a storage device or a display device, which can enable a user to interact with computing device 5500, or any devices (e.g., a network card, a modem, etc.) that can enable computing device 5500 to communicate with one or more other computing devices. Such communication can be enabled via an input / output (I / O) interface (not shown).

[0629] In some embodiments, some or all of the components of computing device 5500 can also be arranged in a cloud computing architecture, rather than being integrated in a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computation, software, data access, and storage services that do not require end-user knowledge of the physical location or configuration of the system that delivers the services. In various embodiments, cloud computing delivers services via the internet using appropriate protocols. For example, a cloud computing provider provides an application via the internet that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture, and corresponding data, can be stored on servers at remote locations. Computing resources in a cloud computing environment can be consolidated or distributed at locations remote to the user. Cloud computing infrastructure can provide services through a shared data center, although they appear as a single point of access to the user. Thus, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server, or installed directly on a client device, either directly or in other ways.

[0630] In embodiments of the disclosure, computing device 5500 can be used to implement video encoding / decoding. Memory 5520 can include one or more video codec modules 5525 having one or more program instructions. These modules are accessible and executable by processing unit 5510 to perform the functions of the various embodiments described herein.

[0631] In example embodiments that perform video encoding, input device 5550 can receive video data as input 5570 to be encoded. The video data can be processed, for example, by video codec module 5525, to generate an encoded bitstream. The encoded bitstream can be provided as output 5580 via output device 5560.

[0632] In example embodiments that perform video decoding, input device 5550 can receive an encoded bitstream as input 5570. The encoded bitstream can be processed, for example, by video codec module 5525, to generate decoded video data. The decoded video data can be provided as output 5580 via output device 5560.

[0633] While the present disclosure has been particularly shown and described with reference to the preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application as defined by the appended claims. Such variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of the application is not intended to be limiting.

Claims

1. A method for video processing, comprising: deriving, for a conversion between a video unit of a video and a bitstream of the video, a block vector (BV) for at least one sub-block of a plurality of sub-blocks of the video unit using an affine model, wherein a same block size is used for all sub-blocks of the plurality of sub-blocks, and / or at least one of parameters of the affine mode or a position of the video unit is applied to a subsequent video unit of the video unit; and performing the conversion based on the BV.

2. The method of claim 1, wherein a filtering manner is applied to a boundary of the plurality of sub-blocks.

3. The method of claim 2, wherein the filtering manner is an overlapped block motion compensation (OBMC).

4. The method of claim 1, wherein a pixel-level refinement process is applied to a sub-block if a number of samples of the sub-block is greater than a threshold, wherein the threshold is equal to 1.

5. The method of claim 1, wherein a block size of a sub-block of the plurality of sub-blocks is one pixel, wherein the block size of a sub-block of the plurality of sub-blocks is WH, and W = H = 1, W representing a width of the sub-block and H representing a height of the sub-block.

6. The method of claim 1, wherein a sub-block partitioning of the plurality of sub-blocks depends on at least one of a color format or a color component.

7. The method of claim 6, wherein Wc is equal to 1 and a maximum of Wl / xS, and Hc is equal to 1 and a maximum of Hl / yS, wherein Wl represents a width of a luma sub-block, Hl represents a height of the luma sub-block, Wc represents a width of a chroma sub-block, Hc represents a height of the chroma sub-block, and xS and yS represent scaling factors depending on the color format.

8. The method of claim 7, wherein for YUV 4:2:0 of the color format, xS = yS = 2.

9. The method of claim 7, wherein for YUV 4:4:4 of the color format, xS = yS = 1.

10. The method of claim 7, wherein for YUV 4:2:2 of the color format, xS = 2 and yS = 1.

11. The method of claim 1, wherein a BV of a chroma sub-block is derived from at least one BV of at least one co-located luma sub-block.

12. The method of claim 11, wherein the at least one co-located luma sub-block comprises a luma sub-block, wherein the luma sub-block comprises at least one co-located luma sample of a chroma sample in the chroma sub-block.

13. The method of claim 11, wherein the at least one co-located luma sub-block comprises a luma sub-block, wherein the luma sub-block comprises at least one of a co-located luma sample of a center of the chroma sub-block, a co-located luma sample of a top-left of the chroma sub-block, a co-located luma sample of a top-right of the chroma sub-block, a co-located luma sample of a bottom-left of the chroma sub-block, or a co-located luma sample of a bottom-right of the chroma sub-block.

14. The method of claim 11, wherein the at least one BV of the at least one collocated luma subblock is weighted averaged.

15. The method of claim 11, wherein the at least one BV of the at least one collocated luma subblock that includes a majority of collocated luma samples is used.

16. The method of claim 11, wherein the at least one BV of the at least one collocated luma subblock at a location is used.

17. The method of claim 11, wherein the location is a center.

18. The method of claim 1, wherein the derived BV is rounded to a precision.

19. The method of claim 18, wherein the precision is at least one of: an integer pixel precision, a half-pixel precision, or a quarter-pixel precision.

20. The method of claim 1, wherein the derived BV is clipped to a region.

21. The method of claim 20, wherein the region includes samples that have been previously reconstructed before a current block.

22. The method of claim 20, wherein the region includes at least one of: samples in a coding tree unit (CTU) or samples in a CTU row.

23. The method of claim 20, wherein the region includes an IBC cache.

24. The method of claim 1, wherein the parameters of IBC affine mode are applied by the subsequent video unit coded with IBC affine.

25. The method of claim 1, wherein the parameters of IBC affine mode are applied by the subsequent video unit coded with inter affine.

26. The method of claim 1, wherein parameters of inter affine mode are applied by the subsequent video unit coded with IBC affine.

27. The method of claim 1, wherein IBC affine is used with a coding tool, and the coding tool includes an in-loop filter mode.

28. The method of claim 27, wherein a boundary between different subblocks is processed by a deblocking filter.

29. The method of claim 1, wherein IBC affine is used with a coding tool, and the coding tool includes a transform mode.

30. The method of claim 29, wherein the transform mode includes at least one of: a discrete cosine transform (DCT), a discrete sine transform (DST), a multiple transform selection (MTS), a low frequency non-separable transform (LFNST), or a non-separable primary transform (NSPT).

31. The method of claim 1, wherein IBC affine is used with a coding tool, and the coding tool includes at least one of: sign prediction or dependent quantization.

32. The method of claim 1, wherein one or more control point block vectors (CPBVs) are refined by at least one of: template matching or bilateral matching.

33. The method of claim 1, wherein an indication of whether and / or how to apply IBC affine depends on coding information, wherein the coding information includes video content.

34. The method according to any of the claims 1 to 33, wherein the video unit comprises at least one of: a color component, a sub-picture, a slice, a tile, a coding tree unit (CTU), a CTU row, a CTU group, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), a block, a sub-block of a block, a sub-region within a block, a region containing more than one sample or pixel.

35. The method according to any of the claims 1 to 33, wherein an indication whether and / or how the affine model is used to derive the BVs of the at least one sub-block is indicated at one of: sequence level, picture group level, picture level, slice level, or tile group level.

36. The method according to any of the claims 1 to 33, wherein an indication whether and / or how the affine model is used to derive the BVs of the at least one sub-block is indicated in one of: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or tile group header.

37. The method according to any of the claims 1 to 33, wherein whether and / or how the affine model is used to derive the BVs of the at least one sub-block depends on at least one of: a message indicated in one of: DPS, SPS, VPS, PPS, APS, picture header, slice header, tile group header, coding tree unit (CTU), coding unit (CU), CTU row, CTU group, TU, PU block, video coding unit, a position of a CU block, a position of a PU block, a position of a TU block, a position of a video coding unit, a block dimension of a current block, a block dimension of a neighboring block of the current block, a coding mode of a block, an indication of a color format, a coding tree structure, a slice group type, a tile group type, a slice picture type, a tile picture type, a color component, an ID of a temporal layer, a profile of a standard, a level of a standard, or a tier of a standard.

38. The method according to claim 37, wherein the coding mode of the block is at least one of: IBC inter mode, non-IBC inter mode, or non-IBC sub-block mode.

39. The method according to claim 37, wherein the indication of a color format is 4:2:0 or 4:4:

4.

40. The method according to claim 37, wherein the color component is applied to a chroma component or a luma component.

41. The method according to any of the claims 1 to 33, wherein the syntax element is binarized as a flag, a fixed length code, an EG(x) code, a unary code, a truncated unary code, or a truncated binary code.

42. The method according to any of the claims 1 to 33, wherein the syntax element is signed or unsigned.

43. The method according to any of claims 1 to 33, wherein the syntax element is coded with at least one context model, or is bypass coded.

44. The method according to any of claims 1 to 33, wherein the syntax element is signaled under at least one of the following conditions: if a corresponding functionality is applicable, or if a dimension of the block fulfills a condition.

45. The method according to claim 44, wherein the dimension is width and / or height.

46. The method according to any of claims 1 to 33, wherein the syntax element is signaled at one of the following: block level, sequence level, group of pictures level, picture level, slice level, or tile group level.

47. The method according to any of claims 1 to 33, wherein the syntax element is signaled in one of the following: coding structure of CTU, coding structure of CU, coding structure of TU, coding structure of PU, coding structure of CTB, coding structure of CB, coding structure of TB, coding structure of PB, sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependent parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or tile group header.

48. The method according to any of claims 1 to 33, wherein IBC with affine model is combined with at least one of the following coding tools: affine, multi-transform selection (MTS), low-frequency non-separable transform (LFNST), merge mode with motion vector difference (MMVD), MIP, ISP, cross-component linear model (CCLM), convolution cross-component model (CCCM), symmetric motion vector difference (SMVD), bi-directional optical flow (BDOF), decoder-side motion vector refinement (DMVR), history-based MVP (HMVP), template matching, intra block copy (IBC), or palette.

49. The method according to any of claims 1 to 33, wherein IBC with affine model is exclusive with at least one of the following coding tools: affine, multi-transform selection (MTS), LFNST, MMVD, MIP, ISP, CCLM, CCCM, SMVD, BDOF, DMVR, HMVP, template matching, IBC, or palette.

50. The method according to claim 49, wherein the exclusive coding tool is disabled without signaling if the IBC with affine model is used.

51. The method according to claim 49, wherein the IBC with affine model is disabled without signaling if the exclusive coding tool is used. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ 52. The method of any of claims 1-51, wherein the converting comprises encoding the video unit into the bitstream.

53. The method of any of claims 1-51, wherein the converting comprises decoding the video unit from the bitstream.

54. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of claims 1-53.

55. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any of claims 1-53.

56. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing, wherein the method comprises: deriving a block vector (BV) for at least one sub-block of a plurality of sub-blocks of a video unit of the video using an affine model, wherein a same block size is used for all of the plurality of sub-blocks, and / or at least one of parameters of the affine mode or a position of the video unit is applied to a subsequent video unit of the video unit; and generating the bitstream based on the BV.

57. A method for storing a bitstream of a video, comprising: deriving a block vector (BV) for at least one sub-block of a plurality of sub-blocks of a video unit of the video using an affine model, wherein a same block size is used for all of the plurality of sub-blocks, and / or at least one of parameters of the affine mode or a position of the video unit is applied to a subsequent video unit of the video unit; generating the bitstream based on the BV; and storing the bitstream in a non-transitory computer-readable recording medium. ​