Method and device for video processing and medium
By introducing the combined mode of intra-block copying and intra-prediction (CIBCIP) in video encoding and decoding technology, the problem of improving encoding and decoding efficiency in the prior art is solved, and more efficient video encoding and decoding performance is achieved.
Patent Information
- Application Number
- CN202380088328.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-21
- Filing Date
- 2023-12-20
- Publication Date
- 2025-09-02
AI Technical Summary
The existing video encoding and decoding technology has room for improvement in encoding and decoding efficiency, especially in the application of combining intra-block copying and intra-prediction modes.
A method is proposed to determine the prediction of the video unit based on the codec information, intra-block copying and intra-prediction combination mode (CIBCIP), thereby improving the codec performance and efficiency.
By combining intra-block copying and intra-prediction modes, the encoding and codec performance and efficiency of video encoding and codec are improved.
Smart Images

Figure CN120584484A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate generally to video processing techniques, and more particularly, to combined intra block copying and intra prediction. Background Art
[0002] Digital video capabilities are now being used in every aspect of our lives. For video encoding and decoding, various video compression technologies have been proposed, including MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Codec (AVC), ITU-T H.265 High Efficiency Video Codec (HEVC), and Versatile Video Codec (VVC). However, there is a general desire to further improve the encoding and decoding efficiency of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is provided. The method includes: determining, for conversion between a video unit of a video and a bitstream of the video unit, whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to the video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; and performing conversion based on the prediction for the video unit. This method can improve codec performance and codec efficiency.
[0005] In a second aspect, a device for video processing is provided. The device includes a processor and a non-volatile memory having instructions. When the instructions are executed by the processor, the processor performs the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores instructions, and the instructions cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a video bitstream, the video bitstream being generated by a method performed by an apparatus for video processing. The method includes: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to a video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; and generating a bitstream based on the prediction for the video unit.
[0008] In a fifth aspect, a method for storing a bitstream of a video is provided. The method includes: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to a video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; generating a bitstream based on the prediction for the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings.In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram illustrating an example video encoding and decoding system is shown according to some embodiments of the present disclosure;
[0012] Figure 2 shows a block diagram illustrating a first example video encoder according to some embodiments of the present disclosure;
[0013] Figure 3 shows a block diagram illustrating an example video decoder according to some embodiments of the present disclosure;
[0014] Figure 4 An example of an encoder block diagram is shown;
[0015] Figure 5 67 intra prediction modes are shown;
[0016] Figure 6Reference samples for wide-angle intra prediction are shown;
[0017] Figure 7 The discontinuity problem is shown when the orientation exceeds 45°.
[0018] Figure 8 Shows the MMVD search point;
[0019] Figure 9 This is an explanation of the symmetric MVD pattern;
[0020] Figure 10 Shows the extended CU area used in BDOF;
[0021] Figure 11 Shows an affine motion model based on control points;
[0022] Figure 12 Show the affine MVF of each sub-block;
[0023] Figure 13 shows the position of the inherited affine motion prediction value;
[0024] Figure 14 Shows control point motion vector inheritance;
[0025] Figure 15 The positions of candidate positions for the constructed affine Merge pattern are shown;
[0026] Figure 16 is a diagram of the use of motion vectors for the proposed combination method;
[0027] Figure 17 The sub-block MV VSB and the pixel Δv(i,j) are shown;
[0028] Figure 18A shows the spatial neighborhood blocks used by ATVMP;
[0029] Figure 18B Shows the derivation of sub-CU motion fields by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-CUs;
[0030] Figure 19 Local illumination compensation is shown;
[0031] Figure 20 It is shown that no downsampling is performed on the short edges;
[0032] Figure 21 shows decoding side motion vector refinement;
[0033] Figure 22 Shows the diamond-shaped area in the search area;
[0034] Figure 23 Shows the location of the spatial merge candidate;
[0035] Figure 24 Shows candidate pairs considered for redundancy check of spatial merge candidates;
[0036] Figure 25 It is an illustration of the motion vector scaling of the time domain Merge candidate;
[0037] Figure 26 Candidate positions C0 and C1 for the time domain Merge candidate are shown;
[0038] Figure 27 Shows the VVC spatial neighboring blocks of the current block;
[0039] Figure 28 is a diagram of the virtual blocks in the i-th search round;
[0040] Figure 29 An example of GPM partitions grouped at the same angle is shown;
[0041] Figure 30 Shows unidirectional prediction MV selection for geometric partitioning mode;
[0042] Figure 31 shows an example generation of warp weights w0 using geometric segmentation mode;
[0043] Figure 32 Show the spatial neighboring blocks used to derive spatial Merge candidates;
[0044] Figure 33 It shows that template matching is performed on the search area around the initial MV;
[0045] Figure 34 is a diagram of a sub-block to which OBMC is applied;
[0046] Figure 35 Shows SBT location, type and transformation type;
[0047] Figure 36 Shows the neighboring sample points used to calculate SAD;
[0048] Figure 37 Shows the neighboring samples used to calculate SAD for sub-CU level motion information;
[0049] Figure 38 Shows the sorting process;
[0050] Figure 39 shows the recorder process in the encoder;
[0051] Figure 40 shows the reordering process in the decoder;
[0052] Figure 41 is a graphic representation of the extended reference area;
[0053] Figure 42 Shows the IBC reference area depending on the current CU position;
[0054] Figure 43 An example showing symmetry in a picture of screen content;
[0055] Figure 44A This is an illustration of BV adjustment for horizontal flipping;
[0056] Figure 44B This is an illustration of the BV adjustment for vertical flipping;
[0057] Figure 45 Shows the intra-frame template matching search area used;
[0058] Figure 46 is a graphic representation of the template area;
[0059] Figure 47 shows the spatial portion of the convolution filter;
[0060] Figure 48 shows the reference region (and its filling) used to derive the filter coefficients;
[0061] Figure 49 Four Sobel-based gradient modes for GLM are shown;
[0062] Figure 50 Examples showing different numbers of samples in different reference rows used for fusion;
[0063] Figure 51 An example of different numbers of samples in different reference lines used for fusion is shown, and samples surrounded by boxes are discarded and not used for fusion of the reference lines;
[0064] Figure 52 Examples of different numbers of samples in different reference lines for fusion are shown, and the reference L n The sample points represented by the blank circles in are filled and used for the fusion of the reference line;
[0065] Figure 53 A flowchart illustrating a method for video processing according to an embodiment of the present disclosure; and
[0066] Figure 54 A block diagram illustrates a computing device in which various embodiments of the present disclosure may be implemented.
[0067] Throughout the drawings, the same or similar reference numbers generally refer to the same or similar elements. DETAILED DESCRIPTION
[0068] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that these embodiments are described only for the purpose of illustrating and helping those skilled in the art to understand and implement the present disclosure, and do not imply any limitation on the scope of the present disclosure. In addition to the methods described below, the disclosure described herein can also be implemented in various ways.
[0069] In the following description and claims, unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0070] References in this disclosure to "one embodiment," "an embodiment," "an example embodiment," and the like indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment is required to include that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an example embodiment, it is intended that such feature, structure, or characteristic, whether or not explicitly described, be applicable to other embodiments and that it is within the knowledge of those skilled in the art to apply such feature, structure, or characteristic.
[0071] It should be understood that although the terms "first" and "second" and the like may be used herein to describe various elements, these elements should not be limited to these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0072] The terms used herein are used only for the purpose of describing specific embodiments and are not intended to limit the example embodiments. As used herein, the singular forms "a," "an," and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprise," "including," "having," "including," and / or "comprising" when used herein indicate the presence of the features, elements, and / or components, etc., but do not exclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Sample Environment
[0073] Figure 1is a block diagram illustrating an example video codec system 100 that can utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0074] The video source 112 may include a source such as a video capture device. Examples of a video capture device include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0075] The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a coded representation of the video data. The bitstream may include coded pictures and associated data. The coded pictures are coded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The coded video data may be directly transmitted to the destination device 120 via the network 130A via the I / O interface 116. The coded video data may also be stored on a storage medium / server 130B for access by the destination device 120.
[0076] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modem. The I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to a user. The display device 122 may be integrated with the destination device 120, or may be external to the destination device 120, the destination device 120 being configured to interface with an external display device.
[0077] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other existing and / or future standards.
[0078] Figure 2is a block diagram illustrating an example of a video encoder 200 according to some embodiments of the present disclosure, which may be Figure 1 An example of the video encoder 114 in the system 100 is shown.
[0079] Video encoder 200 may be configured to implement any or all of the techniques of this disclosure. Figure 2 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0080] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a cache 213 and an entropy coding unit 214, and the prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.
[0081] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0082] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are described in detail in the following sections. Figure 2 are shown separately in the example.
[0083] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0084] The mode selection unit 203 can, for example, select one of a plurality of coding modes (intra-frame coding or inter-frame coding) based on the error result, and provide the resulting intra-frame coded block or inter-frame coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a joint intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on an inter-frame prediction signal and an intra-frame prediction signal. In the case of inter-frame prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel precision or integer pixel precision).
[0085] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing the current video block with one or more reference frames from the cache 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0086] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on the current video block, for example, depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" may refer to a portion of a picture consisting of macroblocks, all of which are based on macroblocks within the same picture. Furthermore, as used herein, in some aspects, "P slices" and "B slices" may refer to portions of a picture consisting of macroblocks that are independent of macroblocks in the same picture.
[0087] In some examples, motion estimation unit 204 may perform unidirectional prediction on the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0088] Alternatively, in other examples, the motion estimation unit 204 may perform bidirectional prediction on the current video block. The motion estimation unit 204 may search the reference pictures in list 0 for a reference video block for the current video block, and may also search the reference pictures in list 1 for another reference video block for the current video block. The motion estimation unit 204 may then generate multiple reference indices and multiple motion vectors, the multiple reference indices indicating multiple reference pictures in list 0 and list 1 containing multiple reference video blocks, and the multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 may output the multiple reference indices and multiple motion vectors for the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0089] In some examples, motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, motion estimation unit 204 may signal the motion information of the current video block with reference to the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of an adjacent video block.
[0090] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.
[0091] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0092] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner.Two examples of prediction signaling techniques that may be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge mode signaling.
[0093] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0094] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0095] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform a subtraction operation.
[0096] Transform processing unit 208 may generate one or more transform coefficient video blocks for a current video block by applying one or more transforms to the residual video block associated with the current video block.
[0097] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0098] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0099] After reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video blockiness artifacts in the video block.
[0100] The entropy encoding unit 214 may receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives the data, the entropy encoding unit 214 may perform one or more entropy encoding operations to generate entropy-encoded data and output a bitstream including the entropy-encoded data.
[0101] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 An example of the video decoder 124 in the system 100 is shown.
[0102] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 3 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0103] exist Figure 3 In the example of FIG. 3 , the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform a decoding process that is generally opposite to the encoding process described with respect to the video encoder 200.
[0104] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which motion information includes motion vectors, motion vector precision, reference picture list indexes, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, which includes deriving several most likely candidates based on data from adjacent PBs and reference pictures. The motion information typically includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indexes, and in the case of prediction regions in B slices, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" may refer to deriving motion information from spatially neighboring blocks or temporally neighboring blocks.
[0105] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filters used with sub-pixel precision may be included in the syntax elements.
[0106] Motion compensation unit 302 may calculate interpolated values for sub-integer pixels of a reference block using interpolation filters used by video encoder 200 during encoding of the video block. Motion compensation unit 302 may determine the interpolation filters used by video encoder 200 based on received syntax information, and motion compensation unit 302 may use the interpolation filters to produce a prediction block.
[0107] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks used to encode the (multiple) frames and / or (multiple) slices of the encoded video sequence, partition information describing how each macroblock of the picture of the encoded video sequence is partitioned, a mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding and decoding, signal prediction, and residual signal reconstruction. A slice can be an entire picture or a region of a picture.
[0108] The intra prediction unit 303 can use, for example, an intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0109] The reconstruction unit 306 can obtain the decoded block, for example, by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra-frame prediction unit 303. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra-frame prediction and also produces the decoded video for presentation on a display device.
[0110] Some exemplary embodiments of the present disclosure are described in detail below. It should be noted that the section headings used in this document are for ease of understanding and do not limit the embodiments disclosed in a section to that section. In addition, although some embodiments are described with reference to a multifunctional video codec or other specific video codecs, the disclosed technology is also applicable to other video coding and decoding technologies. In addition, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for de-encoding will be implemented by the decoder. In addition, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at a different compression bit rate. 1. Brief Overview The present invention relates to video coding techniques. Specifically, it relates to combined intra-block copying, in which a reference (or prediction) block is obtained using samples from the current image, intra-frame prediction, and other codec tools in image / video coding. It can be applied to existing video coding standards, such as HEVC or Versatile Video Codec (VVC). It may also be applicable to future video coding standards or video codecs. 2. Introduction Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. These include H.261 and H.263, produced by ITU-T; MPEG-1 and MPEG-4 Vision, produced by ISO / IEC; and the joint efforts of the two organizations to produce the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction and transform coding. To explore future video codec technologies beyond HEVC, the Joint Video Exploration Team (JVET) was jointly established by VCEG and MPEG in 2015. Since then, JVET has adopted many new approaches and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 / SC29 / WG11 (MPEG) was created with the goal of developing a VVC standard with a 50% bitrate reduction compared to HEVC. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC1 / SC 29 / WG 5) are investigating the potential need for standardization of future video codecs with compression capabilities significantly exceeding those of the current VVC standard. This future standardization effort could take the form of additional extensions to VVC or an entirely new standard. In a joint collaborative effort known as the Joint Video Exploration Team (JVET), working groups collaborate on this exploratory activity to evaluate compression technology designs proposed by experts in the field. New codec features and coding methods implemented in the Enhanced Compression Model (ECM) software, explored by the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG, are being investigated as potential enhancements to the video codec technology that exceed the capabilities of VVC. 2.1. Encoding and decoding process of typical video codecs Figure 4 An example of a VVC encoder block diagram is shown, which contains three loop filtering blocks: deblocking filter (DF), sample adaptive offset (SAO), and ALF. Unlike DF, which uses a predefined filter, SAO and ALF use the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a finite impulse response (FIR) filter, respectively, where the offset and filter coefficients are transmitted by the encoded side information. ALF is located in the last processing stage of each picture and can be seen as a tool that attempts to capture and repair artifacts produced by previous stages. Intra-mode coding with 67 intra-prediction modes To capture arbitrary edge directions present in natural videos, the number of directional intra modes is extended from 33 to 65, as used in HEVC. Figure 5 , and planar and DC modes remain the same. These more dense band-directional intra prediction modes apply to all block sizes and for both luma and chroma intra prediction. In HEVC, each intra-coded block has a square shape, and the length of each side is a power of 2. Therefore, no division operation is required to generate intra prediction values using DC mode. In VVC, in general, blocks can have a rectangular shape, and in general, a division operation is required for each block. To avoid division operations for DC prediction, only the longer side is used to calculate the average value for non-square blocks. 2.2.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. Conventional angular intra prediction directions are defined from 45 degrees to 135 degrees in a clockwise direction. In VVC, for non-square blocks, several traditional angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e. 67, and the intra mode encoding method remains unchanged. To support these predictions, Figure 7 As shown, an upper reference with a length of 2W+1 and a left reference with a length of 2H+1 are defined. The number of alternative modes in the wide-angle direction mode depends on the aspect ratio of the block. The alternative intra-frame prediction modes are shown in Table 2-1. Table 2-1 Intra-frame prediction modes replaced by wide-angle mode like Figure 7 As shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the increased interval Δp α If the wide-angle mode represents a non-fractional offset, there are eight modes in the wide-angle mode that meet this condition: [-14, -12, -10, -6, 72, 76, 78, 80]. When predicting a block using these modes, the samples in the reference buffer are directly copied without any interpolation. This modification reduces the number of samples that need to be smoothed. Furthermore, this approach aligns the design of the traditional prediction mode with the non-fractional mode in the wide-angle mode. In VVC, 4:2:2 chroma format and 4:4:4 chroma format as well as 4:2:0 chroma format are supported. The derivation table of chroma derivation mode (DM) for 4:2:2 chroma format was originally derived from HEVC, and the number of entries was expanded from 35 to 67 to align with the expansion of intra prediction mode. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes ranging from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for 4:2:2 chroma format is updated by replacing some values of the mapping table entries to more accurately convert the prediction angle for the chroma block. 2.3. Inter-frame prediction For each inter-predicted codec unit (CU), the motion parameters include motion vectors, reference picture indices, reference picture list usage indices, and additional information required for the new codec features of VVC to be used for inter-prediction sample generation. The motion parameters can be transmitted through signals in an explicit or implicit manner. When a CU is encoded and decoded in skip mode, the CU is associated with one PU and has no significant residual coefficients, no codec motion vector increments or reference picture indices. A Merge mode is specified in VVC, in which the motion parameters for the current CU are obtained from neighboring CUs, including spatial and temporal candidates, as well as other scheduling introduced in VVC. Merge mode can be applied to any inter-predicted CU, not only to skip mode. An alternative to Merge mode is explicit transmission of motion parameters, in which motion vectors, corresponding reference picture indices for each reference picture list and reference picture list usage flag, and other required information are explicitly signaled for each CU. 2.4. Intra-block copy (IBC) Intra-block copying (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the encoding and decoding efficiency of screen content materials. Since the IBC mode is implemented as a block-level encoding and decoding mode, block matching (BM) is performed at the encoder to find the optimal block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block, which has been reconstructed inside the current picture. The luminance block vector of the CU encoded and decoded by IBC has integer precision. The chrominance block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The IBC-encoded CU is regarded as a third prediction mode in addition to the intra-frame prediction mode or the inter-frame prediction mode. The IBC mode is applicable to CUs whose width and height are both less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks for blocks with a width or height no greater than 16 luma samples. For non-merge mode, a block vector search is first performed using a hash-based search. If the hash search does not return a valid candidate, a local search based on block matching is performed. In the hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The calculation of the hash key for each position in the current picture is based on 44 sub-blocks. For larger current block sizes, the hash key is determined to match the hash key of the reference block when all hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If multiple reference blocks are found whose hash keys match the hash key of the current block, the block vector cost of each matched reference is calculated, and the block vector cost with the minimum cost is selected. In the block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using a flag, and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows: – IBC Skip / Merge mode: The Merge candidate index is used to indicate which block vectors from the list of neighboring candidate IBC coded blocks are used to predict the current block. The Merge list consists of spatial, HMVP and pairwise candidates. – IBC AMVP mode: Block vector differences are encoded and decoded in the same way as motion vector differences. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the top neighbor. (If IBC codec is used.) When any neighbor is not available, the default block vector will be used as the prediction value. A flag is signaled to indicate the block vector prediction index. 2.5. IBC Movement Candidates The term "block" may refer to a codec tree block (CTB), a codec tree unit (CTU), a codec block (CB), a CU, a PU, a TU, a PB, a TB, or a video processing unit including multiple samples / pixels. A block may be rectangular or non-rectangular. For IBC coded blocks, a block vector (BV) is used to indicate the displacement from the current block to a reference block that has been reconstructed inside the current picture. W and H are the width and height of the current block (eg, luma block). The non-adjacent spatial candidates of the current codec block are the adjacent spatial candidates of the virtual block in the i-th search round (e.g. Figure 9). The width and height of the virtual block for the i-th search round are calculated as follows: newWidth = i × 2 × gridX + W, newHeight = i × 2 × gridY + H. Obviously, if the search round i is 0, the virtual block is the current block. In the following, the BV prediction value is also the BV candidate. The skip mode is also the merge mode. BV candidates can be divided into several groups based on some criteria. Each group is called a subgroup. For example, we can use adjacent spatial and temporal BV candidates as the first subgroup and the remaining BV candidates as the second subgroup. In another example, we can use the first N (N ≥ 2) BV candidates as the first subgroup, the next M (M ≥ 2) BV candidates as the second subgroup, and the remaining BV candidates as the third subgroup. 2.6. Merge Mode with MVD (MMVD) In addition to the Merge mode, in the case where the implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is transmitted by signal immediately after the regular Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a merge candidate is selected, it is further refined by signaled MVD information. This further information includes a merge candidate flag, an index specifying the magnitude of motion, and an index indicating the direction of motion. In MMVD mode, one of the first two candidates in the merge list is selected to serve as the MV basis. The MMVD candidate flag is signaled to specify which merge candidate to use between the first and second merge candidates. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 8 As shown in Table 2-2, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2-2. Table 2-2 - Relationship between distance index and predefined offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 2-3. It should be noted that the meaning of the MVD symbol can change according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV (i.e., the POCs of both references are greater than the POC of the current picture, or the POCs of both references are less than the POC of the current picture) and has two lists pointing to the same side of the current picture, the symbols in Table 2-3 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV and has two MVs pointing to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), and the POC difference in list 0 is greater than the POC difference in list 1, the symbols in Table 2-3 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and have opposite values for the signs of list 1 MVs. Otherwise, if the POC difference in List 1 is greater than the POC difference in List 0, then the sign in Table 2-3 specifies the sign of the MV offset added to the List 1 MV component of the starting MV and has the opposite value for the sign of List 0 MV. The MVD is scaled based on the POC difference in each direction. If the POC difference in both lists is the same, no scaling is required. Otherwise, if the POC difference in list 0 is greater than the POC difference in list 1, then the MVD for list 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb, as Figure 26 If the POC difference of L1 is greater than L0, then the MVD for list 0 is scaled in the same way. If the starting MV is unidirectionally predicted, then the MVD is added to the available MVs. Table 2-3 Symbols of MV offsets specified by direction index Direction IDX 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.7. Symmetric MVD Encoding and Decoding In VVC, in addition to the normal unidirectional prediction mode and bidirectional prediction mode MVD signaling, the symmetric MVD mode for bidirectional prediction MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not transmitted by signal but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, BiDirPredFlag is set equal to 0. Otherwise, if the nearest reference picture in list-0 and the nearest reference picture in list-1 form a reference picture forward-backward pair or a reference picture backward-forward pair, then BiDirPredFlag is set to 1, and both list-0 and list-1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2) At the CU level, if the CU is bidirectionally predicted and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is explicitly signaled. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. The reference index for list-0 and list-1 are set equal to the reference picture pair, respectively. MVD1 is set equal to (-MVD0). The final motion vector is shown below. In the encoder, symmetric MVD motion estimation begins with an initial MV evaluation. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, is included in JEM. Compared to the JEM version, BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplication operations and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if it meets all of the following conditions: – The CU is encoded and decoded using the “true” bi-prediction mode, i.e. one of the two reference pictures precedes the current picture in display order and the other reference picture follows the current picture in display order; – The distances from the two reference pictures to the current picture (i.e., POC differences) are the same; – Both reference images are short-term reference images; –CU is not encoded or decoded using Affine mode or SbTMVP Merge mode; –CU has more than 64 luma samples; – Both CU height and CU width are greater than or equal to 8 luma samples; – BCW weight index indicates equal weight; –WP is not enabled for the current CU; –CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes that the motion of the object is smooth. For each 4x4 sub-block, the motion refinement (v X ,v y ). Motion refinement is then used to adjust the bidirectional prediction sample values in the 4x4 sub-block. The following steps are applied in the BDOF process. First, the horizontal and vertical gradients of the two predicted signals are calculated by directly calculating the difference between two adjacent samples. and Right now, Among them, I (k) (i, j) is the sample value at coordinate (i, j) of the prediction signal in list k, k=0,1, and shift1 is calculated based on the luma bit depth bitDepth, because shift1=max(6, bitDepth-6). Then, the autocorrelation and cross-correlation of the gradients S1, S2, S3, S5 and S6 are calculated as in where Ω is a 6×6 window around the 4×4 sub-block, n a and n b The values of are set equal to min(1, bitDepth-11) and min(4, bitDepth-8), respectively. Then using the cross-correlation and autocorrelation terms, the motion refinement (v X ,v y ): in, th′ BIO =2 max(5,BD-7) . is the floor function, and Based on the motion refinement and gradients, the following adjustments are calculated for each sample in the 4x4 sub-block: Finally, the BDOF samples of the CU are calculated by adjusting the bidirectional prediction samples as follows: predBDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift (2-7) These values are chosen so that the multipliers in the BDOF process do not exceed 15 bits and the maximum bit width of the intermediate parameters in the BDOF process remains within 32 bits. In order to derive the gradient value, it is necessary to generate some predicted sample points I in the list k (k = 0, 1) outside the current CU boundary (k) (i,h). Figure 10 As shown, BDOF in VVC uses an extended row / column around the CU boundary. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples in the extended area (white positions) are generated by directly taking the reference samples at the nearest integer position (applying the floor() operation to the coordinates) without interpolation, and a conventional 8-tap motion compensation interpolation filter is used to generate the prediction samples inside the CU (gray positions). These extended sample values are only used for gradient calculations. For the rest of the steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it will be divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are regarded as CU boundaries in the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. The BDOF process can be skipped for each sub-block. When the SAD between the initial L0 prediction sample and the L1 prediction sample is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to equal to (8*W*(H>>1)), where W indicates the sub-block width and H indicates the sub-block height. In order to avoid the additional complexity of the SAD calculation, the SAD between the initial L0 prediction sample and the L1 prediction sample calculated in the DVMR process is reused here. If BCW is enabled for the current block, that is, the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, that is, luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. BDOF is also disabled when the CU is encoded or decoded using symmetric MVD mode or CIIP mode. 2.9. Joint Intra-Frame and Inter-Frame Prediction (CIIP) 2.10. Affine Motion Compensated Prediction In HEVC, only the translation motion model is applied for motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transformation motion compensation prediction is applied. Figure 11 As shown, the affine motion field of a block is described by the motion information of two control points (4-parameters) or the motion vectors of three control points (6-parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as: Among them (mv 0x ,mv 0y ) is the motion vector of the upper left control point, (mv 1x ,mv 1y ) is the motion vector of the upper right control point, and (mv 2x ,mv 2y ) is the motion vector of the lower left control point. In order to simplify the motion compensation prediction, block-based affine transformation prediction is applied. In order to derive the motion vector of each 4×4 luminance sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation, as follows: Figure 12 As shown, and rounded to 1 / 16 fractional precision. A motion compensated interpolation filter is then applied to generate a prediction for each subblock with the derived motion vector. The subblock size of the chroma component is also set to 4×4. The MV of a 4×4 chroma subblock is calculated as the average of the MVs of the four corresponding 4×4 luminance subblocks. As in translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.10.1 Affine Merge Prediction AF_MERGE mode can be applied to CUs with width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of the spatially adjacent CUs. There can be up to 5 CPMVP candidates, and the index is transmitted by signal to indicate a CPMVP candidate to be used for the current CU. The following three types of CPMV candidates are used to form the affine Merge candidate list: – Inherited affine merge candidates inferred from the CPMV of neighboring CUs; – The constructed affine Merge candidate CPMVP derived using the translation MV of the neighboring CU; –Zero MV. In VVC, there are at most two inherited affine candidates derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. Figure 13 . For the left prediction value, the scanning order is A0→A1, and for the upper prediction value, the scanning order is B0→B1→B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMVP candidate in the affine Merge list of the current CU. As shown, if the neighboring lower left block A is encoded and decoded in affine mode, the motion vectors v2, v3 and v4 of the upper left, upper right and lower left corners of the CU containing block A are obtained. When block A is encoded and decoded using a 4-parameter affine model, the two CPMVs of the current CU are calculated based on v2 and v3. In the case where block A is encoded and decoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated based on v2, v3 and v4. A constructed affine candidate is a candidate that is constructed by combining the neighboring translation motion information of each control point. The motion information for the control point is obtained from Figure 15 CPMV is derived from the specified spatial and temporal neighboring blocks shown. k (k=1, 2, 3, 4) represents the kth control point. For CPMV1, the B2→B3→A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1→B0 blocks are checked, and for CPMV3, the A1→A0 blocks are checked. If available, TMVP is used as CPMV4. After obtaining the MVs of the four control points, the affine merge candidate is constructed based on the motion information. The following combinations of control point motion vectors are used in the following order to construct: {CPMV1,CPMV2,CPMV3},{CPMV1,CPMV2,CPMV4},{CPMV1,CPMV3,CPMV4}, {CPMV2,CPMV3,CPMV4},{CPMV1,CPMV2},{CPMV1,CPMV3}. The combination of 3 CPMVs constructs a 6-parameter affine Merge candidate, and the combination of 2 CPMVs constructs a 4-parameter affine Merge candidate. To avoid the motion scaling process, the relevant combination of control point MVs is discarded if the reference indices of the control points are different. After the inherited affine merge candidates and the constructed affine merge candidates are checked, if the list is still not complete, a zero MV is inserted to the end of the list. 2.10.2 Affine AMVP Prediction Affine AMVP mode can be applied to CUs with width and height both greater than or equal to 16. An affine flag in the CU level is signaled in the bitstream to indicate whether affine AMVP mode is used, and another flag is signaled to indicate whether 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value is signaled in the bitstream. The affine AVMP candidate list size is 2 and is generated by using the following four types of CPMV candidates in sequence: – Inherited affine merge candidates inferred from the CPMV of neighboring CUs; – The constructed affine Merge candidate CPMVP derived using the translation MV of the neighboring CU; – Translational MV from neighboring CU; –Zero MV. The order in which inherited affine AMVP candidates are checked is the same as the order in which inherited affine Merge candidates are checked. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as the current block are considered. When the inherited affine motion prediction value is inserted into the candidate list, the deduplication process is not applied. The AMVP candidates are constructed from Figure 15 The same check order is derived from the specified spatial neighbors shown in . The same check order is used in the affine Merge candidate construction. In addition, the reference picture index of the neighboring blocks is also checked. In the check order, the first block that is inter-coded and has the same reference picture as the current CU is used. Only when the current CU is encoded and decoded using a 4-parameter affine mode and both mv0 and mv1 are available, they are added to the affine AMVP list as a candidate. When the current CU is encoded and decoded using a 6-parameter affine mode and all three CPMVs are available, the three CPMVs are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidate is set to unavailable. If the affine AMVP list candidate is still less than 2 after the inherited affine AMVP candidate and the constructed AMVP candidate are checked, mv0, mv1 and mv2 will be added in sequence as translation MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, the affine AMVP list is filled with zero MVs. 2.10.3 Affine Motion Information Storage In VVC, the CPMV of an affine CU is stored in a separate buffer. The stored CPMV is only used to generate the inherited CPMVP for subsequently coded CUs in affine Merge mode and affine AMVP mode. The sub-block MV derived from the CPMV is used for motion compensation, MV derivation of the Merge / AMVP list of translation MVs, and deblocking. To avoid picture row cache for additional CPMV, affine motion data inheritance in CUs from above CTUs is treated differently compared to inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the above CTU row, the bottom left and bottom right sub-block MVs in the row cache are used instead of CPMVs for affine MVP derivation. In this way, CPMVs are only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to a 4-parameter model. Figure 16 As shown, along the top CTU boundary, the bottom left sub-block motion vector and the bottom right sub-block motion vector of the CU are used for affine inheritance of the CU in the CTU below. 2.10.4 Prediction Refinement Using Optical Flow for Affine Mode Compared with pixel-based motion compensation, sub-block based affine motion compensation can save memory access bandwidth and reduce computational complexity at the expense of prediction accuracy loss. In order to achieve finer granularity of motion compensation, prediction refinement using optical flow (PROF) is used to refine sub-block based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after performing sub-block based affine motion compensation, the brightness prediction samples are refined by adding the difference derived from the optical flow equation. PROF is described in the following four steps: Step 1) Perform sub-block based affine motion compensation to generate sub-block prediction I(i,j). Step 2) Use a 3-tap filter [-1, 0, 1] to calculate the spatial gradient g of the sub-block prediction at each sample point x (i,j) and g y (i, j). The gradient calculation is exactly the same as that in BDOF. g x (i,j)=(I(i+1,j)>>shift1)-(I(i-1,j)>>shift1) (2-10) g y (i,j)=(I(i,j+1)>>shift1)-(I(i,j-1)>>shift1) (2-11) Shift1 is used to control the accuracy of the gradient. For gradient calculation, the sub-block (i.e., 4×4) prediction is extended by one sample on each side. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extension boundary are copied from the nearest integer pixel position in the reference picture. Step 3) Calculate the brightness prediction refinement by the following optical flow equation. ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) (2-12) Wherein, Δv(i,j) is the difference between the sample point MV calculated for the sample point position (i,j) (denoted by v(i,j)) and the sub-block MV of the sub-block to which the sample point (i,j) belongs, as Figure 17 As shown in . Δv(i,j) is quantized in units of 1 / 32 luma sample accuracy. Since the affine model parameters and the sample positions relative to the sub-block center are invariant from sub-block to sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let dx(i,j) and dy(i,j) be the distance from the sample position (i,j) to the sub-block (x SB ,y SB )’s horizontal and vertical offsets, Δv(x,y), can be derived from the following equations: In order to maintain accuracy, the sub-block (x SB ,y SB ) is calculated as ((W SB -1) / 2,(H SB -1) / 2), where W SB and H SB are the sub-block width and sub-block height respectively. For a 4-parameter affine model, For the 6-parameter affine model, Among them, (v 0x ,v 0y ),(v 1x ,v 1y ),(v 2x ,v 2y ) are the upper left, upper right and lower left control point motion vectors, w and h are the width and height of the CU. Step 4) Finally, the luma prediction refinement ΔI(i,j) is added to the sub-block prediction I(i,j). The final prediction I' is generated as the following equation. I′(i,j)=I(i,j)+ΔI(i,j) (2-17) PROF is not applied in two cases for affine-coded CUs: 1) all control point MVs are the same, indicating that the CU has only translational motion; 2) the affine motion parameters are larger than the specified limit, because the sub-block based affine MC is downgraded to the CU based MC to avoid large memory access bandwidth requirements. A fast coding method is applied to reduce the coding complexity of affine motion estimation using PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) If the CU is not a root block and its parent block does not select the affine mode as its best mode, PROF is not applied because the possibility of selecting the affine mode as the best mode for the current CU is low; b) If the magnitudes of the four affine parameters (C, D, E, F) are all less than a predefined threshold and the current picture is not a low-latency picture, PROF is not applied because the improvement introduced by PROF is small for this case. In this way, affine motion estimation using PROF can be accelerated. 2.11. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and merge mode for the CU in the current picture. The same co-located picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: –TMVP predicts motion at CU level, but SbTMVP predicts motion at sub-CU level; – While TMVP obtains the temporal motion vector from the co-located block in the co-located picture (the co-located block is the lower right block or the center block relative to the current CU), SbTMVP applies motion displacement before obtaining the temporal motion information from the co-located picture, where the motion displacement is obtained based on the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process Figure 18A and Figure 18B As shown in Figure 18 , SbTMVP predicts the motion vector of a sub-CU within the current CU in two steps. In the first step, the spatial neighbor A1 in Figure 18 is checked. If A1 has a motion vector that uses the co-located picture as its reference picture, then this motion vector is selected as the motion displacement to be applied. If this motion is not identified, the motion displacement is set to (0,0). In the second step, the motion offset identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the co-located picture, as Figure 18B shown. Figure 18B The example in assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block in the co-located image (i.e., the minimum motion grid covering the center sample) is used to derive the motion information for the sub-CU. After the motion information of the co-located sub-CU is identified, the motion information is converted into a motion vector and reference index for the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub-block based Merge candidate list is used for signaling based on the sub-block Merge mode, which contains both SbTMVP candidates and affine Merge candidates. The SbTMVP mode is enabled / disabled by the sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block based Merge candidates, followed by the affine Merge candidate. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed to 8×8, and as for the affine Merge mode, the SbTMVP mode is only applicable to CUs whose width and height are both greater than or equal to 8. The encoding logic for the additional SbTMVP Merge candidate is the same as that for other Merge candidates, ie, for each CU in a P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.12. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (the difference between the CU's motion vector and the predicted motion vector) is signaled in units of quarter-luminance-samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be encoded and decoded with different precisions. Depending on the mode for the current CU (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows: – Normal AMVP mode: quarter-luma-sample, half-luma-sample, integer-luma-sample or quadruple-luma-sample. – Affine AMVP mode: quarter-luma-sample, integer-luma-sample, or 1 / 16-luma-sample. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both horizontal MVD and vertical MVD for reference list L0 and reference list L1) are zero, then quarter-luminance-sample MVD resolution is inferred. For CUs with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luma-sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luma-sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half-luma-sample or other MVD precision (integer-luma-sample or quadruple-luma-sample) is used for normal AMVP CUs. In the case of half-luma-sample, a 6-tap interpolation filter is used for half-luma-sample positions instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer-luma-sample or quadruple-luma-sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, a second flag is used to indicate whether integer-luma-sample MVD precision or 1 / 16-luma-sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luma-sample, half-luma-sample, integer-luma-sample, or quadruple-luma-sample), the motion vector prediction value for the CU will be rounded to the same precision as the MVD before being added together with the MVD. The motion vector prediction value is rounded towards zero (i.e., negative motion vector prediction values are rounded towards positive infinity, while positive motion vector prediction values are rounded towards negative infinity). The encoder uses RD checks to determine the motion vector resolution for the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, VTM11 only conditionally calls RD checks for MVD precisions other than quarter-luma-sample. For normal AVMP mode, the RD costs for quarter-luma-sample and integer-luma-sample MVD precisions are first calculated. The RD cost for integer-luma-sample MVD precision is then compared with the RD cost for quarter-luma-sample MVD precision to determine whether further checking for quadruple-luma-sample MVD precision is necessary. When the RD cost for quarter-luma-sample MVD precision is significantly less than that for integer-luma-sample MVD precision, the RD check for quadruple-luma-sample MVD precision is skipped. Subsequently, if the RD cost for integer-luma-sample MVD precision is significantly greater than the best RD cost of the previously tested MVD precision, the check for half-luma-sample MVD precision is skipped. For the affine AMVP mode, if the affine inter mode is not selected after checking the rate-distortion cost of the affine Merge / Skip mode, Merge / Skip mode, quarter-luma-sample MVD precision normal AMVP mode, and quarter-luma-sample MVD precision affine AMVP mode, the 1 / 16-luma-sample MV precision and 1 pixel MV precision affine inter modes are not checked. In addition, the affine parameters obtained in the quarter-luma-sample MV precision affine inter mode are used as the starting search points in the 1 / 16 luma sample and quarter-luma-sample MV precision affine inter modes. 2.13. Bidirectional Prediction Using CU-Level Weights (BCW) In HEVC, the bidirectional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bidirectional prediction mode is extended to allow weighted averaging of the two prediction signals, rather than just a simple average. P bi-pred =((8-w)*P0+w*P1+4)>>3 (2-18) Five weights, w∈{-2,3,4,5,10}, are allowed in weighted average bidirectional prediction. For each bidirectionally predicted CU, the weight w is determined in one of two ways: 1) For non-merge CUs, the weight index is signaled after the motion vector difference; 2) For merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-latency pictures, all 5 weights are used. For non-low-latency pictures, only 3 weights (w∈{3,4,5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing encoder complexity. These algorithms are summarized below. For further details, refer to the VTM software. When combined with AMVR, unequal weights are only conditionally checked for 1-pixel motion vector accuracy and 4-pixel motion vector accuracy if the current picture is a low-latency picture. When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only checked conditionally. – Unequal weights are not searched when certain conditions are met, which depend on the POC distance between the current picture and its reference pictures, the codec QP, and the temporal level. The BCW weight index is encoded using a context codec bit followed by a bypass codec bit. The first context codec bit indicates whether equal weights are used; and if unequal weights are used, an additional bit is signaled using bypass codec to indicate that unequal weights are used. Weighted prediction (WP) is a codec tool supported by the H.264 / AVC and HEVC standards to efficiently encode and decode video content with attenuation. Support for WP has also been added to the VVC standard. WP allows weighting parameters (weights and offsets) to be signaled for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW, which would complicate the VVC decoder design, if a CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For constructed affine Merge mode, affine motion information is constructed based on motion information of up to 3 blocks. The BCW index for a CU using constructed affine Merge mode is simply set equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be applied jointly to a CU. When a CU is encoded or decoded using CIIP mode, the BCW index of the current CU is set to 2, i.e., equal weight. 2.14. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is a codec tool used to address the problem of local illumination changes between the current picture and its temporal reference picture. LIC is based on a linear model where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples of the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x,y)=α·P r (x+v x ,y+v y )+β Among them, P(x,y) is the prediction signal of the current block at the coordinate (x,y); P r (x+v x ,y+v y ) is the motion vector (v x ,v y ) points to the reference block; α and β are the corresponding scaling factors and offsets applied to the reference block. Figure 19 The LIC process is shown in FIG. Figure 19 In the , when LIC is applied to a block, the minimum mean square error (LMSE) method is used to derive the values of the LIC parameters (i.e., α and β), which is calculated by minimizing the neighboring samples of the current block (i.e., Figure 19The template T in the temporal reference picture) and its corresponding reference sample in the temporal reference picture (ie, Figure 19 In addition, in order to reduce the computational complexity, both the template samples and the reference template samples are downsampled (adaptive downsampling) to derive the LIC parameters, i.e., only Figure 19 α and β are derived from the shaded points in . In order to improve the encoding and decoding performance, the short side is not downsampled, such as Figure 20 shown. 2.15. Decoder-side Motion Vector Refinement (DMVR) In order to improve the accuracy of MV in Merge mode, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In bidirectional prediction operation, the refined MV is searched around the initial MV in reference picture list L0 and reference picture list L1. The BM method calculates the distortion between the two candidate blocks in reference picture list L0 and list L1. Figure 21 As described in , the SAD between two blocks is calculated based on each MV candidate (eg, MV0' and MV1') around the initial MV. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, the application of DMVR is restricted and only applies to CUs that are coded or decoded with the following modes and features: – CU-level Merge mode using bidirectional prediction MV – Relative to the current picture, one reference picture is in the past and the other reference picture is in the future. – The distances from the two reference pictures to the current picture (ie, POC differences) are the same. – Both reference images are short-term reference images. –CU has more than 64 luma samples – CU height and CU width are both greater than or equal to 8 luma samples. –BCW weight index indicates equal weight – The current block is not WP enabled – CIIP mode is not used for the current block. The refined MV derived by the DMVR process is used to generate inter-frame prediction samples and is also used in temporal motion vector prediction for future picture encoding and decoding. The original MV is used in the deblocking process and is also used in spatial motion vector prediction for future CU encoding and decoding. Additional features of DMVR are mentioned in the following sub-items. 2.15.1. Search Scheme In DVMR, the search point is around the initial MV, and the MV offset obeys the MV difference mirror rule. In other words, any point examined by DMVR (represented by the candidate MV pair (MV0, MV1)) obeys the following two equations: MV0′=MV0+MV_offset (2-19) MV1′=MV1-MV_offset (2-20) Where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. A 25-point full search is applied to the integer sample offset search phase. The SAD of the initial MV pair is first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample phase of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search phase. To reduce the adverse effects of DMVR refinement uncertainty, it is proposed to prefer the original MV during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value. The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using the parametric error surface equation, rather than an additional search using SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. When the integer sample search phase terminates at the center with the minimum SAD in either the first or second iteration of the search, fractional sample refinement is further applied. In the sub-pixel offset estimation based on the parametric error surface, the center position cost and the costs at four neighboring positions from the center are used to fit a 2-D parabolic error surface equation of the following form: E(x,y)=A(xx min ) 2 +B(yy min ) 2 +C (2-21) Where (x min ,y min ) corresponds to the fractional position with the minimum cost and C corresponds to the minimum cost value. By solving the above equation using the cost values of the five search points, we calculate: x min =(E(-1,0)-E(1,0)) / (2(E(-1,0)+E(1,0)-2E(0,0))) (2-22) y min=(E(0,-1)-E(0,1)) / (2((E(0,-1)+E(0,1)-2E(0,0))) (2-23) Since all cost values are positive and the minimum is E(0,0), then x min and y min is automatically constrained to be between -8 and 8. This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fraction (x min ,y min ) is added to the integer distance refinement MV to obtain the sub-pixel accurate refinement delta MV. 2.15.2. Bilinear interpolation and sample filling In VVC, the resolution of the video image (MV) is 1 / 16 luma samples. Samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are centered around the original fractional pixel MV with integer sample offsets, so the DMVR search process requires interpolation of samples at those fractional positions. To reduce computational complexity, a bilinear interpolation filter is used to generate the fractional samples used in the DMVR search process. Another important effect is that by using a bilinear filter, DVMR does not access more reference samples than the conventional motion compensation process within a 2-sample search range. After obtaining the refined MV using the DMVR search process, a conventional 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples than the normal motion compensation process, samples that are not required in the interpolation process based on the original MV but are required in the interpolation process based on the refined MV are padded with available samples. 2.15.3 Maximum DMVR Processing Unit When the width and / or height of a CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size for the DMVR search process is limited to 16x16. 2.16. Multi-pass decoder-side motion vector refinement In this contribution, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the codec block. In the second pass, BM is applied to each 16×16 sub-block within the codec block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MV is stored for both spatial and temporal motion vector prediction. 2.16.1. First pass - Block-based bilateral matching MV refinement In the first pass, a refined MV is derived by applying BM to the codec block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MV based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive integer sample precision intDeltaMV and half-pixel sample precision halfDeltaMv. The local search applies a 3×3 square search pattern to loop with the search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension and the maximum value of sHor and sVer is 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV or halfDeltaMV local search is terminated. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for minimum cost continues until it reaches the end of the search range. The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: MV0_pass1=MV0+deltaMV MV1_pass1=MV1-deltaMV. 2.16.2. Second pass - Sub-block based bilateral matching MV refinement In the second pass, refined MVs are derived by applying BM to 16×16 grid sub-blocks. For each sub-block, a refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass on reference picture lists L0 and L1. Refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each subblock, BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range of [-sHor, sHor] in the horizontal direction and a search range of [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimension and the maximum value of sHor and sVer is 8. The bilateral matching cost is calculated by applying the cost factor to the SATD cost between the two reference sub-blocks, such as: bilCost = satdCost * costFactor. The search area (2*sHor+1)*(2*sVer+1) is divided into up to 5 diamond search areas, such as Figure 22 As shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond-shaped area is processed in order starting from the center of the search area. Within each area, the search points are processed in raster scan order, from the upper left corner of the area to the lower right corner of the area. When the minimum bilCost within the current search area is less than the threshold equal to sbW*sbH, the integer pixel (int-pel) full search is terminated, otherwise, the int-pel full search continues to the next search area until all search points are checked. BM performs a local search to derive the half-sample accuracy halfDeltaMv. The search pattern and cost function are the same as those defined in 2.9.1. The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). The refined MV of the second pass is then derived as: ·MV0_pass2(sbIdx2)=MV0_pass1+deltaMV(sbIdx2) ·MV1_pass2(sbIdx2)=MV1_pass1-deltaMV(sbIdx2). 2.16.3 Third pass - Sub-block-based bidirectional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8×8 grid sub-blocks. For each 8×8 sub-block, BDOF refinement is applied to derive scaled Vx and Vy instead of clipping from the refined MV of the parent sub-block in the second pass. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample accuracy and clipped between -32 and 32. The refined MVs (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) at the third pass are derived as: ·MV0_pass3(sbIdx3)=MV0_pass2(sbIdx2)+bioMv ·MV1_pass3(sbIdx3)=MV0_pass2(sbIdx2)-bioMv. 2.17. Sample-based BDOF In sample-based BDOF, instead of deriving the motion refinement (Vx, Vy) on a block basis, the derivation is performed per sample. The codec block is divided into 8×8 sub-blocks. For each sub-block, a decision is made whether to apply BDOF by checking the SAD between two reference sub-blocks against a threshold. If the decision is made to apply BDOF to a sub-block, a sliding 5×5 window is used for each sample in the sub-block, and the existing BDOF process is applied to each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the center sample of the window. 2.18. Extended Merge Prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: (1) Spatial MVP from the adjacent CU (2) Temporal MVP from the same CU (3) History-based MVP from FIFO table (4) Paired Average MVP (5) Zero MV. The size of the merge list is signaled in the sequence parameter set header, and the maximum allowed size of the merge list is 6. For each CU codec in merge mode, the index of the best merge candidate is encoded using truncated unary binarization (TU). The first binary bit of the merge index is coded using context, and bypass coding is used for the remaining binary bits. The derivation process of Merge candidates for each category is provided in this session. As in HEVC, VVC also supports parallel derivation of Merge candidate lists for all CUs in a region of a certain size. 2.18.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as in HEVC, except that the positions of the first two Merge candidates are swapped. Up to four Merge candidates are selected from the candidates located in the depicted positions. The order of derivation is B0, A0, B1, A1 and B2. Position B2 is only considered when one or more CUs in positions B0, A0, B1, A1 are not available (for example, because they belong to another strip or slice) or are intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to a redundancy check that ensures that candidates with the same motion information are excluded from the list, so that the coding efficiency is improved. In order to reduce computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. On the contrary, if the corresponding candidates for the redundancy check do not have the same motion information, only the corresponding candidates are considered. Figure 24 For pairs linked by arrows, only candidates are added to the list. 2.18.2. Time Domain Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on the co-located CU belonging to the co-located reference picture. The reference picture list used for the derivation of the co-located CU is explicitly signaled in the slice header. The scaled motion vector for the temporal Merge candidate is obtained as Figure 25 As shown by the dotted line in , the scaled motion vector is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal merge candidate is set equal to zero. The position for the time domain candidate is selected between candidates C0 and C1, such as Figure 26 If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used to derive the temporal merge candidate. 2.18.3. History-Based Merge Candidate Derivation After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of the previously coded block is stored in a table and used as the MVP for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (cleared). Whenever there is a non-sub-block inter-coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate. The HMVP table size S is set to 6, which indicates that up to 6 historically based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table, and all subsequent HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The latest HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidate. Redundancy check for spatial domain Merge candidates or temporal domain Merge candidates is applied to HMVP candidates. In order to reduce the number of redundancy check operations, the following simplifications are introduced: The number of HMPV candidates used for Merge list generation is set to (N<=4)·M:(8-N), where N indicates the number of existing candidates in the Merge list and M indicates the number of HMVP candidates available in the table. Once the total number of available Merge candidates reaches the maximum allowed Merge candidate number minus 1, the Merge candidate list construction process from HMVP is terminated. 2.18.4. Pairwise Average Merge Candidate Derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing merge candidate list. These predefined pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the merge indexes into the merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged, even if they refer to different reference images. If only one motion vector is available, that motion vector is used directly. If no motion vector is available, the list remains invalid. When the Merge List is not full after pairwise average Merge candidates are added, zero MVPs are inserted at the end until the maximum number of Merge candidates is reached. 2.18.5.Merge Estimation Region Merge Estimation Region (MER) allows independent derivation of Merge candidate lists for CUs in the same Merge Estimation Region (MER). The generation of the Merge candidate list for the current CU does not include candidate blocks in the same MER as the current CU. In addition, the update process for the history-based motion vector prediction candidate list is only updated when the following conditions are met: (xCb+cbWidth)>>Log2ParMrgLevel is greater than xCb>>Log2ParMrgLevel, and (yCb+cbHeight)>>Log2ParMrgLevel is greater than (yCb>>Log2ParMrgLevel), where (xCb, yCb) is the top left luma sample position of the current CU in the picture, and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled as log2_parallel_merge_level_minus2 in the sequence parameter set. 2.19. New Merge Candidates 2.19.1. Non-adjacent Merge Candidate Derivation In VVC, Figure 27 The five spatial neighboring blocks and one temporal neighboring block shown are used to derive Merge candidates. It is proposed to use the same pattern as in VVC to derive additional Merge candidates from locations that are not adjacent to the current block. To achieve this, for each search round i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated as follows: Offsetx=-i×gridX, Offsety=-i×gridY Where Offsetx and Offsetty represent the offset of the upper left corner of the virtual block relative to the upper left corner of the current block, and gridX and gridY are the width and height of the search grid. Second, calculate the width and height of the virtual block by: newWidth=i×2×gridX+currWidth newHeight=i×2×gridY+currHeight Where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block. gridX and gridY are currently set to currWidth and currHeight respectively. Figure 28Shows the relationship between the virtual block and the current block. After generating the virtual block, block A i 、B i 、C i 、D i and E i The VVC spatial neighbors of the virtual block can be considered and their positions are obtained using the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i 、B i 、C i 、D i and E i It is the spatial neighboring block used in VVC Merge mode. When building the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighboring blocks are utilized. After the time domain Merge candidates, non-adjacent spatial domain Merge candidates are inserted into the Merge list in the order of B1 → A1 → C1 → D1 → E1. 2.19.2STMVP It is proposed to use three spatial domain Merge candidates and one temporal domain Merge candidate to derive the average candidate as the STMVP candidate. The STMVP is inserted before the upper left spatial merge candidate. All previous merge candidates in the merge list are used to deduplicate STMVP candidates. For spatial candidates, the first three candidates in the current Merge candidate list are used. For temporal candidates, the same positions as VTM / HEVC co-location are used. For spatial candidates, the first, second and third candidates in the current Merge candidate list are inserted before the STMVP, which are denoted as F, S and T. The temporal candidate having the same position as the VTM / HEVC co-location used in TMVP is denoted as Col. The motion vector of the STMVP candidate in prediction direction X (denoted as mvLX) is derived as follows: 1) If the reference indices of the four Merge candidates are all valid and are all equal to zero in the prediction direction X (X=0 or 1), then mvLX=(mvLX_F+mvLX_S+mvLX_T+mvLX_Col)>>2 2) If the reference indexes of three of the four Merge candidates are valid and in the prediction direction X (X=0 or 1) is equal to zero, then mvLX=(mvLX_F×3+mvLX_S×3+mvLX_Col×2)>>3 or mvLX=(mvLX_F×3+mvLX_T×3+mvLX_Col×2)>>3 or mvLX=(mvLX_S×3+mvLX_T×3+mvLX_Col×2)>>3 3) If the reference indexes of two of the four Merge candidates are valid and in the prediction direction X (X=0 or 1) is equal to zero, then mvLX=(mvLX_F+mvLX_Col)>>1 or mvLX=(mvLX_S+mvLX_Col)>>1 or mvLX=(mvLX_T+mvLX_Col)>>1. NOTE: If time domain candidates are not available, STMVP mode is turned off. 2.19.3.Merge List Size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8. 2.20. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter-frame prediction. Geometric partitioning mode is a type of Merge mode and is signaled using a CU-level flag. Other Merge modes include normal Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. For each possible CU size w×h=2 m ×2 n , where m,n∈{3…6} and excluding 8×64 and 64×8, the geometric partitioning mode supports a total of 64 partitions. When using this mode, the CU is positioned by a straight line ( Figure 29) is divided into two parts. The position of the dividing line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter-frame predicted using its own motion; only unidirectional prediction is allowed for each partition, that is, each part has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with regular bidirectional prediction, only two motion compensated predictions are required for each CU. The unidirectional prediction motion for each partition is derived using the process described in 2.20.1. If geometric partitioning mode is used for the current CU, a geometric partitioning index indicating the partitioning mode (angle and offset) of the geometric partitioning and two Merge indices (one index for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS and the syntax binarization for the GPM Merge index is specified. After predicting each part in the geometric partitioning, a blending process with adaptive weights as in 2.20.2. is used to adjust the sample values along the geometric partitioning edges. This is the prediction signal for the entire CU, and the transform and quantization process will be applied to the entire CU as in other prediction modes. Finally, as in 2.20.3., the motion field of the CU predicted using the geometric partitioning mode is stored. 2.20.1. One-way prediction candidate list construction The unidirectional prediction candidate list is directly derived from the merge candidate list constructed according to the extended merge prediction process in 2.18. n represents the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X is equal to the parity of n) of the nth extended merge candidate is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are Figure 30 In the example, the LX motion vector corresponding to the n-th extended Merge candidate does not exist, the motion vector L(1-X) in the same candidate is used instead as the unidirectional prediction motion vector for the geometric partitioning mode. 2.20.2. Blending Along Geometric Partition Edges After predicting each part of the geometric partition using its own motion, blending is applied to the two prediction signals to derive samples around the geometric partition edges. A blending weight is derived for each position of the CU based on the distance between each position and the partition edge. The distance from the segmentation edge to the position (x, y) is derived as: where i,j are the indices of the angle and offset for the geometric partition, which depend on the geometric partition index transmitted by the signal. x,j and ρ y,jThe sign of depends on the angle index i. The weights of each part of the geometric segmentation are derived as follows: wIdxL(x,y)=partIdx? 32+d(x,y):32-d(x,y) (2-28) w1(x,y)=1-w0(x,y) (2-30) partIdx depends on the angle index i. An example of weight w0 is Figure 31 As shown in . 2.20.3. Motion Field Storage for Geometric Partitioning Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and a combination Mv of Mv1 and Mv2 are stored in the motion region of the geometric partition mode codec CU. The type of stored motion vector for each individual position in the motion field is determined as: sType=abs(motionIdx)<32?2:(motionIdx≤0?(1-partIdx):partIdx) (2-31) Where motionIdx is equal to d(4x+2,4y+2), which is recalculated according to equation (2-18). partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise if sType is equal to 2, the combined Mv from Mv0 and Mv2 is stored. The following process is used to generate the combined Mv: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-directional prediction motion vector. Otherwise, if Mv1 and Mv2 are from the same list, only the unidirectional predicted motion Mv2 is stored. 2.21. Multi-hypothesis prediction In Multi-Hypothesis Prediction (MHP), up to two additional prediction values are signaled on top of Inter AMVP mode, Normal Merge mode, Affine Merge and MMVD mode. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal. p n+1 =(1-α n+1 )p n +α n+1 h n+1 The weighting factor α is specified according to the following Tables 2-4. Table 2-4 - Weighting factors for MHP add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For inter-AMVP mode, if unequal weights in BCW are selected in bi-prediction mode, only MHP is applied. The additional assumptions can be Merge mode or AMVP mode. In the case of Merge mode, the motion information is indicated by the Merge index, and the Merge candidate list is the same as the candidate list in the geometric partitioning mode. In the case of AMVP mode, the reference index, MVP index and MVD are transmitted through the signal. 2.22. Non-adjacent airspace candidates Non-adjacent spatial merge candidates are inserted after TMVP in the regular merge candidate list. The mode of spatial merge candidates is Figure 32 As shown above, the distance between non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. Template Matching (TM) Template Matching (TM) is a decoder-side MV derivation method to refine the motion information of the current CU by finding the closest match between the template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and the blocks in the reference picture (i.e., the same size as the template). Figure 33 As shown, a better MV is searched around the initial motion of the current CU within the search range of [-8, +8] pixels. Two modified template matching methods are proposed: the search step size is determined based on the AMVR mode, and the TM can be cascaded with the bilateral matching process in the Merge mode. In AMVP mode, MVP candidates are determined based on template matching error to pick the MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then TM is performed only for that specific MVP candidate for MV refinement. TM refines the MVP candidate by using an iterative diamond search, starting with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode) within the [-8, +8] pixel search range. The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for 4-pixel AMVR mode), followed by searching sequentially through half-pixels and quarter-pixels depending on the AMVR mode as specified in Table 2-5. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. Table 2-5 Search mode of AMVR and Merge mode with AMVR In Merge mode, a similar search method is applied to the Merge candidate indicated by the Merge index. As shown in Table 2-5, depending on whether the alternative interpolation filter (the interpolation filter used when AMVR is in half-pixel mode) is used based on the Merge motion information, TM can be performed all the way to 1 / 8 pixel MVD accuracy or skip the accuracy after half-pixel accuracy. In addition, when TM mode is enabled, template matching can be run as an independent process or as an MV refinement process between the block-based bilateral matching (BM) method and the sub-block-based BM method, depending on whether BM can be enabled according to its enabling condition check. 2.24. Overlapped Block Motion Compensation (OBMC) Overlapped Block Motion Compensation (OBMC) has been previously used in H.263. In JEM, unlike in H.263., OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, OBMC is performed for all motion compensated (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it is applied to both the luminance component and the chrominance components. In JEM, an MC block corresponds to a codec block. When a CU is encoded and decoded using sub-CU modes (including Sub-CU Merge, Affine, and FRUC modes), each sub-block of the CU is an MC block. In order to handle CU boundaries in a uniform manner, OBMC is performed at the sub-block level for all MC block boundaries, where the sub-block size is set to be equal to 4×4, as shown in FIG. Figure 34 shown. When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of four connected neighboring sub-blocks (if available and different from the current motion vector) are also used to derive a prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block. The prediction block based on the motion vector of the neighboring sub-block is denoted as P N , where N indicates the index for the upper, lower, left, and right sub-blocks, and the prediction block based on the motion vector of the current sub-block is denoted as P C When P N When the motion information of the neighboring sub-block is used, the neighboring sub-block contains the same motion information as the current sub-block and does not contain the motion information of the neighboring sub-block. N Execute OBMC. Otherwise, P N Each sample point is added to P C The same point in the N Four rows / columns are added to P C The weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N, and the weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C The exception is small MC blocks (ie, when the height or width of the codec block is equal to 4 or the CU is coded in sub-CU mode), for which only P N Two rows / columns are added to P C In this case, weighting factors {1 / 4, 1 / 8} are used for P N , and weighting factors {3 / 4,7 / 8} are used for P C For the P generated based on the motion vector of the vertical (horizontal) adjacent sub-block N , P N The samples in the same row (column) of are added to P with the same weighting factor C . In JEM, for CUs with a size less than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied for the current CU. For CUs with a size greater than 256 luma samples or that are not encoded or decoded using AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top and left neighboring blocks is used to compensate for the top and left boundaries of the current CU's original signal, and then the normal motion estimation process is applied. 2.25. Multi-Transform Selection (MTS) for Core Transforms In addition to the DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is used for residual coding of both inter-frame and intra-frame codec blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the basis functions of the selected DST / DCT. Table 2-6 Transform basis functions of DCT-II / VIII and DST VII for N-point input To maintain the orthogonality of the transform matrix, the transform matrix is quantized more accurately than the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within 16 bits, all coefficients will have 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is applied only for luma. MTS signaling is skipped when one of the following conditions applies: – The position of the last significant coefficient for the luma TB is less than 1 (i.e., DC only); – The last significant coefficient of the luminance TB is located within the MTS zero region. If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally signaled to indicate the transform type for the horizontal and vertical directions, respectively. The transform and signaling mapping table is shown in Table 2-7. A unified transform selection for ISP and implicit MTS is used by removing the intra mode and block shape dependencies. If the current block is in ISP mode, or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are turned on, only DST7 is used for both the horizontal transform kernel and the vertical transform kernel. In terms of transform matrix accuracy, an 8-bit main transform kernel is used. Therefore, all transform kernels used in HEVC are kept the same, including 4-point DCT-2 and 4-point DST-7, 8-point DCT-2, 16-point DCT-2 and 32-point DCT-2. In addition, other transform cores including 64-point DCT-2, 4-point DCT-8, 8-point DST-7, 16-point DST-7, 32-point DST-7 and 8-point DCT-8, 16-point DCT-8, 32-point DCT-8 also use the 8-bit main transform core. Table 2-7 Conversion and signaling mapping table To reduce the complexity of large-sized DST-7 and DCT-8, high-frequency transform coefficients are zeroed for DST-7 and DCT-8 blocks with size (width or height, or both) equal to 32. Only coefficients in the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded using transform skip mode. To avoid syntax coding redundancy, the transform skip flag is not signaled when the CU-level MTS_CU_flag is not equal to zero. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for inter-frame coded blocks, implicit MTS can still be enabled. 2.26. Sub-Block Transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, a sub-portion of the residual block is only coded for the CU. When there is an inter-predicted CU with cu_cbf equal to 1, cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-portion of the residual block is coded. In the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a portion of the residual block is coded using the inferred adaptive transform, and the other portion of the residual block is zeroed. When SBT is used for inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, such as Figure 35 As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in 2:2 partitioning or 1:3 / 3:1 partitioning. The 2:2 partitioning is similar to the binary tree (BT) partitioning, while the 1:3 / 3:1 partitioning is similar to the asymmetric binary tree (ABT) partitioning. In the ABT partitioning, only small areas contain non-zero residuals. If one dimension of the CU is 8 in luminance samples, 1:3 / 3:1 partitioning along that dimension is not allowed. There are up to 8 SBT modes for a CU. Position-dependent transform kernel selection is applied to the luma transform block in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, in Figure 35 The horizontal and vertical transforms for each SBT position are specified in [1]. For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DCT-7, respectively. When a side of the residual TU is larger than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slice, cbf, and horizontal and vertical kernel transform types for the residual block. SBT is not applied to CUs coded with combined inter-intra mode. 2.27. Adaptive Merge Candidate Reordering Based on Template Matching To improve encoding and decoding efficiency, after constructing the merge candidate list, the order of each merge candidate is adjusted based on the template matching cost. Merge candidates are arranged in the list according to the template matching cost in ascending order. They are operated on in the form of subgroups. The template matching cost is measured by the SAD (sum of absolute differences) between the neighboring samples of the current CU and its corresponding reference samples. If the Merge candidate includes bidirectional predictive motion information, the corresponding reference sample is the average of the corresponding reference sample in reference list 0 and the corresponding reference sample in reference list 1, such as Figure 36 If the Merge candidate includes sub-CU level motion information, the corresponding reference sample is composed of the neighboring samples of the corresponding reference sub-block, as shown in Figure 37 As shown in . The sorting process is operated in the form of subgroups, such as Figure 38 As shown in the figure, the first three merge candidates are sorted together. The next three merge candidates are sorted together. The template size (width of the left template or height of the upper template) is 1. The subgroup size is 3. 2.28. Adaptive Merge Candidate List It can be assumed that the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first subgroup, and the last 3 Merge candidates are taken as the second subgroup (ie, the last subgroup). For the encoder, after the Merge candidate list is constructed, some Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, such as Figure 39 shown. More specifically, the template matching costs for the Merge candidates in all subgroups except the last subgroup are calculated; then the Merge candidates in their own subgroups except the last subgroup are reordered; finally, the final Merge candidate list is obtained. For the decoder, after the Merge candidate list is constructed, some / no Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, as Figure 40 As shown. Figure 40 In , the subgroup located in the selected (signaled) Merge candidate is called the selected subgroup. More specifically, if the selected Merge candidate is located in the last subgroup, the Merge candidate list construction process is terminated after the selected Merge candidate is derived, no reordering is performed, and the Merge candidate list is not changed; otherwise, the execution process is as follows: After all Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; the template matching costs for the Merge candidates in the selected subgroup are calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained. For both encoder and decoder, The template matching cost is derived as a function of T and RT, where T is a set of samples in the template and RT is a set of reference samples for the template. When deriving the reference samples of the template for the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision. For bidirectional prediction, the reference samples (RT) of the template are derived by taking a weighted average of the reference samples of the template in reference list 0 (RT0) and the reference samples of the template in reference list 1 (RT1), as shown below. RT=((8-w)*RT0+w*RT1+4)>>3 (2-32) The weights of the reference templates in reference list 0 (8-w) and reference list 1 (w) are determined by the BCW index of the merge candidate. BCW indices of {0, 1, 2, 3, 4} correspond to w values of {-2, 3, 4, 5, 10}, respectively. If the local illumination compensation (LIC) flag of the Merge candidate is true, the reference samples of the template are derived using the LIC method. The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT. The template size is 1. That is, the width of the left template and / or the height of the upper template is 1. If the codec mode is MMVD, the Merge candidates used to derive the basic Merge candidate are not reordered. If the coding mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered. 2.29. IBC with extended reference area An IBC reference area design that does not increase the current memory area required by ECM-3 is proposed and its performance is tested. Figure 41 The design is shown. In the figure, the blue square represents the current CTU, and the green square represents the CTU that can be used by the IBC reference. Specifically, assuming that W represents the maximum horizontal CTU index and the current CTU index is (m, n), for the codec unit in the current CTU, the CTUs with indices (0, n)...(m, n) and (m-1, n)...(W, n) define the reference area that can be used by IBC. One reason for this design is that in the current ECM, the left, upper, and upper-left CTUs are being used and therefore need to be preserved. To achieve this, all CTUs to the right of the upper CTU in the upper CTU row (for the CTU to be encoded in the current CTU row) and all CTUs to the left of the current CTU in the current CTU row (for the CTU to be encoded in the next CTU row) must be preserved. This means that this design does not increase the buffer size required for the current ECM. 2.30. IBC with Template Matching It is proposed to also use template matching with IBC for both IBC Merge mode and IBC AMVP mode. Compared to the Merge list used by the regular IBC Merge mode, the IBC-TM Merge list has been modified so that candidates are selected based on a deduplication method with the motion distance between candidates in the regular TM Merge mode. The zero motion padding at the end (meaningless for intra coding) has been replaced with motion vectors pointing to the left (-W, 0), above (0, -H), and top-left (-W, -H) CUs, and the list is filled with left candidates (without deduplication) if necessary. In IBC-TM Merge mode, template matching methods are utilized to refine the selected candidates before the RDO or decoding process.The IBC-TM Merge mode has been competed with the regular IBC Merge mode, and the TM Merge flag is signaled. In IBC-TM AMVP mode, up to 3 candidates are selected from the IBC Merge list. Each of the 3 selected candidates is refined using a template matching method and ranked according to their resulting template matching cost. Only the first two candidates are then typically considered in the motion estimation process. Template matching refinement for both IBC-TM Merge and AMVP modes is very simple since IBC motion vectors are constrained to be integers and within the reference region, e.g. Figure 42 As shown in . Therefore, in IBC-TM Merge mode, all refinements are performed with integer precision, and in IBC-TM AMVP mode, all refinements are performed with integer precision or 4-pixel precision. In both cases, the refined motion vectors in each refinement step must satisfy the constraints of the reference region. 2.31. Reconstruction Reordering IBC (RR-IBC) Screen content codecs similar to intra block copy (IBC) generate prediction blocks by directly copying previously coded reference areas in the same picture. Symmetry is often observed in video content, especially in text character areas and computer-generated graphics in screen content sequences, such as Figure 43 Therefore, a specific screen content codec that takes symmetry into account will effectively compress such video content. A reconstruction-reordering inter-block coding (R-IBC) mode is proposed for screen content video coding. When applied, the samples in the reconstructed block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is derived without flipping. On the decoder side, the reconstructed block is flipped back to restore the original block. The RR-IBC codec block supports two flipping methods, horizontal flipping and vertical flipping. First, for the IBC AMVP codec block, a syntax flag is transmitted by signaling to indicate whether the reconstruction is flipped, and if flipped, another flag is further transmitted by signaling to specify the flipping type. For IBC Merge, the flipping type is inherited from the adjacent block without syntax signaling. Taking into account horizontal symmetry or vertical symmetry, the current block and the reference block are usually aligned horizontally or vertically. Therefore, when horizontal flipping is applied, the vertical component of BV is not transmitted by signaling and is presumed to be equal to 0. Similarly, when vertical flipping is applied, the horizontal component of BV is not transmitted by signaling and is presumed to be equal to 0. In order to better utilize the symmetry characteristics, a flip-aware BV adjustment method is applied to refine the block vector candidates. Figure 44A and Figure 44B As shown, (x nbr ,y nbr ) and (x cur ,y cur ) represent the center sample coordinates of the adjacent blocks and the center sample coordinates of the current block, BV nbr and BV cur Represents the BV of the neighboring block and the BV of the current block respectively. Instead of inheriting the BV directly from the neighboring block, the BV is obtained by adding the BV in the BV when the neighboring block is coded and decoded using horizontal flipping. nbr The horizontal component (denoted as BV nbr h ) to calculate BV cur The horizontal component, BV cur h =2(x nbr -x cur )+BV nbr h Similarly, in the case where the adjacent block is coded using vertical flipping, bynbr The vertical component (denoted as BV nbr v ) to calculate BV cur The vertical component, BV cur v =2(y nbr -y cur )+BV nbr v . 2.32. Intra-frame template matching Intra Template Matching (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed portion of the current frame, whose L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template that is most similar to the current template and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. The prediction signal is obtained by combining the L-shaped causal neighbors of the current block with Figure 45 It is generated by matching another block in a predefined search area in , which consists of the following: R1: current CTU; R2: upper left CTU; R3: upper CTU; R4: left CTU. SAD is used as the cost function. In each region, the decoder searches for the template that has the minimum SAD with respect to the current template and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. That is: SearchRange_w=a*BlkW SearchRange_h=a*BlkH Where 'a' is a constant that controls the gain / complexity tradeoff. In practice, 'a' is equal to 5. Intra template matching tools are enabled on width and height for CUs of size less than or equal to 64. This maximum CU size for intra template matching is configurable. When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level through a dedicated flag. 2.33. Intra-frame prediction fusion The intra prediction fusion method uses multiple prediction values generated from different modes / reference lines. In subtest a, multiple intra prediction values are generated and then fused by weighted averaging. The process of deriving the prediction value to be used in the fusion process is described as follows: 1) For the angular intra prediction mode in the single mode case including TIMD and DIMD, the proposed method derives the intra prediction value by weighting the intra predictions obtained from multiple reference lines, denoted as p fusion =w0p line +w1p line+1 , where p line is intra prediction from the default reference line, and p line+1 is the prediction from the row above the default reference row. The weights are set to w0=3 / 4 and w1=1 / 4. 2) For TIMD mode using hybrid, p line is used in the first mode (w0=1, w1=0) and p line+1 Used in the second mode (w0=0, w1=1). 3) For DIMD mode with hybrid, the number of prediction values selected for weighted averaging is increased from 3 to 6. In subtest b, intra prediction fusion is performed on reference lines instead of prediction blocks. line and r line+1 ) is used for intra prediction fusion. The corresponding intra prediction angle DeltaInt is considered in the fusion process. The fusion reference line (r fusion Each value in [i]) is derived from: r fusion [i]=(3·r line [i]+r line+1 [i+DeltaInt])>>2 When the angular intra mode has a non-integer slope (reference sample interpolation required) and the block size is greater than 16, the proposed intra prediction fusion is applied to the luma block, which is used together with MRL and not applied to ISP codec blocks. In the method studied in subtest a, PDPC is applied to the intra prediction mode that uses the reference line closest to the current block. 2.34. Template-based Multi-reference Line Intra Prediction (TMRL) The proposed TMRL model includes the following aspects: a) Extending the reference row candidate list and the intra prediction mode candidate list. The extended reference row candidate list used in this scheme is {1, 3, 5, 7, 12}. The restriction on the top CTU row remains unchanged. The size of the intra prediction mode candidate list is 10. The construction of the intra prediction mode candidate list is similar to that of the MPM. The differences are: The planar mode is excluded from the proposed intra prediction mode candidate list. If the DC mode is not already included, then the DC mode is added after the modes of the 5 neighboring PUs and the DIMD mode. Angular modes with incremental angles from ±1 to ±4 are added (compared to the existing angular modes in the intra prediction mode candidate list). b) Construction of TMRL candidate list. For a block, there are 5×10=50 combinations between extended reference lines and allowed intra prediction modes. Since the extended reference lines start from reference line 1, the area covered by reference line 0 is used for template matching. Template area (see Figure 46 ) is calculated between the prediction (generated from 50 combinations) and the reconstruction. The 20 combinations with the smallest SAD costs are selected in ascending order to form the TMRL candidate list. c) TMRL signaling Instead of directly encoding and decoding the reference line and intra mode, an index into the TMRL candidate list is encoded to indicate which combination of reference line and prediction mode is used to encode the current block. In the proposed TMRL mode, a truncated Golomb-Rice codec with a divisor of 4 is used to encode and decode the selected combination from the combination list. The binarization process and codeword are shown in Table 2-8. Table 2-8 Binarization of TMRL index d) Modification on the encoder side Encoder side modifications are tested to further improve the codec efficiency. For intra blocks larger than 8×8, additional TMRL RDO is added if no TMRL mode is selected by SATD comparison. 2.35 Convolutional Cross-Component Model (CCCM) for Intra Prediction We propose to apply a convolutional cross-component model (CCCM) to predict chroma samples from reconstructed luma samples in a similar spirit as the current CCLM mode. Like CCLM, when chroma downsampling is used, the reconstructed luma samples are downsampled to match the lower resolution chroma grid. Again, similar to CCLM, there are options to use a single model of CCCM or a multi-model variant of CCCM. The multi-model variant uses two models, one derived for samples above the average luminance reference value and another derived for the remaining samples (following the spirit of CCLM design). The multi-model CCCM mode can be selected for PUs with at least 128 available reference samples. 2.35.1 Convolutional Filters The proposed convolutional 7-tap filter consists of a 5-tap plus a cross-shaped spatial component, a nonlinear term, and a bias term. The input to the spatial 5-tap component of the filter consists of the center (C) luma sample co-located with the chroma sample to be predicted, as well as its upper / north (N) neighbor, lower / south (S) neighbor, left / west (W) neighbor, and right / east (E) neighbor, as shown below in Figure 47 As shown in . The nonlinear term P is expressed as a power of 2 of the center luma sample C and scaled to the sample value range of the content: P=(C*C+midVal)>>bitDepth That is, for 10-bit content, it is calculated as: P=(C*C+512)>>10 The bias term B represents a scalar offset between the input and output (similar to the offset term in CCLM) and is set to the intermediate chrominance value (512 for 10-bit content). The output of the filter is calculated as the filter coefficient c i Convolution with the input value and clipped to the range of valid chroma samples: predChromaVal=c0C+c1N+c2S+c3E+c4W+c5P+c6B. 2.35.2 Calculation of filter coefficients The filter coefficients c are calculated by minimizing the MSE between the predicted chroma samples and the reconstructed chroma samples in the reference region. i . Figure 48 The reference region is shown as consisting of six rows of chroma samples above and to the left of the PU. The reference region extends one PU width to the right and one PU height below the PU boundary. The region is adjusted to include only available samples. The extension shown in blue is required to support "side samples" in the cross-shaped spatial filter and to fill in unavailable areas. MSE minimization is performed by computing the autocorrelation matrix for the luma input and the cross-correlation vector between the luma input and the chroma output. The autocorrelation matrix is LDL-decomposed, and back-substitution is used to compute the final filter coefficients. This process roughly follows the calculation of the ALF filter coefficients in ECM; however, LDL decomposition is chosen instead of Cholesky decomposition to avoid the use of square root operations. The proposed method uses only integer arithmetic. 2.35.3 Bitstream Signaling The use of the mode is signaled using the CABAC codec PU-level flag. A new CABAC context is included to support this process. When it comes to signaling, CCCM is considered a sub-mode of CCLM. That is, if the intra prediction mode is LM_CHROMA_IDX (to enable single-mode CCCM) or MMLM_CHROMA_IDX (to enable multi-mode CCCM), only the CCCM flag is signaled. 2.36 Gradient Linear Model (GLM) Compared to CCLM, GLM uses luma sample gradients to derive linear models instead of downsampled luma values. Specifically, when GLM is applied, the input to the CCLM process (i.e., downsampled luma samples L) is replaced by luma sample gradients G. The other parts of CCLM (e.g., parameter derivation, linear transformation of prediction samples) remain unchanged. C=α·G+β For signaling, when CCLM mode is enabled for the current CU, two flags are transmitted through the signal for the Cb and Cr components respectively to indicate whether GLM is enabled for each component; if GLM is enabled for a component, a syntax element is further signaled to select one of the four gradient filters for gradient calculation. Enable four gradient filters for GLM, such as Figure 49 , which shows four Soble-based gradient modes for GLM. 3. Question In the current design of IBC, the prediction of the current block is obtained using samples in the current picture indicated by the block vector. IBC's encoding and decoding performance is very good for screen content videos with repeated content. However, for natural content videos, due to different characteristics, the coding gain of IBC is much lower than that for screen content videos. 4. Detailed solution The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow sense. In addition, these solutions can be combined in any way. In the present disclosure, intra block copy (IBC) is not limited to the current IBC technology, but can be interpreted as a technology in which the reference (or prediction) block is obtained using samples in the current slice / slice / sub-picture / picture / other video unit (e.g., CTU row), excluding conventional intra prediction methods. The reference row may refer to a row reconstruction sample and / or a column reconstruction sample that is adjacent to or non-adjacent to the current block, and is used to derive intra-frame prediction of the current video unit via an interpolation filter along a specific direction, and the specific direction is determined by an intra-frame prediction mode (e.g., conventional intra-frame prediction using an intra-frame prediction mode), or by weighting reference samples of the reference row using a matrix or vector (e.g., MIP) to derive intra-frame prediction of the current video unit. In this disclosure, IBC-GPM may refer to an encoding tool that uses IBC in a video unit to obtain prediction for at least one sub-partition when the video unit is geometrically divided into more than one sub-partition. In this disclosure, IBC-LIC may refer to a coding tool in which local illumination compensation (LIC) is used to refine video units using IBC codec. In the following discussion, IBC may be replaced by other codec tools that rely on coded / decoded / reconstructed information within the same region, eg, color palette, intra template matching. Combination of intra block copy and intra prediction (CIBCIP or IBC-CIIP) 1. It is proposed to use a combination of intra block copy (IBC) and intra prediction (CIBCIP) to derive the prediction / reconstruction of a video unit, which is obtained by fusing the IBC prediction signal and the intra prediction signal. a. In one example, P(x,y)=w ip1 *IP1(x,y)+w ip2 *IP2(x,y)+…+w ipn *IP n (x,y)+ w ibc1 *IBC1(x,y)+w ibc2 *IBC2(x,y)+…+w ibcm *IBC m (x,y), where P(x,y) is the generated predicted value, IP k is the prediction generated by the k-th intra prediction, IBC j is the prediction generated by the jth IBC, and w ipk and w ibcjis the corresponding weighted value. i. In one example, n=1 or 2, or 3 and m=1. ii. In one example, n=1 or 2, or 3 and m=2. iii. In one example, n=1 and m=1, or 2, or 3. iv. In one example, n=2 and m=1, or 2, or 3. b. In one example, intra prediction can be done by angular intra prediction, DC, planar, cross-component prediction (CCLM), multi-model CCLM, left CCLM, upper CCLM, CCCM, left / upper CCCM, GLM or their variants. i. Intra prediction mode The intra prediction mode can be encoded and decoded using MPM or TIMD or DIMD or any other method to transmit the intra prediction mode through signaling. ii. A specific set of intra prediction modes may be allowed to be used in CIBCIP. c. In one example, IBC Merge mode may be used. i. In one example, one or more BV candidates in the IBC Merge list may be allowed to be used in CIBCIP. ii. In one example, one or more BV offsets may be used for CIBCIP. 1) In one example, a BV offset may be added to the BV candidate before the BV candidate is used to obtain an IBC prediction. 2) In one example, the BV offset may be signaled or derived. iii. In one example, the IBC Merge mode may be at least one of a conventional IBC Merge mode, an IBC-MBVDMerge mode, and an IBC-TM Merge mode. d. In one example, the IBC AMVP mode may be used. i. In one example, one or more BV prediction values in the IBC AMVP list may be allowed to be used in the CIBCIP. 1) In one example, how the BV prediction value is derived and used for a block decoded with CIBCIP codec can be the same as for a block decoded with non-CIBCIP codecs. a) Alternatively, how the BV prediction value is derived and used for blocks decoded with CIBCIP codecs may be different from blocks decoded with non-CIBCIP codecs. ii. In one example, the Block Vector Difference (BVD) used in CIBCIP can be signaled in the same manner as the IBC mode. 1) Alternatively, the BVD may not be signaled but predefined. iii. In one example, the BVD can be derived using codec information. e. In one example, a Merge index (mergeIdx) indicating a BV candidate in the IBC Merge list and / or a BVP index (bvpIdx) indicating a BV prediction value in the IBC AMVP list used to obtain an IBC prediction signal may be transmitted through a signal. 1) In one example, the binarization or signaling method of the Merge index or BVP index may be the same as that in the IBC mode. 2) Alternatively, a Merge index or a BVP index may be predefined, for example, mergeIdx=0 or mergeIdx=1; bvpIdx=0 or bvpIdx=1. 3) Alternatively, the codec information can be used to derive the Merge index or BVP index. 4) Alternatively, template matching (eg, with minimum template matching cost) may be used to derive the Merge index or BVP index. f. In one example, the construction of the IBC Merge list or IBC AMVP list used in CIBCIP mode may be the same as or different from that used in IBC mode. g. In one example, the number of BV candidates (N) that can be used in the IBC Merge (or AMVP) list of CIBCIP is less than or equal to the number of BV candidates (M) that can be used in the IBC Merge (or AMVP) list of IBC. N is an integer greater than 0 and less than or equal to M. i. In one example, N=1, or N=2, or N=3, or N=4, or N=5, or N=6. ii. In one example, the top N BV candidates of the IBC Merge (or AMVP) list may be used for CIBCIP. h. In one example, template matching can be used to derive / refine the BV, which is used to obtain the IBC prediction signal. i. In one example, the BV offset can be derived using template matching, which is added to the BV candidate in the IBCMerge list. ii. In one example, a template matching-based approach can be used to derive the BVD. iii. In one example, a template matching based approach can be used to derive the BVD signature. iv. In one example, the intra prediction mode or intra prediction method used to obtain the intra prediction signal can be used for template matching to derive / refine the BV. i. In one example, the BV list may be reordered before being used in CIBCIP. i. In one example, template matching or bilateral matching costs can be used for re-ranking. ii. In one example, template matching or bilateral matching may be used during the construction of the BV list used for CIBCIP. iii. In one example, the BV list may refer to the IBC Merge list or the IBC AMVP list. iv. In one example, the reordering method for the BV list for CIBCIP can be the same as that for IBC. v. Alternatively, the reordering method for the BV list for CIBCIP may be different from that for IBC. 1) In one example, the number of BV candidates (N1) in the reordered BV list for CIBCIP may be less than or equal to the number of BV candidates (M1) in the reordered BV list for IBC mode. a) In one example, when IBC Merge mode is used for CIBCIP, N1=1 or 2 or 3 or 4. b) In one example, when IBC AMVP is used for CIBCIP, N1=1 or 2 or 3. j. In one example, intra prediction may refer to a conventional intra prediction method (e.g., intra prediction using 35 intra prediction modes in HEVC or 67 intra prediction modes in VVC), or other intra prediction methods that utilize samples in the current slice / slice / sub-picture / picture / other video unit (e.g., CU, PU, TU, CTU, CTU row) excluding IBC to obtain a prediction block. i. In one example, one or more predefined intra prediction modes may be used to obtain an intra prediction signal. 1) In one example, the predefined intra prediction mode may refer to planar mode, DC mode, horizontal mode, and vertical mode. ii. In one example, one or more of the most probable modes (MPMs) may be used to obtain an intra prediction signal. iii. In one example, an intra prediction mode may be used to obtain an intra prediction signal, the intra prediction mode being derived using a block vector used to obtain an IBC prediction signal. iv. In one example, an intra prediction signal may be obtained using an intra prediction mode derived using a template based method such as TIMD. v. In one example, an intra prediction signal may be obtained using an intra prediction mode derived using neighboring samples or gradients of neighboring samples (such as DIMD). vi. In one example, ISP can be used to obtain intra-frame prediction signals. vii. In one example, MIP can be used to obtain the intra prediction signal. viii. In one example, MRL may be used to obtain an intra prediction signal. ix. In one example, the intra prediction signal may be intra template matching prediction (IntraTMP). x. In one example, PDPC / gradient PDPC can be used to obtain the intra prediction signal. xi. In one example, intra prediction fusion can be used to obtain the intra prediction signal. 1) Alternatively, intra prediction fusion may not be used to generate the intra prediction signal. xii. In one example, TMRL may be used to obtain an intra prediction signal. xiii. In one example, at least one codec tool may differ from conventional intra prediction. 1) In one example, the codec tool may specify how to fill the reference samples, or whether and / or how to filter the reference samples, or whether and / or how to apply the filtering process (e.g., PDPC / Gradient PDPC), or whether and / or how to use an interpolation filter. 2) Alternatively, the coding tools used to obtain the intra prediction signal can be the same as conventional intra prediction. k. In one example, the fused / final prediction signal may be refined through a filtering process. i. In one example, the filtering process may refer to PDPC or gradient PDPC. 2. In one example, weighting parameters for fusing the IBC prediction signal and the intra prediction signal may be transmitted via a signal or derived. a. In one example, the weighting parameters can be transmitted via signals. i. In one example, a set of weighting parameters is constructed and an index indicating the weighting parameters can be transmitted through a signal. b. In one example, the weighting parameters can be derived using codec information. i. In one example, the codec information may refer to the codec mode of the neighboring unit. 1) In one example, the weighting parameters may depend on whether one or more neighboring units are coded in intra prediction or IBC mode. ii. In one example, the codec information may refer to an intra prediction mode used to obtain an intra prediction signal. iii. In one example, the codec information may refer to the block size or block dimension of the current video unit and / or the neighboring video units. iv. In one example, a template matching method may be used (eg, with minimum template matching cost) Derive weighting parameters. v. In one example, a template matching based approach can be used to derive weighting parameters. c. In one example, the weighting parameters may be predefined. d. In one example, one intra prediction signal and one IBC prediction signal are used in CIBCIP. i. In one example, P(x,y)=w ip1 *IP1(x,y)+w ibc1 *IBC1(x,y), where w ip1 +w ibc1 =1, where wipl1+wibc1=1. ii. In one example, P(x,y)=(w ip1 *IP1(x,y)+w ibc1 *IBC1(x,y)+offset)>>shift, where w ip1 +w ibc1 =(1< <shift)。 1) In one example, offset=0. 2) In one example, offset=1<<(shift-1). 3) In one example, w ip1 =1,shift=1. 4) In one example, w ip1 =1 / 2 / 3, shift=2. 5) In one example, w ip1 =1 / 2 / 3 / 4 / 5 / 6 / 7, shift=3. 6) In one example, w ip1=1 / 2 / 3 / 4 / 5 / 6 / 7 / 8 / 9 / 10 / 11 / 12 / 13 / 14 / 15, shift=4. 7) In one example, shift = 5 / 6 / 7 / 8. iii. In one example, when using IBC AMVP mode, w ip1 =1, shift=1, or w ip1 = 2, shift=2. iv. In one example, when using IBCMercure mode, w ip1 =3, shift=4. e. In one example, the weighting parameters may be different for IBC AMVP mode and IBC Merge mode. i. Alternatively, the weighting parameters may be the same for IBC AMVP mode and IBC Merge mode. f. In one example, the weighting parameters may depend on the video content. i. In one example, the weighting parameters may be different for natural sequences and screen content sequences. 3. In one example, the reference area of CIBCIP can be smaller than or equal to the reference area of IBC. a. In one example, the reference area of CIBCIP may depend on the coding information of intra prediction. i. In one example, the reference region of CIBCIP may depend on the intra prediction mode. b. Alternatively, the reference region of CIBCIP may be different from that of IBC. 4. Whether and / or how the CIBCIP mode is applied to a video unit may depend on codec information, which may refer to: a. Whether IBC or intra prediction method is allowed b. Whether to use CU skip mode c. Block dimensions and / or block size i. In one example, a block is allowed to be encoded and decoded using CIBCIP when the block size (W×H) is less than or equal to a threshold (T), where W and H represent the block width and block height, respectively. 1) In one example, T=256, or 512, or 1024, or 2048, or 4096. 2) In one example, T may depend on whether IBC AMVP mode or IBC Merge mode is used. a) In one example, when using IBC AMVP mode, T may be set equal to T1. i. In one example, T1=256, or T1=512, T1=1024, T1= 2048, T1=4096. b) In one example, when using IBC Merge mode, T may be set equal to T2. i. In one example, T2=256, or T2=512, T2=1024, T2= 2048, T2=4096. 3) In one example, T may depend on the slice / picture type. ii. In one example, when the block size (W×H) is less than or equal to a threshold (T3), the block is allowed to be encoded and decoded using CIBCIP, where W and H represent the block width and block height, respectively. 1) In one example, T3=16 or 32 or 64 or 128 or 256. 2) In one example, T3 may depend on whether IBC AMVP mode or IBCMercure mode is used. a) In one example, when using IBC AMVP mode, T3 may be set equal to T4. i. In one example, T4=16, or T4=32, T4=64, T4=128, T4=256. b) In one example, when using IBC Merge mode, T may be set equal to T5. i. In one example, T5=16, or T5=32, T5=64, T5=128, T5=256. 3) In one example, T3 may depend on the slice / picture type. iii. In one example, when W is greater than or equal to a threshold (T4) and / or H is greater than or equal to a threshold (T5), the block is allowed to be encoded and decoded using CIBCIP. 1) In one example, T4=4 / 8 / 16 / 32. 2) In one example, T5 = 4 / 8 / 16 / 32. iv. In one example, when W is less than or equal to a threshold (T6) and / or H is less than or equal to a threshold (T7), the block is allowed to be encoded and decoded using CIBCIP. 1) In one example, T6=16 / 32 / 64. 2) In one example, T7=16 / 32 / 64. v. In one example, the block size may refer to the luma block size. d. Block depth e. Slice / picture type and / or partition tree type (single tree or dual tree or partial dual tree) i. In one example, CIBCIP may only be applied to I slices / pictures. 1) Alternatively, CIBCIP may be applied to all slice / picture types. ii. In one example, whether CIBCIP is applied to a particular slice / picture type may depend on whether IBC AMVP mode or IBC Merge mode is used. 1) In one example, when CIBCIP is applied to IBC AMVP mode, it is only applied to I slices / pictures. a) Alternatively, when CIBCIP is applied to IBC AMVP mode, it is applied to all slice / picture types. 2) In one example, when CIBCIP is applied to IBC Merge mode, it is only applied to I slices / pictures. a) Alternatively, when CIBCIP is applied to IBC Merge mode, it is applied to all slice / picture types. f. Time domain layer identification g. Block location h. Color component. 5. In one example, an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs in the list can be used for intra prediction of CIBCIP. a. In one example, one or more conventional IPMs (eg, 35 IPMs in HEVC or 67 IPMs in VVC) may be included in the list. i. In one example, one or more IPMs may be associated with whether a particular intra codec tool is used. 1) In one example, the specific intra-frame codec tool may refer to ISP or MRL. b. In one example, one or more MIP modes may be included in the list. c. In one example, some or all of the MPMs may be included in the list. i. In one example, an MPM can be in the primary MPM list and / or in the secondary MPM list. d. In one example, one or more derived IPMs may be included in a list. i. In one example, the derived IPM may use a DIMD or TIMD based approach. 1) In one example, one or more derived IPMs using TIMD may be added to a list. a) In one example, a first optimal TIMD mode may be added. b) In one example, a second optimal TIMD mode may be added. c) In one example, the top N (eg, N is an integer) best TIMDs may be added model. d) In one example, one or more DIMD patterns may be used to derive TIMD model. i. In one example, one or more DIMD modes can be used to construct an MPM list, and the MPM list is used to derive the TIMD mode. 2) In one example, one or more derived IPMs using DIMD may be added to a list. a) In one example, a first optimal DIMD mode may be added. b) In one example, a second optimal DIMD mode may be added. c) In one example, the top N (eg, N is an integer) best DIMD mode. e. In one example, the IPMs in the list may be reordered. i. In one example, templates of video units can be used for reordering. ii. In one example, a TIMD-based approach can be used for reordering. iii. In one example, one or more BVs may be used for reordering. f. In one example, one or more IPMs in the list before or after reordering may be replaced by one or more IPMs of the reference video unit located by the BV. i. In one example, the Nth IPM in the list may be replaced, such as N=0, or 1, or 2, or 3, or 4, or 5. 1) In one example, deduplication is used before replacement, such as if the IPM located by the BV is already in the list, it is not used to replace the existing one. ii. Alternatively, one or more IPMs of the reference video unit located by the BV may be added to the list. 1) In one example, deduplication is used when adding IPM. g. In one example, at least one IPM derived using the BV may be added to the IPM candidate list. i. In one example, when an IPM derived using BV is already in the IPM candidate list, the IPM may not be added to the IPM candidate list. ii. Alternatively, when the IPM derived using the BV is already in the IPM candidate list, the default IPM may be added to the IPM candidate list. h. In one example, at least one TIMD mode is used to construct an IPM candidate list. Let the size of the list be N. i. In one example, N=2. 1) In one example, the best TIMD mode is used as the first IPM in the list. 2) In one example, the IPM derived from the BV is used as the second IPM in the list. a) Alternatively, a default IPM is used as the second IPM in the list. i. In one example, the default IPM may refer to planar, DC, horizontal mode, or vertical mode. 3) In one example, when the second IPM is the same as the first IPM in the list, the second IPM may be replaced by another IPM. a) In one example, another IPM may refer to planar, DC, horizontal mode, or vertical mode. ii. In one example, N=3. 1) In one example, the best TIMD mode is used as the first IPM in the list. 2) In one example, the second best TIMD mode is used as the second IPM in the list. 3) In one example, the IPM derived from the BV is used as the second IPM or the third IPM in the list. a) Alternatively, one or more default IPMs are used as the second or third IPM in the list. IPM. i. In one example, the default IPM may refer to planar, DC, horizontal mode, or vertical mode. 4) In one example, when the second / third IPM is the same as the first / first two IPMs in the list, the second / third IPM may be replaced by another IPM. a) In one example, another IPM may refer to planar, DC, horizontal mode, or vertical mode. iii. In one example, one or more of the above TIMD modes may be replaced by a DIMD mode. iv. In one example, when DIMD and / or TIMD are not allowed to be used, the TIMD / DIMD mode may be replaced by the default IPM. i. In one example, the IPM candidate list constructed using the above method can be applied to the IBC AMVP mode and / or the IBC Merge mode. i. In one example, the IPM candidate list may be the same for IBC AMVP mode and IBC Merge mode. ii. Alternatively, the IPM candidate list may be different for IBC AMVP mode and IBC Merge mode. 6. In one example, the determination of generating an IBC prediction signal and / or generating an intra prediction signal and / or fusing IBC and intra prediction signals for CIBCIP may be derived. a. In one example, the determination to generate an IBC prediction signal may refer to a specific IBC encoding tool and / or a specific IBC Merge candidate type and / or AMVP / Merge index. i. In one example, the specific IBC codec tool may refer to RR-IBC or IBC-TM or IBC-MBVD or IBC-GPM or IBC-LIC. ii. In one example, the specific IBC Merge candidate type may refer to a normal Merge candidate, an IBC-TM Merge candidate, an IBC-MBVD candidate, or an RR-IBC candidate. b. In one example, the determination to generate an intra prediction signal may refer to a specific intra prediction method and / or a specific IPM, and / or an IPM index indicating which IPM to use. c. In one example, the determination of fusing IBC and intra prediction signals may refer to weighting parameters and / or the number of IBC and / or intra prediction signals used for fusion. d. In one example, a list of “CIBCIP candidates” may be constructed, and an index indicating one “CIBCIP candidate” of the list may be derived or signaled. i. In one example, a “CIBCIP candidate” may include determining to generate an IBC prediction signal and / or determining to generate an intra prediction signal and / or determining to fuse IBC and intra prediction signals. 7. In one example, the codec information used in CIBCIP can be used to encode and decode subsequent video units. a. In one example, the block vector used to generate the IBC prediction signal can be considered as the block vector of normal IBC. i. In one example, the block vector may be added to the IBC HMVP table. ii. In one example, the block vector may be used to construct an IBCAVP / Merge candidate list for subsequent video units. b. In one example, the IPM used to generate the intra prediction signal can be regarded as the IPM of normal intra prediction. i. In one example, the IPM can be used to construct the MPM list for subsequent video units. ii. In one example, IPM can be used for chroma prediction. iii. In one example, IPM can be propagated for non-intra-coded video units. c. Alternatively, the codec information may not be used for subsequent video units. i. In one example, instead of the IPM for the current video unit, a predefined IPM (eg, DC or Planar) may be used for subsequent video units. 8. In one example, CIBCIP may not be allowed to be used with one or more specific codecs. a. In one example, a specific codec tool may refer to IBC AMVP mode, or IBC Merge mode, or IBC-TM mode, or IBC-MBVD mode, or RR-IBC mode, or IBC-LIC, or IBC-GPM, or AMVR for IBC. i. In one example, CIBCIP may not be allowed to be used with IBC-LIC or RR-IBC or IBC-GPM. b. Alternatively, CIBCIP can be used with one or more of the above encoding tools. i. In one example, CIBCIP may be allowed to be used with IBC AMVP mode or IBC Merge mode or IBC-TM mode or IBC-MBVD mode or IBC's AMVR. 9. In one example, whether and / or how CIBCIP is applied may depend on the color format and / or color components. a. In one example, CIBCIP can be applied to all color components. b. In one example, when CIBCIP is applied to chroma components, the derivation of intra prediction may be different from that of luma components. i. In one example, CCLM, or MMLM, or CCCM, or Chroma-DIMD, or Chroma-TIMD, or a fusion of CCLM / MMLM / CCCM with angular mode can be used to obtain intra prediction. c. In one example, whether and / or how CIBCIP is applied to the first component may depend on whether CIBCIP is applied to the second component. i. In one example, the first component may refer to a chrominance component (eg, Cb and / or Cr), and the second component may refer to a luma component (eg, Y). ii. In one example, CIBCIP may be applied to the first component in the same manner as the second component. 1) Alternatively, the way CIBCIP is applied to the first component may be different from the second component. a) In one example, the weighting parameters may be different. d. In one example, CIBCIP may be applied to the luma component but not to the chroma components. i. In one example, the luma component may refer to Y in the YCbCr color space or G in the RGB color space. ii. In one example, the chroma components may refer to Cb and / or Cr in the YCbCr color space, or R and / or B in the RGB color space. Signaling of a combination of intra block copy and intra prediction 10. Indications of CIBCIP patterns can be derived instantly. 11. The indication of CIBCIP mode may be conditionally signaled, where the condition may include: a. Whether IBC or intra prediction method is allowed b. Block dimensions and / or block size i. In one example, when the block size (W×H) is less than or equal to a threshold (T), the indication of CIBCIP mode may not be signaled, where W and H represent block width and block height, respectively. 1) In one example, T=256, or 512, or 1024, or 2048, or 4096. 2) In one example, T may depend on whether IBC AMVP mode or IBC Merge mode is used. a) In one example, when using IBC AMVP mode, T may be set equal to T1. i. In one example, T1=256, or T1=512, T1=1024, T1= 2048, T1=4096. b) In one example, when using IBC Merge mode, T may be set equal to T2. i. In one example, T2=256, or T2=512, T2=1024, T2= 2048, T2=4096. 3) In one example, T may depend on the slice / picture type. ii. In one example, when the block size (W×H) is less than or equal to a threshold ( T3 ), the indication of the CIBCIP mode may not be signaled, where W and H represent the block width and block height, respectively. 1) In one example, T3=16 or 32 or 64 or 128 or 256. 2) In one example, T3 may depend on whether IBC AMVP mode or IBCMercure mode is used. a) In one example, when using IBC AMVP mode, T3 may be set equal to T4. i. In one example, T4=16, or T4=32, T4=64, T4=128, T4=256. b) In one example, when using IBC Merge mode, T may be set equal to T5. i. In one example, T5=16, or T5=32, T5=64, T5=128, T5=256. 3) In one example, T3 may depend on the slice / picture type. iii. In one example, when W is greater than or equal to a threshold (T6) and / or H is greater than or equal to a threshold (T7), an indication of CIBCIP may be signaled. 1) In one example, T6 = 4 / 8 / 16 / 32. 2) In one example, T7=4 / 8 / 16 / 32. iv. In one example, when W is less than or equal to a threshold (T8) and / or H is less than or equal to a threshold (T9), an indication of CIBCIP may be signaled. 1) In one example, T8=16 / 32 / 64. 2) In one example, T9=16 / 32 / 64. v. In one example, the block size may refer to the luma block size. c. Block depth d. Slice / picture type and / or partition tree type (single tree or dual tree or partial dual tree) i. In one example, the indication of CIBCIP mode may be signaled for only I slices / pictures. 1) Alternatively, the indication of CIBCIP mode may be signaled for all slice / picture types. ii. In one example, whether an indication of CIBCIP mode is signaled for a particular slice / picture type may depend on whether IBC AMVP mode or IBC Merge mode is used. 1) In one example, when CIBCIP is applied to IBC AMVP mode, an indication of the CIBCIP mode is signaled for I slice / picture. a) Alternatively, when CIBCIP is applied to IBC AMVP mode, an indication of the CIBCIP mode is signaled for all slice / picture types. 2) In one example, when CIBCIP is applied to IBC Merge mode, an indication of the CIBCIP mode is signaled for I slice / picture. a) Alternatively, when CIBCIP is applied to IBC Merge mode, an indication of the CIBCIP mode is signaled for all slice / picture types. e. Time domain layer identification f. Block location i. In one example, the indication of CIBCIP mode is not signaled for blocks positioned at the top left of a slice / picture. g. Color component. h. In one example, if an indication of CIBCIP mode is not signaled, it may be inferred as a default value. i. In one example, if an indication of CIBCIP mode is not signaled, it can be presumed to be false. i. In one example, if an indication of CIBCIP mode is not signaled, it may be presumed to be true. 12. One or more syntax elements may be used to signal whether the current block is coded using CIBCIP mode. a. In one example, the syntax element may be binarized using fixed length codec or truncated unary codec or unary codec or EG codec or a coded flag. b. In one example, syntax elements may be bypass coded or context coded. i. The context may depend on coded information such as block dimensions and / or block size, and / or slice / Picture type, and / or information of neighboring blocks (adjacent or non-adjacent), and / or information of other coding tools used for the current block, and / or information of the temporal layer. ii. In one example, syntax elements may use different contexts when encoded for IBC AMVP mode and IBC Merge mode. 1) Alternatively, when encoding and decoding a syntax element for both IBC AMVP mode and IBC Merge mode, the syntax element may use the same context. c. In one example, when the current video unit is encoded and decoded by IBC, an indication of CIBCIP mode may be signaled. d. In one example, when CU skipping is used, the indication of CIBCIP mode may not be signaled. i. Alternatively, when CU skipping is used, an indication of CIBCIP mode may be signaled. e. In one example, it is possible to operate in IBC-TM mode or IBC-MBVD mode or RR-IBC mode or Before or after the indication of IBC-LIC or IBC-GPM, syntax elements are transmitted through signals. i. In one example, whether and / or how syntax elements are signaled may depend on the IBC mode or IBC-TM mode or IBC-MBVD mode or IBC-IBC mode or IBC-LIC or IBC- Whether GPM is enabled for video units. f. In one example, one or more syntax elements may be in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header are transmitted through signals. g. In one example, syntax elements may be coded in a predictive manner. h. For example, the syntax elements of the current block can be predicted from the syntax elements of neighboring blocks. 13. In one example, for the above project, RR-IBC or symmetric IBC methods can be used in CIBCIP. 14. In one example, for the above project, the RR-IBC or symmetric IBC method can be disabled in the CIBCIP. 15. In one example, the flip type of the IBC prediction portion may be set to NO_FLIP (eg, 0). 16. When the current block is encoded and decoded using the CIBCIP mode, which one or more IPMs in the IPM candidate list to use to obtain the intra prediction signal can be signaled using one or more syntax elements. a. In one example, the syntax element may be binarized using fixed length codec or truncated unary codec or unary codec or EG codec or codec flag. b. In one example, syntax elements may be bypass coded or context coded. i. The context may depend on coded information such as block dimensions and / or block size, and / or slice / Picture type, and / or information of neighboring blocks (adjacent or non-adjacent), and / or information of other coding tools used for the current block, and / or information of the temporal layer. ii. In one example, syntax elements may use different contexts when encoding and decoding for IBC AMVP mode and IBC Merge mode. 1) Alternatively, when encoding and decoding a syntax element for both IBC AMVP mode and IBC Merge mode, the syntax element may use the same context. Intra prediction using fused reference lines 17. It is proposed to fuse more than one reference line before using it to derive intra prediction for a video unit. a. In one example, the number of reference lines (N) and the reference lines to be merged can be predefined, signaled in the bitstream, or derived on the fly, where N is an integer greater than 1. i. In one example, N may be predefined, such as N=2 or N=3. ii. In one example, N may be signaled in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / slice group header. iii. In one example, N may be determined based on codec information. 1) In one example, the codec information may refer to block size, block dimension, or block position, or codec mode, or intra prediction mode. iv. In one example, which reference rows to use in the fusion may be indicated by a reference row index. 1) The reference row index may be predefined, signaled in the bitstream, or derived on the fly. 2) In one example, one of the reference row indices may be predefined, and the other remaining N-1 reference row indices may be signaled. 3) In one example, one of the reference row indices may be signaled, and the other remaining N-1 reference row indices may be predefined or derived on the fly. b. In one example, a weighting parameter may be used to fuse N reference rows. i. In one example, L = W1*L1+W2*L2+…+W N-1 *L N , where L i and W i denotes the i-th reference row used in fusion and the corresponding weight, and L denotes the fused reference row for intra prediction. 1) In another example, L = (W'1*L1+W'2*L2+...+W' N-1 *L N )>>Shift1, where W'1+W'2+…+W' N-1 =2^Shift1. ii. In one example, the weighting parameters may be predefined, or signaled in the bitstream, or derived on the fly. iii. In one example, when reference line L a Compared to reference line L b When it is closer to the current video unit, L a The corresponding weight parameter W a or W' a Can be equal to or greater than L b W b or W' b . iv. In one example, when N=2, W1=3 / 4 and W2=1 / 4, or W1=5 / 8 and W2= 3 / 8, or W1=1 / 2 and W2=1 / 2. v. In one example, when N=3, W1=1 / 2 and W2=1 / 4 and W3=1 / 4, or W1= 5 / 8 and W2=1 / 4 and W3=1 / 8. c. In one example, the number of samples in one reference row can be the same as the number of samples in another reference row. The fusion of samples is expressed as follows: P(x,y)=W1*P(x1,y1)1+W2*P(x2,y2)2+...+ W N-1 *P(x N-1 ,y N-1 ) N-1 , where P(x i ,y i ) i Represents a sample point in the i-th reference row. i. In one example, when fusing the upper portion of a reference row, samples in different reference rows with the same horizontal position can be fused. 1) In one example, x1=x2=…=x N-1 . ii. In one example, when blending the left portion of a reference line, samples in different reference lines with the same vertical position can be fused. 1) In one example, y1=y2=…=y N-1 . d. In one example, the number of samples in one reference row may be different from the number of samples in another reference row. m and reference line L n The number of samples in is denoted as S m and S n The example is shown as Figure 50 . i. In one example, the number of samples in the fused reference row may be the same as the number of samples of the reference row with the least number of samples. 1) In one example, in order to fuse the samples of the fused reference line, the samples of the reference line L can be m Use the reference line L n Many samples, among which L m The total number of samples in is greater than L n . a) In one example, it is possible to m Two or more sample points are used in L n Use one sample point in . 2) In one example, reference line L m S in n The sample points can be compared with the reference line L n S in n The example is depicted as Figure 51 . ii. In one example, the number of samples in the fused reference row may be the same as the number of samples of the reference row with the largest number of samples. 1) In one example, when S n Less than S m When (S m -S n ) sample points can use the reference line S n The samples in are filled or derived and used for fusion. The example is depicted as Figure 52 . e. In one example, fusion of reference rows may be performed after each reference row is derived. i. Alternatively, the fusion of the reference lines can be performed during their derivation. f. In one example, the derivation of reference samples used in the fused reference row can be the same as the derivation of reference samples not used for fusion of the reference row. i. Alternatively, the derivation may be different. 1) In one example, how to handle unavailable reference samples can be different. g. In one example, reference sample filtering may be performed after blending of reference lines. i. Alternatively, reference sample filtering can be performed before fusion of reference lines. 1) In one example, the reference sample filtering may be different for different reference rows. 18. Whether and how to use the fused reference lines to derive intra prediction for the current video unit may depend on codec information. a. In one example, the codec information may refer to one or more intra prediction methods. i. In one example, the fused reference rows can be used for conventional intra prediction. ii. In one example, the fused reference line can be used in MRL / ISP / MIP / DIMD / TIMD. iii. In one example, the fused reference rows can be used for conventional chroma intra prediction. iv. In one example, the fused reference lines can be used for fusion of LM and angle of chrominance. v. In one example, the fused reference lines can be used in addition to or as a replacement for the current intra prediction method. b. In one example, the codec information may refer to a color component. i. In one example, the fused reference rows can be used for intra prediction of the luma component. ii. In one example, the fused reference lines can be used for intra prediction of chroma components. c. In one example, the codec information may refer to an intra prediction mode. i. In one example, when using DC mode, a fused reference line may be used. ii. In one example, when using planar mode, a fused reference line may be used. iii. In one example, when angular intra prediction mode is used, fused reference lines may be used. iv. In one example, when angular intra prediction mode has a non-integer slope, fused reference lines may be used. v. In one example, the fused reference row can be used for more than one intra prediction mode. d. In one example, the codec information may refer to the block size / dimension of the current block and / or neighboring blocks. i. In one example, when the block size of the current block is greater than and equal to T1, the fused reference row may be used. ii. In another example, when the block size of the current block is smaller than T2, the fused reference rows may be used. e. In one example, the codec information may refer to a slice type and / or a temporal layer and / or a QP. f. In one example, fused reference rows may not be allowed for video units in different CTUs. 19. In one example, how reference lines are fused, and whether and how the fused reference lines are used to derive intra prediction for the current video unit can be signaled in the bitstream. General aspects 20. In the above examples, video unit can refer to color component / sub-picture / slice / slice / codec tree unit (CTU) / CTU row / CTU group / Codec unit (CU) / Prediction unit (PU) / Transform unit (TU) / Codec tree block (CTB) / Codec tree block (CTB) / Codec block (CB) / Prediction block (PB) / Transform block (TB) / block / subblock of a block / subregion within a block / or any other region containing multiple samples or pixels. 21. Whether and / or how to apply the method disclosed above can be transmitted through signals at the sequence level / picture level group / picture level / slice level / slice group level, for example in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / PPS / slice header / slice group header. 22. Whether and / or how the above methods are applied may depend on the following information: a. Messages transmitted by signal in DPS / SPS / VPS / PPS / APS / picture header / slice header / slice group header / largest codec unit (LCU) / coding unit (CU) / LCU line / LCU group / LCU block / video coding unit b. Location of CU / PU / TU / block / video codec unit c. Block dimensions of the current block and / or its neighboring blocks d. Block shape of the current block and / or its neighboring blocks e. Block codec mode, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode f. Indication of color format (e.g. 4:2:0, 4:4:4) g. Codec tree structure h. Slice / slice group type and / or picture type i. Color components (e.g., may be applied only to chroma components or luma components) j. Time domain layer ID k. The grade, level or layer of the standard 5. Example Embodiments 5.1 Example 1 In this contribution, three aspects of extending the use of IBC are proposed: Aspect #1: Combined IBC and intra prediction (IBC-CIIP); Aspect #2: IBC with geometric partitioning (IBC-GPM); Aspect #3: IBC with Local Illumination Compensation (IBC-LIC). Combined IBC and intra prediction (IBC-CIIP) When IBC-CIIP is applied to a CU, IBC and intra prediction are used to obtain two prediction signals. These two prediction signals are weighted and summed to generate the final prediction. IBC-CIIP can be applied in both IBC AMVP mode and IBC Merge mode. A CU flag is signaled to indicate the use of IBC-CIIP. IBC using geometric partitioning (IBC-GPM) When IBC GPM is applied to a CU, the CU is geometrically partitioned into two subpartitions. Prediction signals for the two subpartitions are generated using IBC and intra prediction. IBC GPM can be applied in IBC Merge mode. A CU flag is signaled to indicate the use of IBC GPM. IBC with Local Illumination Compensation (IBC-LIC) When the IBC LIC is applied to a CU, the local illumination variation between the CU and its prediction block is modeled as a linear equation. The parameters of the linear equation are derived in a similar manner to the LIC for inter-frame prediction. The IBC LIC can be applied to both IBC AMVP mode and IBC Merge mode. For IBC AMVP mode, an IBC LIC flag is signaled to indicate the use of the IBC LIC. For IBC Merge mode, the IBC-LIC flag is inferred from the Merge candidate. 5.2 Example 2 Combined IBC and intra prediction (IBC-CIIP) is a codec tool for CU that uses IBC and intra prediction to obtain two prediction signals, and the two prediction signals are weighted and summed to generate the final prediction as follows: P=(w ibc *P ibc +((1<<shift)-w ibc )*P intra +1<(shift-1))>>shift Among them, P ibc and P intra Represents the IBC prediction signal and intra-frame prediction signal respectively. For IBC Merge mode, w ibc and displacement are set equal to 13 and 14, and for IBC AMVP mode, they are set equal to 1 and 1. The intra prediction mode (IPM) candidate list is used to generate the intra prediction signal, and the IPM candidate list size is predefined as 2. The IPM index is signaled to indicate which IPM to use. The first IPM in the IPM candidate list is derived using TIMD, and the second IPM used in IBC Merge mode is derived using the block vector of the Merge candidate, which is used to generate the IBC prediction signal. In IBC AMVP mode, horizontal mode is used as the second IPM, and when the derived TIMD mode is horizontal mode, it is replaced by planar. IBC-CIIP is used for CUs with width * height >= 32 and maximum (width, height) <= 32. When IBC-CIIP is applied to IBC AMVP mode, the same method as the IBC AMVP mode is used to generate the IBC prediction signal. When RR-IBC is enabled, IBC-CIIP is not applied to the IBC AMVP mode. When IBC-CIIP is applied to IBC Merge mode, conventional IBC Merge mode, IBC TM Merge mode or IBC-MBVD mode is used to generate the IBC prediction signal. A CU flag is transmitted by signal for IBC AMVP mode and IBC Merge mode to indicate whether IBC-CIIP is applied. When CU skip is used for the current CU, IBC-CIIP is not applied. When IBC LIC or IBC GPM is enabled for the CU, IBC-CIIP is not applied. 5.3 Example 3 IBC with Local Illumination Compensation (IBC-LIC) is a codec tool that compensates for local illumination variations between a CU encoded with IBC and its prediction block within a picture using a linear equation. The parameters of the linear equation are the same as those for the LIC used for inter prediction, except that the reference template is generated using the block vector in the IBC LIC. The IBC LIC can be applied to IBC AMVP mode and IBC Merge mode. For the IBC AMVP mode, the IBC LIC flag is signaled to indicate the use of the IBC LIC. For the IBC Merge mode, the IBC-LIC flag is inferred from the Merge candidate. The IBC-LIC is also used during the ARMC process of the IBC Normal Merge mode and the IBC TM Merge mode, but is not used during the ARMC process of the IBC-MBVD mode. The IBC LIC is used for CUs with width*height>=32 and width*height<=256.
[0111] As used herein, the term "video unit" or "video block" may be a sequence, a picture, a slice, a tile, a sub-picture, a codec tree unit (CTU) / codec tree block (CTB), a CTU / CTB row, one or more codec units (CU) / codec blocks (CB), one or more CTUs / CTBs, one or more virtual pipeline data units (VPDUs), or a sub-region within a picture / slice / slice / tile. The term "reference row" may refer to row and / or column reconstruction samples that are adjacent or non-adjacent to a current block, and are used to derive intra prediction of the current video unit via an interpolation filter along a specific direction, and the specific direction is determined by an intra prediction mode (e.g., conventional intra prediction with an intra prediction mode), or by weighting reference samples of a reference row using a matrix or vector (e.g., MIP).
[0112] Figure 53 A flow chart of a method 5300 for video processing according to an embodiment of the present invention is shown. The method 5300 is implemented during conversion between a video unit of a video and a bitstream of the video.
[0113] At box 5310, for conversion between a video unit of a video and a bitstream of the video unit, it is determined whether a combined intra block copy (IBC) and intra prediction (CIBCIP) mode is applied to the video unit based on at least one of the following: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements.
[0114] At block 5320, a prediction for the video unit is derived by combining the IBC prediction signal and the intra prediction signal.
[0115] At block 5330, conversion is performed based on the prediction of the video unit. In some embodiments, the conversion may include encoding the video unit into a bitstream. Additionally or alternatively, the conversion may include decoding the video unit from the bitstream. In this way, codec efficiency and codec performance can be improved.
[0116] In some embodiments, the codec information includes at least one of the following: whether IBC or intra prediction mode is allowed, whether codec unit (CU) skip mode is used, block dimension and / or block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position, or color component.
[0117] In some embodiments, a video unit is allowed to be encoded and decoded using CIBCIP if the width of the video unit is greater than or equal to a first threshold and / or the height of the video unit is greater than or equal to a second threshold. In some embodiments, the first threshold is one of the following: 4, 8, 16, or 32. In some embodiments, the second threshold is one of the following: 4, 8, 16, or 32.
[0118] In some embodiments, if the width of the video unit is less than or equal to a third threshold and / or the height of the video unit is less than or equal to a fourth threshold, then the video unit is allowed to be encoded and decoded using CIBCIP. In some embodiments, the third threshold is one of the following: 16, 32, or 64. In some embodiments, the fourth threshold is one of the following: 16, 32, or 64.
[0119] In some embodiments, an intra-prediction mode (IPM) candidate list is constructed, and one or more IPMs in the IPM candidate list are used for intra-prediction of CIBCIP. In some embodiments, one or more derived IPMs using template-based intra-mode derivation (TIMD) are added to the IPM candidate list. In some embodiments, the top N best TIMD modes are added to the IPM candidate list, where N is an integer.
[0120] In some embodiments, one or more derived IPMs using chroma decoder side intra mode derivation (DIMD) are added to the IPM candidate list. In some embodiments, the top N best DIMD modes are added to the IPM candidate list, where N is an integer.
[0121] In some embodiments, the indication of CIBCIP mode is indicated based on a condition, where the condition includes at least one of the following: whether IBC or intra prediction mode is allowed, block dimension and / or block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position, or color component.
[0122] In some embodiments, if the width of the video unit is greater than or equal to a fifth threshold and / or the height of the video unit is greater than or equal to a sixth threshold, then the CIBCIP mode is indicated. In some embodiments, the fifth threshold is one of the following: 4, 8, 16, or 32. In some embodiments, the sixth threshold is one of the following: 4, 8, 16, or 32.
[0123] In some embodiments, if the width of the video unit is less than or equal to a seventh threshold and / or the height of the video unit is less than or equal to an eighth threshold, then the CIBCIP mode is indicated. In some embodiments, the seventh threshold is one of the following: 16, 32, 64. In some embodiments, the eighth threshold is one of the following: 16, 32, 64.
[0124] In some embodiments, if the indication of CIBCIP mode is not signaled, the indication of CIBCIP mode is presumed to be a default value. In some embodiments, if the indication of CIBCIP mode is not signaled, the indication of CIBCIP mode is presumed to be false. Alternatively, if the indication of CIBCIP mode is not signaled, the indication of CIBCIP mode is presumed to be true.
[0125] In some embodiments, if CU skipping is used, then the indication of CIBCIP mode is not signaled. Alternatively, if CU skipping is used, then the indication of CIBCIP mode is signaled.
[0126] In some embodiments, a video unit includes at least one of the following: a color component, a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec unit (CU), a codec tree unit (CTU), a CTU row, a CTU group, a slice, a sub-picture, a block, a sub-region within a block, or a region containing more than one sample or pixel.
[0127] In some embodiments, an indication of whether and / or how to derive prediction for a video unit by including an IBC prediction signal and an intra prediction signal is indicated at one of the following: a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level. In some embodiments, an indication of whether and / or how to derive prediction for a video unit by including an IBC prediction signal and an intra prediction signal is indicated in one of the following: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0128] In some embodiments, method 5300 further includes: determining whether and / or how to derive the prediction of the video unit by including an IBC prediction signal and an intra-frame prediction signal based on at least one of the following: a message indicated in one of a DPS, SPS, VPS, PPS, APS, a picture header, a slice header, a slice group header, a largest codec unit (LCU), a codec unit (CU), an LCU row, an LCU group, a TU, a PU block, a video codec unit, a position of one of a CU, a PU, a TU, a block, a video codec unit, a block dimension of a current block and / or a block dimension of a neighboring block of a current block, a block shape of a current block and / or a block shape of a neighboring block of a current block, a codec mode of the video unit; an indication of a color format, a codec tree structure slice type, a slice group type, a picture type, a color component, a temporal layer identifier, a profile or level or layer of a standard.
[0129] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video, and the bitstream of the video is generated by a method performed by an apparatus for video processing. The method includes: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to a video unit based on at least one of the following: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; and generating a bitstream based on the prediction of the video unit.
[0130] According to further embodiments of the present disclosure, a method for storing a video bitstream is provided, comprising: deriving a prediction of a video unit by combining an IBC prediction signal and an intra-frame prediction signal; generating a bitstream based on the prediction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0131] Embodiments of the present disclosure may be described according to the following clauses, the features of which may be combined in any reasonable way.
[0132] Item 1. A method of video processing, comprising: for conversion between a video unit of a video and a bitstream of the video unit, determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to the video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; and performing conversion based on the prediction for the video unit.
[0133] Item 2. A method according to Item 1, wherein the codec information includes at least one of the following: whether IBC or intra-frame prediction mode is allowed, whether codec unit (CU) skip mode is used, block dimension and / or block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position, or color component.
[0134] Item 3. The method of Item 2, wherein the video unit is allowed to be encoded using CIBCIP if a width of the video unit is greater than or equal to a first threshold and / or a height of the video unit is greater than or equal to a second threshold.
[0135] Item 4. The method of Item 3, wherein the first threshold is one of the following: 4, 8, 16, 32.
[0136] Item 5. The method of Item 3, wherein the second threshold is one of the following: 4, 8, 16, 32.
[0137] Item 6. The method of Item 2, wherein the video unit is allowed to be encoded using CIBCIP if a width of the video unit is less than or equal to a third threshold and / or a height of the video unit is less than or equal to a fourth threshold.
[0138] Item 7. The method of Item 6, wherein the third threshold is one of the following: 16, 32, 64.
[0139] Item 8. The method of Item 6, wherein the fourth threshold is one of the following: 16, 32, 64.
[0140] Item 9. The method of Item 1, wherein an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs in the IPM candidate list are used for intra prediction of CIBCIP.
[0141] Clause 10. The method of clause 9, wherein one or more derived IPMs using template-based intra mode derivation (TIMD) are added to the IPM candidate list.
[0142] Item 11. The method of Item 10, wherein the top N best TIMD modes are added to an IPM candidate list, where N is an integer.
[0143] Item 12. The method of Item 9, wherein one or more derived IPMs using chroma decoder side intra mode derivation (DIMD) are added to the IPM candidate list.
[0144] Item 13. The method of Item 12, wherein the top N best DIMD modes are added to an IPM candidate list, where N is an integer.
[0145] Item 14. A method according to Item 1, wherein the indication of the CIBCIP mode is indicated based on a condition, wherein the condition includes at least one of the following: whether IBC or intra-frame prediction mode is allowed, block dimension and / or block size, block depth, slice type, picture type, partition tree type, temporal layer identifier, block position, or color component.
[0146] Item 15. The method of Item 14, wherein the indication of the CIBCIP mode is indicated if the width of the video unit is greater than or equal to a fifth threshold and / or the height of the video unit is greater than or equal to a sixth threshold.
[0147] Item 16. The method of Item 15, wherein the fifth threshold is one of the following: 4, 8, 16, 32.
[0148] Item 17. The method of Item 15, wherein the sixth threshold is one of the following: 4, 8, 16, 32.
[0149] Item 18. The method of Item 14, wherein the indication of the CIBCIP mode is indicated if the width of the video unit is less than or equal to a seventh threshold and / or the height of the video unit is less than or equal to an eighth threshold.
[0150] Item 19. The method of Item 18, wherein the seventh threshold is one of the following: 16, 32, 64.
[0151] Item 20. The method of Item 18, wherein the eighth threshold is one of the following: 16, 32, 64.
[0152] Clause 21. The method of clause 1, wherein if the indication of the CIBCIP mode is not signaled, the indication of the CIBCIP mode is presumed to be a default value.
[0153] Clause 22. The method of clause 1, wherein if the indication of CIBCIP mode is not signaled, the indication of CIBCIP mode is presumed to be false.
[0154] Clause 23. The method of clause 1, wherein the indication of CIBCIP mode is presumed to be true if the indication of CIBCIP mode is not signaled.
[0155] Item 24. The method of Item 1, wherein if CU skipping is used, the indication of CIBCIP mode is not signaled.
[0156] Clause 25. The method of clause 1, wherein if CU skipping is used, an indication of CIBCIP mode is signaled.
[0157] Item 26. A method according to any one of Items 1 to 25, wherein the video unit comprises at least one of the following: a color component, a prediction block (PB), a transform block (TB), a codec block (CB), a prediction unit (PU), a transform unit (TU), a codec tree block (CTB), a codec unit (CU), a codec tree unit (CTU), a CTU row, a CTU group, a slice, a sub-picture, a block, a sub-region within a block, or a region containing more than one sample or pixel.
[0158] Item 27. A method according to any one of items 1 to 25, wherein an indication of whether and / or how to derive a prediction for a video unit by including an IBC prediction signal and an intra prediction signal is indicated at one of: a sequence level, a group of pictures level, a picture level, a slice level, or a slice group level.
[0159] Item 28. A method according to any one of items 1 to 25, wherein an indication of whether and / or how the prediction of the video unit is derived by including an IBC prediction signal and an intra prediction signal is indicated in one of the following items: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependency parameter set (DPS), decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a slice group header.
[0160] Item 29. The method according to any one of items 1 to 25 further includes: determining whether and / or how to derive the prediction of the video unit by including an IBC prediction signal and an intra-frame prediction signal is based on at least one of the following: a message indicated in one of DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest codec unit (LCU), codec unit (CU), LCU row, LCU group, TU, PU block, video codec unit, the position of one of the CU, PU, TU, block, video codec unit, the block dimensions of the current block and / or the block dimensions of the neighboring blocks of the current block, the block shape of the current block and / or the block shapes of the neighboring blocks of the current block, the codec mode of the video unit; an indication of the color format, the codec tree structure slice type, slice group type, picture type, color component, temporal layer identifier, standard grade or level or layer.
[0161] Item 30. The method of any one of Items 1 to 29, wherein converting comprises encoding the video unit into a bitstream.
[0162] Item 31. The method of any one of Items 1 to 29, wherein converting comprises decoding the video unit from a bitstream.
[0163] Item 32. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform a method according to any one of Items 1 to 31.
[0164] Item 33. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 31.
[0165] Item 34. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream of the video being generated by a method performed by an apparatus for video processing, wherein the method comprises: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to a video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; and generating a bitstream based on the prediction for the video unit.
[0166] Item 35. A method for storing a bitstream of a video, comprising: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to a video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction for the video unit by combining an IBC prediction signal and an intra prediction signal; generating a bitstream based on the prediction for the video unit; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0167] Figure 54 A block diagram of a computing device 5400 in which various embodiments of the present disclosure may be implemented is shown. The computing device 5400 may be implemented as a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300), or may be included in a source device 110 (or video encoder 114 or 200) or a destination device 120 (or video decoder 124 or 300).
[0168] It should be understood that Figure 54 The computing device 5400 shown in FIG. 5 is for illustrative purposes only and is not intended to in any way imply any limitation on the functionality and scope of the embodiments of the present disclosure.
[0169] like Figure 54 As shown, computing device 5400 comprises a general computing device 5400. Computing device 5400 may include at least one or more processors or processing units 5410, memory 5420, storage unit 5430, one or more communication units 5440, one or more input devices 5450, and one or more output devices 5460.
[0170] In some embodiments, the computing device 5400 can be implemented as any user terminal or server terminal with computing capabilities. The server terminal can be a server, a large computing device, etc. provided by a service provider. The user terminal can be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including a mobile phone, a station, a unit, a device, a multimedia computer, a multimedia tablet computer, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. It is conceivable that the computing device 5400 can support any type of interface to the user (such as a "wearable" circuit device, etc.).
[0171] The processing unit 5410 may be a physical processor or a virtual processor and may implement various processes based on a program stored in the memory 5420. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capability of the computing device 5400. The processing unit 5410 may also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0172] The computing device 5400 typically includes various computer storage media. Such media can be any media accessible by the computing device 5400, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 5420 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM) or flash memory), or any combination thereof. The storage unit 5430 can be any removable or non-removable medium and can include machine-readable media, such as memory, a flash drive, a disk, or other media that can be used to store information and / or data and can be accessed in the computing device 5400.
[0173] The computing device 5400 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Figure 54 Although not shown, a magnetic disk drive for reading from and / or writing to a removable nonvolatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable nonvolatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data medium interfaces.
[0174] The communication unit 5440 communicates with another computing device via a communication medium. In addition, the functions of the components in the computing device 5400 can be implemented by a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 5400 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0175] Input device 5450 may be one or more of various input devices, such as a mouse, keyboard, trackball, voice input device, and the like. Output device 5460 may be one or more of various output devices, such as a display, speaker, printer, and the like. With the aid of communication unit 5440, computing device 5400 may also communicate with one or more external devices (not shown), such as storage devices and display devices, one or more devices that enable a user to interact with computing device 5400, or, if desired, any device that enables computing device 5400 to communicate with one or more other computing devices (e.g., a network card, a modem, and the like). Such communication may be performed via an input / output (I / O) interface (not shown).
[0176] In some embodiments, some or all components of the computing device 5400 may also be arranged in a cloud computing architecture rather than being integrated into a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services without requiring the end user to know the physical location or configuration of the systems or hardware providing these services. In various embodiments, cloud computing provides services via a wide area network (such as the Internet) using appropriate protocols. For example, a cloud computing provider provides an application via a wide area network that can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data may be stored on servers in a remote location. Computing resources in a cloud computing environment may be consolidated or distributed across remote data centers. Cloud computing infrastructure can provide services through shared data centers, although to users, they appear as a single access point. Therefore, cloud computing architecture can be used to provide the components and functionality described herein from a service provider in a remote location. Alternatively, the components and functionality described herein may be provided by a conventional server or installed directly or otherwise on a client device.
[0177] In an embodiment of the present disclosure, the computing device 5400 may be used to implement video encoding / decoding. The memory 5420 may include one or more video encoding / decoding modules 5425 having one or more program instructions. These modules are accessible and executable by the processing unit 5410 to perform the functions of the various embodiments described herein.
[0178] In an example embodiment performing video encoding, an input device 5450 may receive video data as input to be encoded 5470. The video data may be processed, for example, by a video codec module 5425 to generate an encoded bitstream. The encoded bitstream may be provided as output 5480 via an output device 5460.
[0179] In an example embodiment performing video decoding, an input device 5450 may receive an encoded bitstream as input 5470. The encoded bitstream may be processed, for example, by a video codec module 5425 to generate decoded video data. The decoded video data may be provided as output 5480 via an output device 5460.
[0180] Although the present disclosure has been specifically shown and described with reference to the preferred embodiments of the present disclosure, it will be understood by those skilled in the art that various changes in form and details may be made without departing from the spirit and scope of the present application as defined by the appended claims. Such variations are intended to be encompassed by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A video processing method, comprising: For conversion between a video unit of a video and a bitstream of the video unit, determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to the video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; as well as The conversion is performed based on the prediction of the video unit.
2. The method according to claim 1, wherein the codec information includes at least one of the following: Whether IBC or intra prediction is allowed, Whether the codec unit (CU) skip mode is used, block dimensions and / or block size, Block depth, Strip type, Image type, Split tree type, Time domain layer identification, block location, or Color component. 3 . The method of claim 2 , wherein if a width of the video unit is greater than or equal to a first threshold and / or a height of the video unit is greater than or equal to a second threshold, the video unit is allowed to be encoded and decoded using the CIBCIP. The method according to claim 3 , wherein the first threshold is one of the following: 4, 8, 16, 32. The method according to claim 3 , wherein the second threshold is one of the following: 4, 8, 16, 32. 6 . The method of claim 2 , wherein if a width of the video unit is less than or equal to a third threshold and / or a height of the video unit is less than or equal to a fourth threshold, the video unit is allowed to be encoded and decoded using the CIBCIP. The method according to claim 6 , wherein the third threshold is one of the following: 16, 32, 64. The method according to claim 6 , wherein the fourth threshold is one of the following: 16, 32, 64.
9. The method of claim 1, wherein an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs in the IPM candidate list are used for intra prediction of the CIBCIP.
10. The method of claim 9, wherein one or more derived IPMs using template-based intra mode derivation (TIMD) are added to the IPM candidate list. The method of claim 10 , wherein the top N best TIMD modes are added to the IPM candidate list, where N is an integer.
12. The method of claim 9, wherein one or more derived IPMs using chroma decoder side intra mode derivation (DIMD) are added to the IPM candidate list.
13. The method of claim 12, wherein the top N best DIMD modes are added to the IPM candidate list, where N is an integer.
14. The method of claim 1 , wherein the indication of the CIBCIP mode is indicated based on a condition, wherein the condition comprises at least one of the following: Whether IBC or intra prediction is allowed, block dimensions and / or block size, Block depth, Strip type, Image type, Split tree type, Time domain layer identification, block location, or Color component.
15. The method of claim 14, wherein the indication of the CIBCIP mode is indicated if a width of the video unit is greater than or equal to a fifth threshold and / or a height of the video unit is greater than or equal to a sixth threshold. The method according to claim 15 , wherein the fifth threshold is one of the following: 4, 8, 16, 32. The method according to claim 15 , wherein the sixth threshold is one of the following: 4, 8, 16, 32.
18. The method of claim 14, wherein the indication of the CIBCIP mode is indicated if a width of the video unit is less than or equal to a seventh threshold and / or a height of the video unit is less than or equal to an eighth threshold. The method of claim 18 , wherein the seventh threshold is one of the following: 16, 32, 64.
20. The method of claim 18, wherein the eighth threshold is one of the following: 16, 32, 64.
21. The method of claim 1, wherein if the indication of the CIBCIP mode is not signaled, the indication of the CIBCIP mode is inferred to be a default value.
22. The method of claim 1, wherein if the indication of the CIBCIP mode is not signaled, then the indication of the CIBCIP mode is presumed to be false.
23. The method of claim 1, wherein the indication of the CIBCIP mode is presumed to be true if the indication of the CIBCIP mode is not signaled.
24. The method of claim 1, wherein if CU skipping is used, the indication of the CIBCIP mode is not signaled.
25. The method of claim 1, wherein the indication of the CIBCIP mode is signaled if CU skipping is used.
26. The method according to any one of claims 1 to 25, wherein the video unit comprises at least one of the following: Color component, Prediction Block (PB), Transform Block (TB), Codec Block (CB), Prediction Unit (PU), Transformation Unit (TU), Codec Tree Block (CTB) Codec Unit (CU), Codec Tree Unit (CTU), CTU line, CTU group, strips, piece, sub-images, piece, a sub-region within a block, or An area containing more than one sample or pixel.
27. The method of any one of claims 1 to 25, wherein the indication of whether and / or how to derive the prediction for the video unit by including the IBC prediction signal and the intra prediction signal is indicated at one of: Sequence level, Picture group level, Picture level, Stripe level, or Film group level.
28. The method of any one of claims 1 to 25, wherein the indication of whether and / or how to derive the prediction for the video unit by including the IBC prediction signal and the intra prediction signal is indicated in one of: Sequence header, Picture header, Sequence Parameter Set (SPS), Video Parameter Set (VPS), Dependent Parameter Set (DPS), Decoding Capability Information (DCI), Picture Parameter Set (PPS), Adaptive Parameter Set (APS), Strip header, or Film group header.
29. The method according to any one of claims 1 to 25, further comprising: Determining whether and / or how to derive the prediction for the video unit by including the IBC prediction signal and the intra prediction signal is based on at least one of: Message indicated in one of DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest codec unit (LCU), codec unit (CU), LCU row, LCU group, TU, PU block, video codec unit The position of one of the CU, PU, TU, block, video codec unit, the block dimensions of the current block and / or the block dimensions of the neighboring blocks of the current block, the block shape of the current block and / or the block shapes of the neighboring blocks of the current block, The encoding and decoding mode of the video unit; Indication of color format, Codec tree structure Strip type, Chipset type, Image type, Color component, Time domain layer identification, A grade, level, or tier of standard.
30. The method of any one of claims 1 to 29, wherein the converting comprises encoding the video unit into the bitstream.
31. The method of any one of claims 1 to 29, wherein the converting comprises decoding the video unit from the bitstream.
32. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 31.
33. A non-transitory computer-readable storage medium storing instructions, wherein the instructions cause a processor to execute the method according to any one of claims 1 to 31.
34. A non-transitory computer-readable recording medium storing a bit stream of a video, wherein the bit stream of the video is generated by a method performed by an apparatus for video processing, wherein the method comprises: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to the video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; as well as The bitstream is generated based on the prediction of the video unit.
35. A method for storing a bitstream of a video, comprising: determining whether to apply a combined intra block copy (IBC) and intra prediction (CIBCIP) mode to the video unit based on at least one of: codec information, whether an indication of CIBCIP is indicated, or one or more syntax elements; deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; generating the bitstream based on the prediction of the video unit; as well as The bitstream is stored in a non-transitory computer-readable recording medium.