Method and device for video processing and medium
By determining the combination of intra-block copying and intra-prediction in the video unit and deriving based on specific information of the video unit, the problem of insufficient video encoding and decoding efficiency and performance in the prior art is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202380072646.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-12
- Filing Date
- 2023-10-11
- Publication Date
- 2025-05-30
AI Technical Summary
The existing video encoding and decoding technology has shortcomings in improving the encoding and decoding efficiency and performance, especially when processing complex video content, it is difficult to effectively utilize the combination of intra-block copying and intra-prediction.
A method is proposed to improve the prediction efficiency of the video unit by determining whether to apply the combination of intra-block copying (IBC) and intra-prediction (CIBCIP) in the video unit and deriving based on the codec information, color format, color components or syntax elements of the video unit.
By combining IBC prediction signals and intra prediction signals, this method significantly improves the efficiency and performance of video encoding and decoding, and can process complex video content more effectively.
Smart Images

Figure CN120077652A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to combined intra-block copy and intra prediction. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is generally a desire to further improve the encoding / decoding efficiency of video encoding / decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: determining whether to apply a combination of intra-block copy (IBC) and intra prediction (CIBCIP) to a video unit based on at least one of the following for the conversion between the video unit of a video and the bitstream of the video unit: the encoding / decoding information, color format, color component, or syntax element of the video unit; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; and performing the conversion based on the prediction of the video unit. This can improve the encoding / decoding efficiency and encoding / decoding performance.
[0005] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to execute the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing of a video. The method includes: determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of the following: codec information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; and generating a bitstream based on the prediction of the video unit.
[0008] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes: determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of the following: codec information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; generating a bitstream based on the prediction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] The present invention content is provided to introduce a selection of concepts further described below in the detailed description in a simplified form. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the exemplary embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0011] Figure 1 A block diagram showing an exemplary video codec system according to some embodiments of the present disclosure is shown;
[0012] Figure 2 A block diagram showing a first exemplary video encoder according to some embodiments of the present disclosure is shown;
[0013] Figure 3 A block diagram showing an exemplary video decoder according to some embodiments of the present disclosure is shown;
[0014] Figure 4 An example of an encoder block diagram is shown;
[0015] Figure 5 Sixty-seven intra prediction modes are shown;
[0016] Figure 6 Shows reference samples for wide-angle intra prediction;
[0017] Figure 7 Shows the problem of discontinuity in the case of directions exceeding 45°;
[0018] Figure 8 Shows MMVD search points;
[0019] Figure 9 Is an illustration for symmetric MVD mode;
[0020] Figure 10 Shows the extended CU region used in BDOF;
[0021] Figure 11 Shows an affine motion model based on control points;
[0022] Figure 12 Shows the affine MVF for each sub-block;
[0023] Figure 13 Shows the position of the inherited affine motion prediction value;
[0024] Figure 14 Shows control point motion vector inheritance;
[0025] Figure 15 Shows the positioning of candidate positions for constructing the affine Merge mode;
[0026] Figure 16 Is an illustration of the motion vector usage of the proposed combination method;
[0027] Figure 17 Shows sub-block MV VSB and pixel Δv(i,j);
[0028] Figure 18A Shows the spatial neighboring blocks used by ATVMP;
[0029] Figure 18B Shows the derivation of the sub-CU motion field by applying the motion displacement from the spatial neighbors and scaling the motion information from the corresponding co-located sub-CU;
[0030] Figure 19 Shows local illumination compensation;
[0031] Figure 20 Shows not performing subsampling for the short side;
[0032] Figure 21 Shows decoder-side motion vector refinement;
[0033] Figure 22A diamond-shaped area in the search area is shown;
[0034] Figure 23 The location of the spatial Merge candidate is shown;
[0035] Figure 24 shows candidate pairs considered for redundancy check of spatial merge candidates;
[0036] Figure 25 It is a diagram of motion vector scaling for temporal Merge candidates;
[0037] Figure 26 The candidate positions for the time domain Merge candidates C0 and C1 are shown;
[0038] Figure 27 The VVC spatial domain neighboring blocks of the current block are shown;
[0039] Figure 28 is a diagram of the virtual blocks in the i-th search round;
[0040] Figure 29 An example of GPM partitions grouped at the same angle is shown;
[0041] Figure 30 Unidirectional prediction MV selection for geometric partitioning mode is shown;
[0042] Figure 31 shows the blending weights w using the geometric partitioning mode 0 Example generation of;
[0043] Figure 32 The spatial neighboring blocks used to derive spatial Merge candidates are shown;
[0044] Figure 33 It shows that template matching is performed on the search area around the initial MV;
[0045] Figure 34 is a diagram of the sub-blocks of the OBMC application;
[0046] Figure 35 The SBT position, type and transformation type are shown;
[0047] Figure 36 The neighboring sample points used to calculate the SAD are shown;
[0048] Figure 37 shows neighboring samples used to calculate SAD for sub-CU level motion information;
[0049] Figure 38 The classification process is shown;
[0050] Figure 39 Shows the reordering process in the encoder;
[0051] Figure 40 Shows the reordering process in the decoder;
[0052] Figure 41 Is an illustration of the extended reference area;
[0053] Figure 42 Shows the IBC reference area depending on the current CU position;
[0054] Figure 43 Shows an example of symmetry in a screen content picture;
[0055] Figure 44A Is an illustration of the BV adjustment for horizontal flipping;
[0056] Figure 44B Is an illustration of the BV adjustment for vertical flipping;
[0057] Figure 45 Shows the intra-template matching search area used;
[0058] Figure 46 Shows an example of different numbers of samples in different reference rows for fusion;
[0059] Figure 47 Shows an example of different numbers of samples in different reference rows for fusion, and the samples surrounded by the square box are discarded and not used for the fusion of the reference row;
[0060] Figure 48 Shows an example of different numbers of samples in different reference rows for fusion, and the samples represented by the blank circles in the reference Ln are filled and used for the fusion of the reference row;
[0061] Figure 49 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and
[0062] Figure 50 Shows a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.
[0063] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0064] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways in addition to the ways described below.
[0065] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0066] As used herein, the terms "one embodiment", "an embodiment", "example embodiment", etc. indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0067] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be termed a second element, and similarly, a second element may be termed a first element, without departing from the scope of the example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0068] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the example embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof.
[0069] Example Environment
[0070] Figure 1FIG. 0 is a block diagram showing an example video codec system 100 that may utilize the techniques of the present disclosure. As shown, video codec system 100 may include a source device 110 and a destination device 120. Source device 110 may also be referred to as a video encoding device, and destination device 120 may also be referred to as a video decoding device. In operation, source device 110 may be configured to generate encoded video data, and destination device 120 may be configured to decode the encoded video data generated by source device 110. Source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0071] Video source 112 may include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or a combination thereof.
[0072] The video data may include one or more pictures. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form an encoded representation of the video data. The bitstream may include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to destination device 120 via I / O interface 116 over network 130A. The encoded video data may also be stored on storage medium / server 130B for access by destination device 120.
[0073] Destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 may include a receiver and / or a modulator. I / O interface 126 may obtain the encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120 or may be external to destination device 120, which is configured to interface with an external display device.
[0074] Video encoder 114 and video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0075] Figure 2is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 can be Figure 1 an example of the video encoder 114 in the system 100 shown.
[0076] The video encoder 200 can be configured to implement any or all of the techniques of the present disclosure. In Figure 2 an example, the video encoder 200 includes multiple functional components. The techniques described in the present disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to execute any or all of the techniques described in the present disclosure.
[0077] In some embodiments, the video encoder 200 can include a splitting unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 can include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0078] In other examples, the video encoder 200 can include more, fewer, or different functional components. In one example, the prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0079] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) can be integrated, for the purpose of explanation, these components are shown separately in Figure 2 an example.
[0080] The splitting unit 201 can split a picture into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.
[0081] The mode selection unit 203 can select one coding mode (intra coding or inter coding) from a variety of coding modes, for example, based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra and inter prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0082] To perform inter - frame prediction on a current video block, the motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from the cache 213 with the current video block. The motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and the decoded samples of pictures from the cache 213 other than the picture associated with the current video block.
[0083] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, e.g., depending on whether the current video block is in an I - slice, a P - slice, or a B - slice. As used herein, an "I - slice" can refer to a part of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P - slice" and a "B - slice" can refer to parts of a picture composed of macroblocks independent of the macroblocks in the same picture.
[0084] In some examples, the motion estimation unit 204 can perform uni - directional prediction on the current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0085] Alternatively, in other examples, the motion estimation unit 204 can perform bi - directional prediction on the current video block. The motion estimation unit 204 can search the reference pictures in list 0 to find one reference video block for the current video block, and can also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 can then generate multiple reference indices and multiple motion vectors, where the multiple reference indices indicate the multiple reference pictures in list 0 and list 1 that contain the multiple reference video blocks, and the multiple motion vectors indicate the multiple spatial displacements between the multiple reference video blocks and the current video block. The motion estimation unit 204 can output the multiple reference indices and the multiple motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information of the current video block.
[0086] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.
[0087] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block, the value indicating to the video decoder 300 that the current video block has the same motion information as another video block.
[0088] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0089] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0090] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0091] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0092] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0093] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0094] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0095] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform to the transform coefficient video block respectively to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0096] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block effect artifacts in the video block.
[0097] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.
[0098] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 may be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0099] The video decoder 300 may be configured to perform any or all of the techniques of the present disclosure. In Figure 3 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video decoder 300. In some examples, a processor may be configured to perform any or all of the techniques described in the present disclosure.
[0100] In Figure 3 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 may perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.
[0101] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-encoded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.
[0102] The motion compensation unit 302 can generate a motion-compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.
[0103] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0104] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy encoding / decoding, signal prediction, and residual signal reconstruction. A slice can be the entire picture or can also be a region of the picture.
[0105] The intra prediction unit 303 can use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0106] The reconstruction unit 306 can obtain the decoded block, for example, by adding a residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the cache 307, and the cache 307 provides reference blocks for subsequent motion compensation / intra prediction, and the cache 307 also generates the decoded video for presentation on a display device.
[0107] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. In addition, although some embodiments describe the video coding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by the decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates.
[0108] 1. Brief Overview
[0109] The present disclosure relates to video coding and decoding techniques. Specifically, it relates to combined intra block copy, where the reference (or prediction) block is obtained using samples in the current picture, and intra prediction, as well as other coding and decoding tools in image / video coding and decoding. The present disclosure can be applied to existing video coding and decoding standards such as HEVC or multi-functional video coding (VVC). The present disclosure can also be applicable to future video coding and decoding standards or video codecs.
[0110] 2. Introduction
[0111] Video coding standards have mainly evolved from the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, and ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed H.264 / MPEG-2 Video and H.264 / MMPEG-4 Advanced Video Coding (AVC) as well as the H.264 / HEVC standard. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard, with the goal of reducing the bitrate by 50% compared to HEVC. ITU-T VCEG (Q6 / 16) and ISO / IEC MPEG (JTC 1 / SC 29 / WG 5) are studying the potential requirements for standardizing future video coding technologies whose compression capabilities significantly exceed the current VVC standard. Such future standardization actions may take the form of (one or more) additional extensions to VVC or a completely new standard. The group jointly collaborating on this exploration activity is called the Joint Video Exploration Team (JVET) to evaluate the compression technology designs proposed by experts in the field. The newly developed decoding features and coding methods that have been implemented in the Enhanced Compression Model (ECM) software are being coordinated and explored by the Joint Video Exploration Team (JVET) of ITU-T VCEG and ISO / IEC MPEG as potential enhanced video coding technologies beyond the capabilities of VVC.
[0112] 2.1. Coding and Decoding Processes of Typical Video Codecs
[0113] Figure 4 An example of the encoder block diagram of VVC is shown, which includes three loop filter blocks: the Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding an offset and applying a Finite Impulse Response (FIR) filter, respectively, where the encoded side information signals the offset and filter coefficients. ALF is located in the final processing stage of each picture and can be regarded as a tool to attempt to capture and fix the artifacts generated in the previous stage.
[0114] 2.2 Intra mode coding and decoding with 67 intra prediction modes
[0115] To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65 as Figure 5 shown, and the planar mode and DC mode remain unchanged. These denser directional intra prediction modes are applicable to all block sizes and to both luma intra prediction and chroma intra prediction.
[0116] In HEVC, each intra-coded block has a square shape and the length of each of its sides is a power of 2. Therefore, no division operation is required to generate the intra prediction value using the DC mode. In VVC, a block can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of a non-square block.
[0117] 2.2.1 Wide-angle intra prediction
[0118] Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index further depends on the block shape. The traditional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VVC, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The original mode index is used to signal the replaced mode, and the original mode index is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode coding and decoding method remains unchanged.
[0119] To support these prediction directions, a top reference of length 2W + 1 and a left reference of length 2H + 1 are defined as Figure 6 shown.
[0120] The number of modes replaced in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2-1.
[0121] Table 2-1 Intra prediction modes replaced by wide-angle modes Aspect ratio Replaced intra prediction mode W / H == 16 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 W / H == 8 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 W / H == 4 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 W / H == 2 Modes 2, 3, 4, 5, 6, 7, 8, 9 W / H == 1 None W / H == 1 / 2 Modes 59, 60, 61, 62, 63, 64, 65, 66 W / H == 1 / 4 Modes 57, 58, 59, 60, 61, 62, 63, 64, 65, 66 W / H == 1 / 8 Modes 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66 W / H == 1 / 16 Modes 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66
[0122] As Figure 7 shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the increased gap Δp αNegative impacts brought. If the wide-angle mode represents non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, which are [-14, -12, -10, -6, 72, 76, 78, 80]. When predicting blocks through these modes, the samples in the reference cache are directly copied without applying any interpolation. Through this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of the traditional prediction mode and the non-fractional mode in the wide-angle mode.
[0123] In VVC, 4:2:2 and 4:4:4 chrominance formats as well as 4:2:0 chrominance format are supported. The chrominance derivation mode (DM) derivation table for the 4:2:2 chrominance format was initially transplanted from HEVC, and the number of entries was extended from 35 to 67 to align with the extension of the intra prediction mode. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luma intra prediction modes with ranges from 2 to 5 are mapped to 2. Therefore, the chrominance DM derivation table for the 4:2:2 chrominance format is updated by replacing some values of the entries in the mapping table to more precisely target the prediction angles for chrominance blocks.
[0124] 2.3. Inter-frame prediction
[0125] For each inter-frame prediction CU, the motion parameters include the motion vector, the reference picture index and the reference picture list usage index, and additional information required for the new decoding features of VVC that will be used for inter-frame prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When the CU is coded and decoded in the skip mode, the CU is associated with a PU and has no significant residual coefficients, no coded motion vector difference (delta) or reference picture index. A Merge mode is specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, including spatial candidates and temporal candidates, as well as additional items introduced in VVC. The Merge mode can be applied to any inter-frame prediction CU, not just the skip mode. An alternative to the Merge mode is the explicit transmission of the motion parameters, where the motion vector, the corresponding reference picture index and the reference picture list usage flag for each reference picture list, and other required information are signaled explicitly for each CU.
[0126] 2.4. Intra block copy (IBC)
[0127] Intra Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. As is well known, it significantly improves the codec efficiency of screen content materials. Since the IBC mode is implemented as a block-level codec mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has been reconstructed within the current picture. The luminance block vector of the CU coded by IBC has integer precision. The chrominance block vector is also rounded to integer precision. When used in combination with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precisions. The CU coded by IBC is regarded as a third prediction mode in addition to the intra or inter prediction modes. The IBC mode is applicable to CUs with a width and height both less than or equal to 64 luma samples.
[0128] On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height not greater than 16 luma samples. For non-Merge modes, a block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a block-matching based local search will be performed.
[0129] In the hash-based search, the hash key matching (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For a current block with a larger size, the hash key is determined to match the hash key of the reference block when the hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector costs of each matching reference are calculated, and the one with the minimum cost is selected.
[0130] In the block-matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode uses flags to signal that it can be signaled as the IBC AMVP mode or the IBC skip / Merge mode, as follows:
[0131] - IBC skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC-coded blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and paired candidates.
[0132] -IBC AMVP mode: The block vector difference is coded and decoded in the same way as the motion vector difference. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the upper neighbor (if coded by IBC). When either neighbor is not available, the default block vector is used as the prediction value. A flag is signaled to indicate the block vector prediction value index.
[0133] 2.5. IBC Motion Candidates
[0134] The term "block" can represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB, or a video processing unit including multiple samples / pixels. The block can be rectangular or non-rectangular.
[0135] For an IBC-coded block, a block vector (BV) is used to indicate the displacement from the current block to a reference block that has been reconstructed within the current picture.
[0136] W and H are the width and height of the current block (e.g., a luminance block).
[0137] The non-adjacent spatial candidates of the current coded block are the adjacent spatial candidates of the virtual block in the i-th search round (as Figure 9 shown). The width and height of the virtual block in the i-th search round are calculated by the following formulas: newWidth = i × 2 × gridX + W, newHeight = i × 2 × gridY + H. Obviously, if the search round i is 0, the virtual block is the current block.
[0138] Hereinafter, the BV prediction value is also a BV candidate. The skip mode is also the Merge mode. BV candidates can be divided into several groups according to some criteria. Each group is called a subgroup. For example, we can take the adjacent spatial and temporal BV candidates as the first subgroup, and the remaining BV candidates as the second subgroup; in another example, we can also take the first N (N ≥ 2) BV candidates as the first subgroup, the next M (M ≥ 2) BV candidates as the second subgroup, and the remaining BV candidates as the third subgroup.
[0139] 2.6. Merge Mode with MVD (MMVD)
[0140] In addition to the Merge mode, in the case where implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag to specify whether the MMVD mode is used for the CU.
[0141] In MMVD, after selecting Merge candidates, they are further refined by MVD information transmitted via signals. The further information includes a Merge candidate flag, an index for specifying the motion amplitude, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The MMVD candidate flag is transmitted via signals to specify which one to use between the first Merge candidate and the second Merge candidate.
[0142] The distance index specifies the motion amplitude information and indicates a predefined offset from the starting point. As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2-2.
[0143] Table 2-2 - Relationship between the distance index and the predefined offset Distance index 0 1 2 3 4 5 6 7 Offset (in luminance samples) 1 / 4 1 / 2 1 2 4 8 16 32
[0144] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 2-3. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a uni-directional prediction MV or a bi-directional prediction MV where both lists point to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 2-3 specify the signs of the MV offsets added to the starting MV. When the starting MV is a bi-directional prediction MV with two MVs pointing to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture) and the POC difference in List 0 is greater than the POC difference in List 1, then the symbols in Table 2-3 specify the signs of the MV offsets added to the List 0 MV component of the starting MV, and the signs for the List 1 MV have opposite values. Otherwise, if the POC difference in List 1 is greater than the POC difference in List 0, then the symbols in Table 2-3 specify the signs of the MV offsets added to the List 1 MV component of the starting MV, and the signs for the List 0 MV have opposite values.
[0145] The MVD is scaled according to the POC differences in each direction. If the POC differences in both lists are the same, no scaling is required. Otherwise, if the POC difference in List 0 is greater than the POC difference in List 1, then as Figure 26 described, by defining the POC difference of L0 as td and the POC difference of L1 as tb, the MVD of List 1 is scaled. If the POC difference of L1 is greater than the POC difference of L0, then the MVD of List 0 is scaled in the same way. MVD. If the starting MV is unidirectionally predicted, the MVD is added to the available MVs.
[0146] Table 2-3 - Signs of MV Offsets Specified by Direction Index Direction index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + -
[0147] 2.7. Symmetric MVD Coding and Decoding
[0148] In VVC, in addition to the regular unidirectional prediction mode MVD signaling and bidirectional prediction mode MVD signaling, the symmetric MVD mode is applied to the bidirectional prediction MVD signaling. In the symmetric MVD mode, the motion information including the reference picture indexes of both list 0 and list 1 and the MVD of list 1 is not signaled but derived.
[0149] The decoding process of the symmetric MVD mode is as follows:
[0150] 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows:
[0151] — If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0.
[0152] — Otherwise, if the nearest reference picture in list -0 and the nearest reference picture in list -1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the list -0 reference picture and the list -1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.
[0153] 2) At the CU level, if the CU is bidirectionally coded and BiDirPredFlag is equal to 1, the symmetric mode flag indicating whether the symmetric mode is used is signaled explicitly.
[0154] When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indexes of list 0 and list 1 are respectively set to be equal to the reference picture pair. MVD1 is set to be equal to (-MVD0). The final motion vector is as shown in the following formula.
[0155] In an encoder, symmetric MVD motion estimation starts from an initial MV assessment. A set of initial MV candidates includes MVs obtained from unidirectional prediction search, MVs obtained from bidirectional prediction search, and MVs from the AMVP list. One with the lowest distortion rate cost is selected as the initial MV for symmetric MVD motion search.
[0156] 2.8. Bidirectional Optical Flow (BDOF)
[0157] The Bidirectional Optical Flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, was included in JEM. Compared with the JEM version, BDOF in VVC is a simpler version and requires much less computation, especially in terms of the number of multiplications and the size of the multiplier.
[0158] BDOF is used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions:
[0159] — The CU is encoded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in the display order, and the other of the two reference pictures is after the current picture in the display order;
[0160] — The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same;
[0161] — Both of the two reference pictures are short-term reference pictures;
[0162] — The CU is not encoded and decoded using the affine mode or the SbTMVP Merge mode;
[0163] — The CU has more than 64 luma samples;
[0164] — Both the CU height and the CU width are greater than or equal to 8 luma samples;
[0165] — The BCW weight index indicates equal weights;
[0166] — WP is not enabled for the current CU;
[0167] — The CIIP mode is not used for the current CU.
[0168] BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4x4 sub-block, motion refinement (v x ,v y) It is calculated by minimizing the difference between the L0 predicted sample points and the L1 predicted sample points. Then motion refinement is used to adjust the bidirectional predicted sample point values in the 4x4 sub-block. The following steps are applied during the BDOF process. By directly calculating the difference between two neighboring sample points, the horizontal and vertical gradients of the two prediction signals, and k = 0, 1, are calculated, that is,
[0169] where I (k) (i, j) is the sample point value at the coordinate (i, j) of the prediction signal in the list k, k = 0, 1, and shift1 is calculated based on the luminance bit depth bitDepth as shift1 = max(6, bitDepth - 6).
[0170] Then, the autocorrelations and cross-correlations of the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated as
[0171] where
[0172] where, Ω is a 6x6 window around the 4x4 sub-block, and the values of n a and n b are respectively set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8).
[0173] Then, using the cross-correlation terms and autocorrelation terms, the motion refinement (v x , v y ) is derived using the following method:
[0174] where, is the floor function, and
[0175] Based on the motion refinement and gradients, the following adjustments are calculated for each sample point in the 4x4 sub-block:
[0176] Finally, in the way shown below to adjust the bidirectional predicted sample points, the BDOF sample points of the CU are calculated: pred BDOF (x, y) = (I(0) (x, y) + I (1) (x, y) + b(x, y) + o offset ) >> shift(2 - 7)
[0177] These values are selected such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.
[0178] To derive the gradient values, some predicted samples I in list k (k = 0, 1) outside the current CU boundary (k) (i, j) need to be generated. As Figure 10 shown, BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating predicted samples outside the boundary, the predicted samples (white positions) in the extended region are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without using interpolation, and the regular 8 - tap motion - compensated interpolation filter is used to generate the predicted samples (gray positions) inside the CU. These extended sample values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e., repeated) from their nearest neighbors.
[0179] When the width and / or height of the CU is greater than 16 luma samples, the CU is divided into sub - blocks with width and / or height equal to 16 luma samples, and the sub - block boundaries are considered as the CU boundaries in the BDOF process. The maximum unit size of the BDOF process is limited to 16x16. For each sub - block, the BDOF process can be skipped. When the SAD between the initial L0 predicted sample and the L1 predicted sample is less than the threshold, the BDOF process is not applied to the sub - block. The threshold is set to be equal to (8 * W * (H >> 1)), where W represents the sub - block width and H represents the sub - block height. To avoid the additional complexity of SAD calculation, the SAD calculated in the DVMR process between the initial L0 predicted sample and the L1 predicted sample is reused here.
[0180] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., for either of the two reference pictures, luma_weight_lx_flag is 1, then BDOF is also disabled. When the CU is encoded / decoded in the symmetric MVD mode or the CIIP mode, BDOF is also disabled.
[0181] 2.9. Combined Inter - and Intra - prediction (CIIP)
[0182] 2.10. Affine Motion Compensation Prediction
[0183] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are various motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As Figure 11 shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters).
[0184] For the 4-parameter affine motion model, the motion vector at the sampling position (x, y) in the block is derived as:
[0185] For the 6-parameter affine motion model, the motion vector at the sampling position (x, y) in the block is derived as:
[0186] where (mv 0x , mv 0y ) is the motion vector of the upper-left control point, (mv 1x , mv 1y ) is the motion vector of the upper-right control point, and (mv 2x , mv 2y ) is the motion vector of the lower-left control point.
[0187] To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector of each 4x4 luminance sub-block, the motion vector of the central sample of each sub-block (as Figure 13 shown) is calculated according to the above equations and rounded to 1 / 16 fractional precision. Then a motion compensation interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size of the chrominance component is also set to 4x4. The MV of a 4x4 chrominance sub-block is calculated as the average of the MVs of four corresponding 4x4 luminance sub-blocks.
[0188] Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode.
[0189] 2.10.1. Affine Merge Prediction
[0190] The AF_MERGE mode can be applied to CUs with both width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of neighboring CUs in the spatial domain. There can be up to five CPMV candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPMV candidates are used to form the affine Merge candidate list:
[0191] — Inherited affine Merge candidates inferred from the CPMV of neighboring CUs;
[0192] — Constructed affine Merge candidate CPMV derived using the translational MVs of neighboring CUs;
[0193] — Zero MV.
[0194] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are shown in Figure 13 . For the left prediction value, the scan order is A0 -> A1, and for the upper prediction value, the scan order is B0 -> B1 -> B2. Only the first inherited candidate from each side is selected. Duplicate removal checking is not performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive the CPMV candidates in the affine Merge list of the current CU. As shown, if the neighboring bottom-left block A is coded in the affine mode, the motion vectors v 2 , v 3 and v 4 of the top-left, top-right, and bottom-left corners of the CU containing block A are obtained. When block A is coded using the 4-parameter affine model, two CPMVs of the current CU are calculated based on v 2 and v 3 . In the case where block A is coded using the 6-parameter affine model, three CPMVs of the current CU are calculated based on v 2 , v 3 and v 4 .
[0195] The constructed affine candidates refer to candidates constructed by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the specified spatial neighbors and temporal neighbors shown in Figure 15 . CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV 1 , the B2 -> B3 -> A2 block is checked, and the MV of the first available block is used. For CPMV 2 , the B1 > B0 block is checked, and for CPMV 3The A1 > A0 block is checked. TMVP is used as CPMV 4 (if available).
[0196] After the MVs at the four control points are obtained, an affine Merge candidate is constructed based on the motion information. The following combinations of the control point MVs are used for sequential construction:
[0197] {CPMV 1 , CPMV 2 , CPMV 3}}, {CPMV 1 , CPMV 2 , CPMV 4}}, {CPMV 1 , CPMV 3 , CPMV 4}}, {CPMV 2 , CPMV 3 , CPMV 4}}, {CPMV 1 , CPMV 2}}, {CPMV 1 , CPMV 3}.
[0198] Combinations of 3 CPMVs construct 6-parameter affine Merge candidates, and combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of the control point MVs are discarded.
[0199] After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list.
[0200] 2.10.2. Affine AMVP Prediction
[0201] The affine AMVP mode can be applied to CUs with width and height both greater than or equal to 16. The affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and it is generated by sequentially using the following four types of CPMV candidates:
[0202] — Inherited affine AMVP candidates speculated from the CPMVs of neighboring CUs;
[0203] — Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs;
[0204] — Translation MVs from neighboring CUs;
[0205] — Zero MVs.
[0206] The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture in the current block are considered. When inserting the inherited affine motion prediction values into the candidate list, the deduplication process is not applied.
[0207] The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 15 . The same checking order as in the affine Merge candidate construction is used. Additionally, the reference picture indices of the neighboring blocks are also checked. The first block in the checking order is used, which is inter-coded and has the same reference picture as the current CU. Only when the current CU is coded using the 4-parameter affine mode and both mv 0 and mv 1 are available, then they are added as a candidate to the affine AMVP list. When the current CU is coded using the 6-parameter affine mode and all three CPMVs are available, then they are added as a candidate to the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0208] If, after checking the inherited affine AMVP candidates and the constructed AMVP candidates, the affine AMVP list candidates are still less than 2, then mv 0 , mv 1 and mv 2 are added in order as translation MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list.
[0209] 2.10.3. Affine Motion Information Storage
[0210] In VVC, the CPMVs of affine CUs are stored in a separate cache. The stored CPMVs are only used for the inherited CPMVs in the affine Merge mode and the inherited CPMVs in the affine AMVP mode for the most recently coded CUs. The sub-block MVs derived from the CPMVs are used for motion compensation, MV derivation for the Merge / AMVP lists of translation MVs, and deblocking.
[0211] To avoid the picture line buffer for additional CPMV, the affine motion data inheritance from the CU above the CTU is processed differently from that from the normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the row above the CTU, the left-bottom and right-bottom sub-block MVs in the line buffer are used for affine MVP derivation instead of the CPMV. In this way, the CPMV is only stored in the local buffer. If the candidate CU is for 6-parameter affine coding / decoding, the affine model is degraded to a 4-parameter model. As Figure 16 shown, along the top boundary of the CTU, the left-bottom and right-bottom sub-block motion vectors of the CU are used for affine inheritance of the CU in the bottom of the CTU.
[0212] 2.10.4. Prediction refinement using optical flow for affine mode
[0213] Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a more refined motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after the sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the differences derived from the optical flow equations. PROF is described as the following four steps:
[0214] Step 1) Sub-block-based affine motion compensation is performed to generate sub-block prediction I(i, j).
[0215] Step 2) Using a 3-tap filter [-1, 0, 1], the spatial gradients g x (i, j) and g y (i, j) are calculated at each sample position. The gradient calculation is exactly the same as that in BDOF. g x (i, j) = (I(i + 1, j) >> shift1) - (I(i - 1, j) >> shfft1) (2 - 10) g y (i, j) = (I(i, j + 1) >> shift1) - (I(i, j - 1) >> shift1) (2 - 11)
[0216] shift1 is used to control the accuracy of the gradient. For gradient calculation, the sub-block (i.e., 4x4) prediction is extended by one sample on each side. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extended boundaries are copied from the nearest integer pixel positions in the reference picture.
[0217] Step 3) Luminance prediction refinement is calculated by the following optical flow equation. ΔI(i, j) = g x (i, j) * Δv x (i, j) + g y (i, j) * Δv y (i, j) (2 - 12)
[0218] Where Δv(i, j) is the difference between the sample MV (represented by v(i, j)) calculated for the sample position (i, j) and the sub - block MV of the sub - block to which the sample (i, j) belongs, as shown in Figure 17 . Δv(i, j) is quantized in units of 1 / 32 luminance sample precision.
[0219] Since the affine model parameters and the sample position relative to the sub - block center do not change from sub - block to sub - block, Δv(i, j) can be calculated for the first sub - block and reused for other sub - blocks in the same CU. Let dx(i, j) and dy(i, j) be the horizontal and vertical offsets from the sample position (i, j) to the center (X SB , y SB ) of the sub - block. Δv(x, y) can be derived by the following equation:
[0220] To maintain accuracy, the input of the sub - block (X SB , y SB ) is calculated as ((W SB - 1) / 2, (H SB - 1) / 2), where W SB and H SB are the sub - block width and sub - block height respectively.
[0221] For the 4 - parameter affine model,
[0222] For the 6 - parameter affine model,
[0223] where (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the top - left, top - right and bottom - left control - point motion vectors, and w and h are the width and height of the CU.
[0224] Step 4) Finally, the luminance prediction refinement ΔI(i, j) is added to the sub - block prediction I(i, j). The final prediction I’ is generated by the following equation. I′(i, j) = I(i, j) + ΔI(i, j) (2-17)
[0225] PROF is not applicable to two cases of affine coded / decoded CUs: 1) all control point MVs are the same, which indicates that the CU only has translational motion; 2) the affine motion parameters are greater than the specified limit because the sub-block based affine MC is degraded to CU based MC to avoid large memory access bandwidth requirements.
[0226] Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) if the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low-delay picture, then PROF is not applied because the improvement introduced by PROF for this case is small. In this way, the affine motion estimation using PROF can be accelerated.
[0227] 2.11. Sub-block based Temporal Motion Vector Prediction (SbTMVP)
[0228] VVC supports the sub-block based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the reference picture to improve the motion vector prediction and Merge mode of the CUs in the current picture. The same reference picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects:
[0229] - TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level;
[0230] - TMVP extracts the temporal motion vector from the reference block in the reference picture (the reference block is the bottom-right block or the center block relative to the current CU), while SbTMVP applies a motion displacement before extracting the temporal motion information from the reference picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0231] The SbTVMP process is shown in Figure 18A and Figure 18B SbTMVP predicts the motion vectors of the sub-CUs within the current CU in two steps. In the first step, Figure 18AThe spatial neighborhood A1 in is examined. If A1 has a motion vector that uses the collocated picture as its reference picture, this motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0).
[0232] In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain sub-CU level motion information (motion vector and reference index) from the collocated picture as Figure 18B shown. Figure 18B The example in assumes that the motion displacement is set to the motion of block A1. Then, the motion information of the corresponding block (the smallest motion grid covering the central sample point) for each sub-CU in the collocated picture is used to derive the motion information of the sub-CU. After the motion information of the collocated sub-CU is identified, the motion information is converted to the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0233] In VVC, a combined sub-block based Merge list containing both SbTMVP candidates and affine Merge candidates is used for signaling the sub-block based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub-block based Merge candidates, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5.
[0234] The sub-CU size used in SbTMVP is fixed to 8x8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs with a width and height both greater than or equal to 8.
[0235] The encoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P-slice or B-slice, additional RD checks are performed to decide whether to use the SbTMVP candidate.
[0236] 2.12. Adaptive Motion Vector Resolution (AMVR)
[0237] In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in units of quarter luminance samples. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be coded and decoded with different precisions. Depending on the mode of the current CU (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows:
[0238] — Normal AMVP mode: quarter luminance samples, half luminance samples, integer luminance samples, or four luminance samples.
[0239] — Affine AMVP mode: quarter luminance samples, integer luminance samples, or 1 / 16 luminance samples.
[0240] If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both the horizontal MVD and the vertical MVD of reference list L0 and reference list L1) are zero, a quarter luminance sample MVD resolution is assumed.
[0241] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luminance sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter luminance sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half luminance sample or other MVD precision (integer or four luminance samples) is used for normal AMVP CUs. In the case of half luminance samples, the half luminance sample position uses a 6-tap interpolation filter instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer luminance sample or four luminance sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer luminance sample MVD precision or 1 / 16 luminance sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter luminance samples, half luminance samples, integer luminance samples, or four luminance samples), the motion vector prediction value of the CU will be rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded to zero (i.e., negative motion vector prediction values are rounded to positive infinity and positive motion vector prediction values are rounded to negative infinity).
[0242] The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM11, the RD check for MVD accuracy outside the quarter-luma samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for the MVD accuracy of the quarter-luma samples and the RD cost for the MV accuracy of the integer-luma samples are calculated. Then, the RD cost for the MVD accuracy of the integer-luma samples is compared with the RD cost for the MVD accuracy of the quarter-luma samples to decide whether it is necessary to further check the RD cost for the MVD accuracy of the four-luma samples. When the RD cost for the MVD accuracy of the quarter-luma samples is much smaller than the RD cost for the MVD accuracy of the integer-luma samples, the RD check for the MVD accuracy of the four-luma samples is skipped. Then, if the RD cost for the MVD accuracy of the integer-luma samples is significantly greater than the best RD cost for the previously tested MVD accuracy, the check for the MVD accuracy of the half-luma samples is skipped. For the affine AMVP mode, if the affine inter prediction mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, the Merge / skip mode, the normal AMVP mode with the MVD accuracy of the quarter-luma samples, and the affine AMVP mode with the MVD accuracy of the quarter-luma samples, the MV accuracy of the 1 / 16-luma samples and the affine inter prediction mode with the 1-pixel MV accuracy are not checked. Additionally, in the affine inter prediction modes with the 1 / 16-luma samples and the quarter-luma samples MV accuracy, the affine parameters obtained in the affine inter prediction mode with the quarter-luma samples MV accuracy are used as the starting search points.
[0243] 2.13. Bi-directional prediction with CU-level weights (BCW)
[0244] In HEVC, the bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (2 - 18)
[0245] In weighted average bi-prediction, five weights are allowed, w ∈ {-2, 3, 4, 5, 10}. For each bi-predicted CU, the weight w is determined in one of two ways: 1) for non-Merge CUs, the weight index is signaled after the motion vector difference; 2) for Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights are used (w ∈ {3, 4, 5}).
[0246] At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized below. For further details, the reader may refer to the VTM software. When combined with AMVR, if the current picture is a low-delay picture, unequal weights are only conditionally checked for 1-pixel and 4-pixel motion vector precisions.
[0247] When combined with affine, affine ME is performed for unequal weights if and only if the affine mode is selected as the current best mode.
[0248] When the two reference pictures in bi-prediction are the same, unequal weights are only conditionally checked.
[0249] When certain conditions are met, unequal weights are not searched, depending on the POC distance between the current picture and its reference pictures, the coding / decoding QP, and the temporal level.
[0250] The BCW weight index is coded using one context-coded bit followed by bypass-coded bits. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, additional bits are signaled using bypass coding to indicate which unequal weight is used.
[0251] Weight prediction (WP) is a coding tool supported by the H.264 / AVC and HEVC standards for efficient coding and decoding of video content in fading situations. Support for WP has also been added in the VVC standard. WP allows signaling of weighting parameters (weights and offsets) for each reference picture in each of reference picture lists L0 and L1. Then, during motion compensation, the (multiple) weights and (multiple) offsets corresponding to the (multiple) reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For MergeCUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV.
[0252] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded and decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights.
[0253] 2.14. Local Illumination Compensation (LIC)
[0254] Local Illumination Compensation (LIC) is a coding tool for solving the problem of local illumination change between the current picture and its temporal reference pictures. LIC is based on a linear model, where a scaling factor and an offset are applied to the reference samples to obtain the predicted samples of the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x, y) = α · P r (x + v x , y + v y ) + β
[0255] where P(x,y) is the predicted signal of the current block at coordinates (x,y); P r (x + Vx, y + Vy) is the reference block pointed to by the motion vector (v x , v y ); and α and β are the corresponding scaling factor and offset applied to the reference block. Figure 19 shows the LIC process. In Figure 19 , when LIC is applied to a block, the least mean square error (LMSE) method is adopted, by minimizing the neighboring samples of the current block (i.e., Figure 19the corresponding reference sample points in the time-domain reference picture (i.e., Figure 19 T0 or T1 in ) to derive the values of the LIC parameters (i.e., α and β). Additionally, to reduce the computational complexity, both the template samples and the reference template samples are subsampled (adaptive subsampling) to derive the LIC parameters, i.e., only Figure 19 the shaded samples in are used to derive α and β.
[0256] To improve the coding and decoding performance, the short side is not subsampled, as Figure 20 shown.
[0257] 2.15. Decoder-side Motion Vector Refinement (DMVR)
[0258] To improve the accuracy of the MVs in the Merge mode, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. During the bi-prediction operation, refined MVs are searched around the initial MVs in the reference picture list LI and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture lists LI and L1. As Figure 21 shown, the SAD between two blocks based on each MV candidate (e.g., MV0’ and MV1’) around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-prediction signal.
[0259] In VVC, the application of DMVR is restricted and is only applied to the CUs encoded and decoded using the following modes and features:
[0260] - CU-level Merge mode with bi-prediction MVs
[0261] - For the current picture, one reference picture is past and the other reference picture is future
[0262] - The distances from the two reference pictures to the current picture (i.e., POC differences) are the same
[0263] - Both reference pictures are short-term reference pictures
[0264] - The CU has more than 64 luma samples
[0265] - Both the CU height and the CU width are greater than or equal to 8 luma samples
[0266] - The BCW weight index indicates equal weights
[0267] - WP is not enabled for the current block
[0268] - The CIIP mode is not used for the current block.
[0269] The refined MVs derived through the DMVR process are used to generate inter - predicted samples and are also used for temporal motion vector prediction in future picture coding. The original MVs are used for the de - blocking process and are also used for spatial motion vector prediction in future CU coding.
[0270] Additional features of DMVR are mentioned in the following sub - articles.
[0271] 2.15.1. Search Scheme
[0272] In DVMR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point examined by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0 + MV_offset (2 - 19) MV1′ = MV1 - MV_offset (2 - 20)
[0273] where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luminance samples starting from the initial MV. The search includes an integer - sample offset search stage and a fractional - sample refinement stage.
[0274] The integer - sample offset search uses a 25 - point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer - sample stage of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and examined in raster - scan order. The point with the minimum SAD is selected as the output of the integer - sample offset search stage. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between the reference blocks pointed to by the initial MV candidates reduces the SAD value by 1 / 4.
[0275] After the integer - sample search, there is fractional - sample refinement. To save computational complexity, the fractional - sample refinement is derived using the parametric error - surface equation instead of using SAD comparison for additional search. The fractional - sample refinement is conditionally invoked based on the output of the integer - sample search stage. When the integer - sample search stage ends at the center with the minimum SAD in the first or second iteration search, the fractional - sample refinement is further applied.
[0276] In the sub - pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two - dimensional parabolic error - surface equation of the following form E(x, y) = A(x - x min ) 2 +B(y - y min )2 +C (2-21)
[0277] where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1, 0) - E(1, 0)) / (2(E(-1, 0) + E(1, 0) - 2E(0, 0))) (2-22) y min = (E(0, -1) - E(0, 1)) / (2((E(0, -1) + E(0, 1) - 2E(0, 0))) (2-23).
[0278] x min and y min values are automatically limited between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined differential MV.
[0279] 2.15.2. Bilinear Interpolation and Sample Padding
[0280] In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate the samples at the fractional positions. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples during the search process in DMVR. Another important effect is that by using the bilinear filter, within the 2-sample search range, compared with the normal motion compensation process, DVMR does not access more reference samples. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples of the normal MC process, samples will be padded from those available samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV.
[0281] 2.15.3. Maximum DMVR Processing Unit
[0282] When the width and / or height of a CU is greater than 16 luma samples, it is further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16.
[0283] 2.16. Multi-pass decoder-side motion vector refinement
[0284] In this contribution, multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the coded / decoded block. In the second pass, BM is applied to each 16x16 sub-block within the coded / decoded block. In the third pass, the MVs in each 8x8 sub-block are refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction.
[0285] 2.16.1. First pass - Block-based bilateral matching MV refinement
[0286] In the first pass, refined MVs are derived by applying BM to the coded / decoded block. Similar to decoder-side motion vector refinement (DMVR), refined MVs are searched around two initial MVs (MV0 and MV1) in reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between two reference blocks in L0 and L1.
[0287] BM performs a local search to derive integer sample accuracy intDeltaMV and half-pixel sample accuracy halfDeltaMV. The local search applies a 3×3 square search pattern to loop within a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8.
[0288] The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the distorted DC effect between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV or halfDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range.
[0289] Further applying the current fractional sample refinement to derive the final deltaMV. Then, the refined MVs after the first pass are derived as:
[0290] ● MV0_pass1 = MV0 + deltaMV
[0291] ● MV1_pass1 = MV1 - deltaMV。
[0292] 2.16.2. Second pass - Sub - block - based bilateral matching MV refinement
[0293] In the second pass, refined MVs are derived by applying BM to 16×16 grid sub - blocks. For each sub - block, refined MVs are searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between two reference sub - blocks in L0 and L1.
[0294] For each sub - block, BM full - search is performed to derive integer - pixel accuracy intDeltaMV. The full - search has a search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8.
[0295] The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub - blocks, i.e., bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into 5 diamond - shaped search zones, as Figure 22 shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond - shaped area is processed in the order starting from the center of the search zone. In each area, search points are processed in raster - scan order from the upper - left corner of the area to the lower - right corner. When the minimum bilCost in the current search zone is less than a threshold equal to sbW * sbH, the integer - pixel full - search is terminated; otherwise, the integer - pixel full - search continues to the next search zone until all search points are checked.
[0296] BM performs local search to derive half - pixel accuracy halfDeltaMv. The search pattern and cost function are the same as those defined in Section 2.9.1.
[0297] Existing VVC DMVR fractional - pixel refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MVs in the second pass are derived as:
[0298] ● MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2)
[0299] ● MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2).
[0300] 2.16.3. Third Pass - Sub - block - based Bidirectional Optical Flow MV Refinement
[0301] In the third pass, refined MVs are derived by applying BDOF to 8x8 grid sub - blocks. For each 8×8 sub - block, starting from the refined MVs of the parent - child blocks in the second pass, BDOF refinement is applied to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between - 32 and 32.
[0302] The refined MVs in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as follows:
[0303] ● MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv
[0304] ● MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) - bioMv.
[0305] 2.17. Sample - based BDOF
[0306] In sample - based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample.
[0307] The coding / decoding block is divided into 8 8×8 sub - blocks. For each sub - block, it is determined whether to apply BDOF by checking the SAD between two reference sub - blocks against a threshold. If it is decided to apply BDOF to the sub - block, for each sample in the sub - block, a sliding 5x5 window is used, and the existing BDOF process is applied to each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the central sample of the window.
[0308] 2.18. Extended Merge Prediction
[0309] In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates:
[0310] (1) Spatial MVPs from spatial neighboring CUs
[0311] (2) Temporal MVP from collocated CU
[0312] (3) History-based MVP from FIFO table
[0313] (4) Paired-average MVP
[0314] (5) Zero MV.
[0315] The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU code in the Merge mode, the index of the best Merge candidate is encoded using rounding binary unary (TU). The first binary bit of the Merge index is coded using context, while bypass coding is used for the other binary bits.
[0316] The derivation process of the Merge candidates for each category is provided in this section. Similar to what is done in HEVC, VVC also supports the parallel derivation of the Merge candidate lists for all CUs within a certain size region.
[0317] 2.18.1. Spatial candidate derivation
[0318] The derivation of the spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates at the indicated positions, up to four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more CUs at positions B0, A0, B1, A1 are not available (e.g., because it belongs to another strip or slice) or are intra-coded. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thus improving the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrows in Figure 24 are considered, and the candidate is added to the list only if the corresponding candidates used for the redundancy check do not have the same motion information.
[0319] 2.18.2. Temporal candidate derivation
[0320] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal Merge candidate, the scaled motion vectors are derived based on the collocated CUs belonging to the collocated reference pictures. The reference picture list to be used for deriving the collocated CUs is explicitly signaled in the slice header. As Figure 25As shown by the dashed line in, a scaled motion vector for a temporal Merge candidate is obtained, and the scaled motion vector is scaled from the motion vector of a co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero.
[0321] The position of the temporal candidate is selected between candidates C0 and C1, as Figure 26 shown. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate.
[0322] 2.18.3. History-based Merge Candidate Derivation
[0323] After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of previously decoded blocks is stored in a table and used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a non-sub-block inter-coded CU, the associated motion information is added as a new HMVP candidate to the last entry of the table.
[0324] The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to find if there is the same HMVP in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are moved forward.
[0325] HMVP candidates can be used in the Merge candidate list construction process. The several most recent HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. For spatial or temporal Merge candidates, a redundancy check is applied to the HMVP candidates.
[0326] To reduce the number of redundancy check operations, the following simplification is introduced:
[0327] The number of HMPV candidates used for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table.
[0328] Once the total number of available Merge candidates reaches one less than the maximum allowed Merge candidates, the process of constructing the Merge candidate list from the HMVP is terminated.
[0329] 2.18.4. Pairwise Average Merge Candidate Derivation
[0330] Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the numbers represent the Merge indices of the Merge candidate list. The average motion vector is calculated separately for each reference list. If both motion vectors are available in a list, they are averaged even if the two motion vectors point to different reference pictures; if only one motion vector is available, that motion vector is used directly; if no motion vector is available, this list is kept invalid.
[0331] When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum number of Merge candidates is reached.
[0332] 2.18.5. Merge Estimation Region
[0333] The Merge Estimation Region (MER) allows for the independent derivation of the Merge candidate list for a CU within the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included for generating the Merge candidate list for the current CU. Additionally, the update process for the candidate list of the history-based motion vector prediction values is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2parMrglevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set in the form of log2_parallel_merge_level_minus2.
[0334] 2.19. New Merge Candidates
[0335] 2.19.1. Non-Adjacent Merge Candidate Derivation
[0336] In VVC, Figure 27The five spatially neighboring blocks and one temporally neighboring block shown are used to derive Merge candidates.
[0337] It is proposed to derive additional Merge candidates from positions not adjacent to the current block using the same pattern as in VVC. To achieve this, for each search round i, a virtual block is generated based on the current block as follows:
[0338] First, the relative position of the virtual block with respect to the current block is calculated by the following formula:
[0339] Offsetx = -i × gridX, Offsety = -i × gridY
[0340] where Offsetx and Offsety represent the offsets of the upper left corner of the virtual block relative to the upper left corner of the current block, and gridX and gridY are the width and height of the search grid.
[0341] Second, the width and height of the virtual block are calculated by the following formula:
[0342] newWidth = i × 2 × gridX + currWidth newHeight = i × 2 × gridY + currHeight.
[0343] where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.
[0344] gridX and gridY are currently set to currWidth and currHeight respectively.
[0345] Figure 28 The relationship between the virtual block and the current block is shown.
[0346] After generating the virtual block, blocks A i , B i , C i , D i and E i can be regarded as the spatially neighboring blocks of the virtual block in VVC, and their positions are obtained using the same pattern as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, blocks A i , B i , C i , D i and E i are the spatially neighboring blocks used in the VVC Merge mode.
[0347] When building the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighbors are used.
[0348] Press B for non-adjacent airspace Merge candidates 1 ->A 1 ->C 1 ->D 1 ->E 1 The order is inserted into the Merge list after the time domain Merge candidate.
[0349] 2.19.2.STMVP
[0350] It is proposed to use three spatial domain Merge candidates and one temporal domain Merge candidate to derive the average candidate as the STMVP candidate.
[0351] The STMVP is inserted before the spatial merge candidate in the upper left corner.
[0352] The STMVP candidate is deduplicated along with all previous merge candidates in the merge list.
[0353] For spatial candidates, the first three candidates in the current Merge candidate list are used.
[0354] For the temporal candidates, the same positions as the VTM / HEVC co-location positions are used.
[0355] For spatial candidates, the first, second, and third candidates inserted before STMVP in the current Merge candidate list are denoted as F, S, and T.
[0356] The temporal candidate having the same position as the VTM / HEVC co-location position used in TMVP is denoted as Col.
[0357] The motion vector of the STMVP candidate in prediction direction X (denoted as mvLX) is derived as follows:
[0358] 1) If the reference indices of the four Merge candidates are all valid and equal to zero in the prediction direction X (X = 0 or 1), then
[0359] mvLX=(mvLX_F+mvLX_S+mvLX_T+mvLX_Col)>>2
[0360] 2) If the reference indexes of three of the four Merge candidates are valid and equal to zero in the prediction direction X (X=0 or 1), then
[0361] mvLX = (mvLX_F × 3 + mvLX_S × 3 + mvLX_Col × 2) >> 3 or
[0362] mvLX = (mvLX_F × 3 + mvLX_T × 3 + mvLX_Col × 2) >> 3 or
[0363] mvLX = (mvLX_S × 3 + mvLX_T × 3 + mvLX_Col × 2) >> 3
[0364] 3) If the reference indices of two of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1), then
[0365] mvLX = (mvLX_F + mvLX_Col) >> 1 or
[0366] mvLX = (mvLX_S + mvLX_Col) >> 1 or
[0367] mvLX = (mvLX_T + mvLX_Col) >> 1
[0368] Note: If the time-domain candidate is not available, the STMVP mode is turned off.
[0369] 2.19.3. Merge List Size
[0370] If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8.
[0371] 2.20. Geometric Partitioning Mode (GPM)
[0372] In VVC, the geometric partitioning mode is supported for inter prediction. The CU-level flag is used as a Merge mode to signal the geometric partitioning mode, and other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. For each possible CU size w×h = 2 m ×2 n , where m,n ∈ {3…6} excluding 8x64 and 64x8, the geometric partitioning mode supports a total of 64 partitions.
[0373] When using this mode, the CU is divided into two parts by a geometrically positioned line ( Figure 29)。The position of the dividing line is derived mathematically from the perspective of a specific partition and an offset parameter. Each part in the geometric partition of the CU is inter - frame predicted using its own motion; only unidirectional prediction is allowed for each partition, that is, each part has a motion vector and a reference index. The unidirectional prediction motion constraint is applied to ensure that, like in traditional bidirectional prediction, each CU only requires two motion - compensated predictions. The unidirectional prediction motion for each partition is derived using the process described in 2.20.1.
[0374] If the geometric partition mode is used for the current CU, the geometric partition index (angle and offset) indicating the partition mode of the geometric partition and two Merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization for the GPM Merge indices is specified. After predicting each part of the geometric partition, as in 2.20.2, a hybrid process with adaptive weights is used to adjust the sample values along the geometric partition edge. This is the prediction signal for the entire CU, and the transform process and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partition mode is stored, as described in 2.20.3.
[0375] 2.20.1. Unidirectional Prediction Candidate List Construction
[0376] In 2.18, the unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (X equals the parity of n) of the n - th extended Merge candidate is used as the n - th unidirectional prediction motion vector of the geometric partition mode. These motion vectors are marked with "x" in Figure 30 . If the corresponding LX motion vector of the n - th extended Merge candidate does not exist, the L(1 - X) motion vector of the same candidate is used as the unidirectional prediction motion vector of the geometric partition mode.
[0377] 2.20.2. Hybrid Along Geometric Partition Edge
[0378] After predicting each part of the geometric partition using its own motion, a hybrid is applied to the two prediction signals to derive the samples around the geometric partition edge. The hybrid weight for each position of the CU is derived based on the distance between the independent position and the partition edge.
[0379] The distance from the position (x, y) to the partition edge is derived as:
[0380] where i and j are the indices of the angle and offset of the geometric segmentation, which depend on the geometric segmentation index transmitted through the signal. p x,j and P y,j The sign of depends on the angle index i.
[0381] The weight of each part of the geometric segmentation is derived as follows: wIdxL(x, y) = partIdx? 32 + d(x, y) : 32 - d(x, y) (2-28) w 1 (x, y) = 1 - w 0 (x, y) (2-30).
[0382] partIdx depends on the angle index i. The weight w 0 An example of is shown in Figure 31 .
[0383] 2.20.3. Motion vector field storage for geometric segmentation mode
[0384] Mv1 from the first part of the geometric segmentation, Mv2 from the second part of the geometric segmentation, and the combination Mv of Mv1 and Mv2 are stored in the motion vector field of the CU encoded / decoded in the geometric segmentation mode.
[0385] The type of motion vector stored for each independent position in the motion vector field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx ≤ 0? (1 - partIdx) : partIdx) (2-31)
[0386] where motionIdx is equal to d(4x + 2, 4y + 2), which is recalculated according to Equation (2-18). partIdx depends on the angle index i.
[0387] If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion vector field, otherwise, if sType is equal to 2, then the combination Mv of Mv1 and Mv2 is stored. The combined Mv is generated using the following procedure:
[0388] 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bidirectional prediction motion vector.
[0389] Otherwise, if Mv1 and Mv2 are from the same list, then only the unidirectional prediction motion Mv2 is stored.
[0390] 2.21. Multiple Hypothesis Prediction
[0391] In Multiple Hypothesis Prediction (MHP), up to two additional prediction values are signaled on top of the inter-frame AMVP mode, the regular Merge mode, the affine Merge, and the MMVD mode. The resulting overall prediction signal is iteratively accumulated using each additional prediction signal. p n+1 = (1 - α n+1 ) p n + α n+1 h n+1
[0392] The weight factor α is specified according to Table 2-4 below.
[0393] Table 2-4 - Weight Factors for MHP add_hyp_weight_idx α 0 1 / 4 1 -1 / 8
[0394] For the inter-frame AMVP mode, MHP is applied only when unequal weights in the BCW are selected in the bi-prediction mode.
[0395] The additional hypothesis can be the Merge mode or the AMVP mode. In the case of the Merge mode, the motion information is indicated by the Merge index, and the Merge candidate list is the same as in the geometric partitioning mode. In the case of the AMVP mode, the reference index, the MVP index, and the MVD are signaled.
[0396] 2.22. Non-Adjacent Spatial Candidates
[0397] Non-adjacent spatial Merge candidates are inserted after the TMVP in the regular Merge candidate list. The pattern of the spatial Merge candidates is shown in Figure 32 above. The distance between the non-adjacent spatial candidates and the current coding block is based on the width and height of the current coding block.
[0398] 2.23. Template Matching (TM)
[0399] Template Matching (TM) is a decoder-side MV derivation method for refining the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference picture (i.e., of the same size as the template). As Figure 33 shown, within the [-8, +8] pixel search range, a better MV is searched around the initial motion of the current CU. Two modified template matchings are proposed: determining the search step size based on the AMVR mode, and TM in the Merge mode can be cascaded using the bilateral matching process.
[0400] In the AMVP mode, an MVP candidate is determined by selecting one that achieves the minimum difference between the current block template and the reference block template based on the template matching error, and then the TM only performs MV refinement on this specific MVP candidate. The TM refines this MVP candidate by using an iterative diamond search starting from the full-pixel MVD accuracy within the [-8, +8] pixel search range (or 4 pixels for the 4-pixel AMVR mode). The AMVP candidate can be further refined by using a cross search with full-pixel MVD accuracy (or 4 pixels for the 4-pixel AMVR mode), and then depending on the AMVR mode specified in Table 2-5, half-pixel and quarter-pixel are used in sequence. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process.
[0401] Table 2-5 - Search patterns for AMVR and Merge mode using AMVR.
[0402] In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 2-5, the TM can be performed all the way to 1 / 8 pixel MVD accuracy, or skip those accuracies that exceed half-pixel MVD accuracy, depending on whether an alternative interpolation filter is used according to the merged motion information (i.e., used when AMVR is in the half-pixel mode). In addition, when the TM mode is enabled, the template matching can work in an independent process or an additional MV refinement process between the block-based bilateral matching (BM) method and the sub-block-based bilateral matching method, depending on whether the BM can be enabled according to its enabling condition check.
[0403] 2.24. Overlapped Block Motion Compensation (OBMC)
[0404] Overlapped Block Motion Compensation (OBMC) has been used in H.263 before. In JEM, different from H.263, OBMC can be turned on and off using the CU-level syntax. When OBMC is used in JEM, OBMC is performed on all motion compensation (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it also applies to the luminance and chrominance components. In JEM, the MC block corresponds to the coding / decoding block. When a CU is coded / decoded in the sub-CU mode (including sub-CUMerge, affine, and FRUC modes), each sub-block of the CU is an MC block. To uniformly process the CU boundary, OBMC is performed on all MC block boundaries at the sub-block level, where the sub-block size is set to be equal to 4×4, as Figure 34 shown.
[0405] When the OBMC is applied to the current sub - block, in addition to the current motion vector, the motion vectors of four connected neighboring sub - blocks, if available and different from the current motion vector, are also used to derive the prediction block of the current sub - block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal of the current sub - block.
[0406] Represent the prediction block based on the motion vector of the neighboring sub - block as P N , where N indicates the indices of the neighboring upper, lower, left, and right sub - blocks, and the prediction block based on the motion vector of the current sub - block is represented as P C . When P N is based on the motion information of neighboring sub - blocks that contain the same motion information as the current sub - block, OBMC is not performed starting from P N . Otherwise, each sample of P N is added to the same sample in P C , that is, the four rows / columns of P N are added to P C . The weight factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and the weight factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . Except for small MC blocks (i.e., when the height or width of the coded / decoded block is equal to 4, or when the CU is coded / decoded using the sub - CU mode), where only two rows / columns of P N are added to P C . In this case, the weight factors {1 / 4, 1 / 8} are used for P N , while the weight factors {3 / 4, 7 / 8} are used for P C . For P N generated based on the motion vectors of vertical (horizontal) neighboring sub - blocks, the samples in the same row (column) of P N are added to P C with the same weight factor.
[0407] In JEM, for a CU with a size less than or equal to 256 luma samples, a signal - transmitted CU - level flag is used to indicate whether OBMC is applied to the current CU. For a CU with a size greater than 256 luma samples or not coded / decoded using the AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is taken into account during the motion - estimation stage. The prediction signal formed using the motion information of the top neighboring block and the left - hand neighboring block by OBMC is used to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion - estimation process is applied.
[0408] 2.25. Multiple - transform selection (MTS) for kernel transform
[0409] In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of inter- and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the basis functions of the selected DST / DCT.
[0410] Table 2-6 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input
[0411] To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms.
[0412] To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applicable to luminance. The MTS signaling is skipped when one of the following conditions is met.
[0413] — The position of the last significant coefficient of the luminance TB is less than 1 (i.e., only DC);
[0414] — The last significant coefficient of the luminance TB is within the MTS zeroing region.
[0415] If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 2-7. A unified transform selection for ISP and implicit MTS is used by eliminating intra-mode and block-shape dependencies. If the current block is in ISP mode or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 is used for horizontal and vertical transform kernels. In terms of transform matrix precision, an 8-bit primary transform kernel is used. Thus, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8) all use an 8-bit primary transform kernel.
[0416] Table 2-7 - Transform and signaling mapping table
[0417] To reduce the complexity of large-sized DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are set to zero. Only the coefficients within the 16×16 low-frequency region are retained.
[0418] As in HEVC, the residual of a block can be coded and decoded using the transform skip mode. To avoid redundancy in syntax coding, when the MTS_CU_flag at the CU level is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. Additionally, when MTS is enabled for an inter-coded block, the implicit MTS can still be enabled.
[0419] 2.26. Sub-Block Transform (SBT)
[0420] In VTM, a sub-block transform is introduced for inter-predicted CUs. In this transform mode, for a CU, only a sub-part of the residual block is coded and decoded. When the cu_cbf of an inter-predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is coded and decoded. For the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is coded and decoded by a presumed adaptive transform, while the other part of the residual block is set to zero.
[0421] When SBT is used for an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, as Figure 35 shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to a binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to an asymmetric binary tree (ABT) partition. In the ABT partition, only the small region contains non-zero residuals. If one dimension of a CU is 8 (in terms of luminance samples), the 1:3 / 3:1 partition is not allowed along that dimension. A CU can have up to 8 SBT modes.
[0422] Position-dependent transform kernel selection is applied to the luminance transform blocks in SBT-V and SBT-H (the chrominance TB always uses DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are in Figure 35Specified in, for example, the horizontal and vertical transforms at SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transforms in both dimensions are set to DCT-2. Thus, the sub-block transforms jointly specify the TU slicing, cbf, and the horizontal and vertical kernel transform types of the residual block.
[0423] SBT is not applied to the CU encoded or decoded in the combined inter-intra mode.
[0424] 2.27. Adaptive Merge candidate reordering based on template matching.
[0425] To improve the encoding and decoding efficiency, after constructing the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list according to the ascending template matching cost. It is operated in the form of subgroups.
[0426] The template matching cost is measured by the SAD (sum of absolute differences) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes the motion information of bidirectional prediction, the corresponding reference sample is the average of the corresponding reference samples in reference list 0 and the corresponding reference samples in reference list 1, as Figure 36 shown. If the Merge candidate contains the motion information at the sub-CU level, the corresponding reference sample consists of the neighboring samples of the corresponding reference sub-block, as Figure 37 shown.
[0427] The sorting process is operated in the form of subgroups, as Figure 38 shown. The first three Merge candidates are sorted together. The subsequent three Merge candidates are sorted together. The template size (the width of the left template or the height of the upper template) is 1. The subgroup size is 3.
[0428] 2.28. Adaptive Merge candidate list
[0429] We can assume that the number of Merge candidates is 8. We take the first 5 Merge candidates as the first subgroup and the subsequent 3 Merge candidates as the second subgroup (i.e., the last subgroup).
[0430] For the encoder, after the Merge candidate list is constructed, some Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, as Figure 39 shown. More specifically, the template matching costs of the Merge candidates in all subgroups except the last subgroup are calculated; then, the Merge candidates in their own subgroups except the last subgroup are reordered; finally, the final Merge candidate list will be obtained.
[0431] For the decoder, after the Merge candidate list is constructed, as Figure 40 shown, some / none of the Merge candidates are adaptively re-ordered in ascending order of the Merge candidate cost. In Figure 40 , the subgroup where the selected (for signal transmission) Merge candidate is located is called the selected subgroup.
[0432] More specifically, if the selected Merge candidate is in the last subgroup, after the selected Merge candidate is derived, the Merge candidate list construction process is terminated, no re-ordering is performed, and the Merge candidate list is not changed; otherwise, the process is as follows:
[0433] After all the Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; the template matching cost of the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are re-ordered; finally, a new Merge candidate list will be obtained.
[0434] For both the encoder and the decoder, the template matching cost is derived as a function of T and RT, where T is the set of samples in the template and RT is the set of reference samples for the template.
[0435] When deriving the reference samples of the template for a Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.
[0436] The reference samples (RT) of the template for bidirectional prediction are derived by weighted averaging the reference samples (RT 0 ) of the template in reference list 0 and the reference samples (RT 1 ) of the template in reference list 1. RT = ((8 - w) * RT 0 + w * RT 1 + 4) >> 3 (2 - 32)
[0437] where the weight (8 - w) of the reference template in reference list 0 and the weight (w) of the reference template in reference list 1 are determined by the BCW index of the Merge candidate. BCW indices equal to {0, 1, 2, 3, 4} correspond to w equal to {-2, 3, 4, 5, 10} respectively.
[0438] If the local illumination compensation (LIC) flag of the Merge candidate is true, the LIC method is used to derive the reference samples of the template.
[0439] The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT.
[0440] The template size is 1. This means that the width of the left template and / or the height of the upper template is 1.
[0441] If the codec mode is MMVD, the merge candidates used to derive the base merge candidates are not reordered.
[0442] If the codec mode is GPM, the merge candidates used to derive the unidirectional prediction candidate list are not reordered.
[0443] 2.29. IBC with Extended Reference Region
[0444] An IBC reference region design that does not increase the current memory area required by ECM-3 is proposed, and the performance is tested.
[0445] Figure 41 The design is shown. In the figure, the blue square represents the current CTU, and the green square represents the CTUs that can be used by IBC reference. Specifically, assuming that W represents the maximum horizontal CTU index and the current CTU index is (m, n), for the codec units in the current CTU, the CTUs with indices (0, n)…(m, n) and (m-1, n)…(W, n) define the reference region that can be used by IBC.
[0446] One reason for having such a design is that in the current ECM, the left, upper, and upper-left CTUs are being used, so savings are needed. To achieve this, all CTUs to the right of the upper CTU in the upper CTU row (for the CTUs to be coded in the current CTU row) and all CTUs to the left of the current CTU in the current CTU row (for the CTUs to be coded in the next CTU row) must be retained. This means that this design does not increase the cache size required by the current ECM.
[0447] 2.30. IBC with Template Matching
[0448] The proposal also uses template matching with IBC for both the IBC Merge mode and the IBC AMVP mode.
[0449] The IBC-TM Merge list has been modified compared to that used by the regular IBC Merge mode such that candidates are selected according to a deduplication method with the motion distances between candidates in the regular TM Merge mode. The end-zero motion completion (which is meaningless for intra-coding) has been replaced by motion vectors to the left (-W, 0), top (0, -H), and upper-left (-W, -H) CUs, and then the list is executed using the left CU without deduplication if needed.
[0450] In the IBC-TM Merge mode, the selected candidates are refined using a template matching method before the RDO or decoding process. The IBC-TM Merge mode has been competing with the regular IBC Merge mode, and the TM Merge flag is signaled.
[0451] In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC Merge list. Each of these 3 selected candidates is refined using a template matching method and sorted according to their resulting template matching costs. Then usually only the first two are considered during the motion estimation process.
[0452] The template matching refinement for both the IBC-TM Merge and AMVP modes is very simple because the IBC motion vectors are constrained to be integers and within the reference region as Figure 42 shown. Thus, in the IBC-TM Merge mode, all refinements are performed with integer precision, and in the IBC-TM AMVP mode, all refinements are performed with integer or 4-pixel precision. In both cases, the refined motion vectors in each refinement step must comply with the constraints of the reference region.
[0453] 2.31. Reconstruction Reordering IBC (RR-IBC)
[0454] Screen content coding and decoding tools similar to Intra Block Copy (IBC) generate a predicted block by directly copying a previously coded and decoded reference region in the same picture. Symmetry is often observed in video content, especially in text character regions and computer-generated graphics in screen content sequences, as Figure 43 shown. Therefore, specific screen content coding and decoding tools that consider symmetry will effectively compress such videos.
[0455] The Reconstruction-Reordering IBC (RR-IBC) mode for screen content video coding and decoding is proposed. When RR-IBC is applied, the samples in the block are flipped according to the flip type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the predicted block is derived without flipping. On the decoder side, the reconstructed block is flipped to restore the original block.
[0456] Two flipping methods, horizontal flipping and vertical flipping, are supported for the blocks coded for RR-IBC. First, the syntax flag for the blocks coded for IBC AMVP is signaled, indicating whether the reconstruction is flipped. If it is flipped, another flag specifying the flipping type is further signaled. For IBC Merge, in the absence of syntax signaling, the flipping type is inherited from neighboring blocks. Considering the horizontal or vertical symmetry, the current block and the reference block are usually horizontally or vertically aligned. Therefore, when horizontal flipping is applied, the vertical component of the BV is not signaled and is presumed to be equal to 0. Similarly, when vertical flipping is applied, the horizontal component of the BV is not signaled and is presumed to be equal to 0.
[0457] To better utilize the symmetry property, a flipping-aware BV adjustment method is applied to refine the block vector candidates. For example, as Figure 44A and Figure 44B shown, (x nbr, y nbr ) and (x cur ,y cur ) represent the coordinates of the center samples of the neighboring block and the current block respectively, BV nbr and BV cur represent the BV of the neighboring block and the current block respectively. Instead of directly inheriting the BV from the neighboring block, the horizontal component of BV nbr (denoted as BV nbr h ) is calculated by adding the motion displacement to the horizontal component of BV cur when the neighboring block is coded with horizontal flipping, i.e., BV cur h = 2(x nbr - x cur ) + BV nbr h . Similarly, the vertical component of BV nbr (denoted as BV nbr v ) is calculated by adding the motion displacement to the vertical component of BV cur when the neighboring block is coded with vertical flipping, i.e., BV cur v = 2(y nbr - y cur ) + BV nbr v .
[0458] 2.32. Intra Prediction Template Matching
[0459] Intra Template Matching Prediction (Intra TMP) is a special intra prediction mode that copies the best prediction block from the reconstructed part of the current frame, and its L-shaped template matches the current template. For a predefined search range, the encoder searches the reconstructed part of the current frame for the template that is most similar to the current template and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side.
[0460] By matching the L-shaped temporary neighbor of the current block with another block in the predefined search area consisting of Figure 45 the prediction signal is generated:
[0461] R1: the current CTU;
[0462] R2: the top-left CTU;
[0463] R3: the upper CTU;
[0464] R4: the left CTU.
[0465] SAD is used as the cost function.
[0466] Within each region, the decoder searches for the template with the minimum SAD relative to the current one and uses its corresponding block as the prediction block.
[0467] The dimensions of all regions (SearchRange_w, SearchRange_h) are proportional to the block dimensions (BlkW, BlkH) with a fixed number of SADs per pixel. That is:
[0468] SearchRange_w = a * BlkW
[0469] SearchRange_h = a * BlkH
[0470] where 'a' is a constant that controls the gain / complexity trade-off. In fact, 'a' is equal to 5. The intra template matching tool is enabled for CUs with dimensions less than or equal to 64 in width and height. This maximum CU size for intra template matching is configurable.
[0471] When DIMD is not used for the current CU, the intra template matching prediction mode is signaled at the CU level by a dedicated flag.
[0472] 3. Problems
[0473] In the current design of IBC, the prediction of the current block is obtained using the samples in the current picture indicated by the block vector. The coding / decoding performance of IBC is very good for screen content videos with repetitive content. However, for natural content videos, due to different characteristics, the coding / decoding gain of IBC is much lower than that for screen content videos.
[0474] 4. Detailed solutions
[0475] The following embodiments should be considered as examples to explain general concepts. These embodiments should not be interpreted in a narrow way. In addition, these embodiments can be combined in any way. In the present disclosure, Intra Block Copy (IBC) may not be limited to the current IBC technology, but can be interpreted as a technology in which the reference (or prediction) block is obtained using the samples in the current strip / slice / subpicture / picture / other video units (e.g., CTU rows), excluding conventional intra prediction methods.
[0476] The reference lines can refer to the reconstructed samples of the rows and / or columns that are adjacent or not adjacent to the current block. The reference lines are used to derive the intra prediction of the current video unit via an interpolation filter along a specific direction, and the specific direction is determined by the intra prediction mode (e.g., conventional intra prediction with an intra prediction mode), or to derive the intra prediction of the current video unit by weighting the reference samples of the reference lines with a matrix or a vector (e.g., MIP).
[0477] In the present disclosure, IBC-GPM may refer to a coding / decoding tool that uses IBC in a video unit to obtain the prediction of at least one sub-division when the video unit is geometrically divided into more than one sub-division.
[0478] In the present disclosure, IBC-LIC may refer to a coding / decoding tool in which local illumination compensation is used to refine the video unit coded / decoded by IBC.
[0479] In the following discussion, IBC can be replaced by other coding / decoding tools (e.g., palette, intra template matching) that rely on the encoded / decoded / reconstructed information in the same region. Combination of intra block copy and intra prediction (CIBCIP or IBC - CIIP) 1. It is proposed to use a combination of Intra Block Copy (IBC) and Intra Prediction (CIB-CIP) to derive the prediction / reconstruction of a video unit obtained by fusing the IBC prediction signal and the intra prediction signal. a. In one example, P(x,y) = w ipl *IP 1 (x,y) + w ip2 *IP 2 (x,y) + … + w ipn *IP n (x, y) + w ibc1 *IBC 1 (x, y) + w ibc2 *IBC 2 (x, y) + … + w ibcm *IBC m (x, y), where P(x, y) is the generated predicted value, IP k is the prediction generated by intra prediction within the k-th frame, IBC j is the prediction generated by the j-th IBC, and w ipk and w ibcj are the corresponding weight values. b. In one example, intra prediction can be generated by angular intra prediction, DC, planar, cross-component prediction (CCLM), multi-model CCLM, left CCLM, top CCLM, etc. ⅰ. The intra prediction mode can be encoded or decoded using MPM or TIMD or DIMD or any other method to signal the intra prediction mode. ⅱ. A specific set of intra prediction modes can be allowed to be used in CIBCIP. c. In one example, the IBC Merge mode can be used. ⅰ. In one example, one or more BV candidates in the IBC Merge list can be allowed to be used in CIBCIP. ⅱ. In one example, one or more BV offsets can be used for CIBCIP. 1) In one example, the BV offset can be added to the BV candidate before the BV candidate is used to obtain the IBC prediction. 2) In one example, the BV offset can be signaled or derived. ⅲ. In one example, the IBC Merge mode can be at least one of the following: regular IBC Merge mode and IBC-MBVD Merge mode and IBC-TM Merge mode. d. In one example, the IBC AMVP mode can be used. ⅰ. In one example, one or more BV predicted values in the IBC AMVP list can be allowed to be used in CIBCIP. ⅱ. In one example, the block vector difference (BVD) used in CIBCIP can be signaled in the same way as the IBC mode. 1) Alternatively, the BVD can be not signaled but predefined. ⅲ. In one example, the BVD can be derived using the coding information. e. In one example, the merge index (mergeIdx) indicating the BV candidate in the IBC Merge list and / or the BVP index (bvpIdx) indicating the BV prediction value in the IBC AMVP list for obtaining the IBC prediction signal can be signaled. 1) In one example, the binarization or signaling method of the Merge index or BVP index can be the same as that in the IBC mode. 2) Alternatively, the Merge index or BVP index can be predefined, e.g., mergeIdx = 0 or mergeIdx = 1; bvpIdx = 0 or bvpIdx = 1. 3) Alternatively, the Merge index or BVP index can be derived using coding information. 4) Alternatively, the Merge index or BVP index can be derived using template matching (e.g., with the minimum template matching cost). f. In one example, the construction of the IBC Merge list or IBC AMVP list used in the CIBCIP mode can be the same as or different from that used in the IBC mode. g. In one example, the number (N) of BV candidates in the IBC Merge (or AMVP) list that can be used for CIBCIP is less than or equal to the number (M) of BV candidates in the IBC Merge (or AMVP) list that can be used for IBC. N is an integer greater than 0 and less than or equal to M. ⅰ. In one example, N = 1, or N = 2, or N = 3, or N = 4, or N = 5, or N = 6. ⅱ. In one example, the first N BV candidates in the IBC Merge (or AMVP) list can be used for CIBCIP. h. In one example, template matching can be used to derive / refine the BV, and the BV is used to obtain the IBC prediction signal. ⅰ. In one example, the BV offset can be derived using template matching, and the BV offset is added to the BV candidate in the IBC Merge list. ⅱ. In one example, the BVD can be derived using a template matching-based method. ⅲ. In one example, the BVD sign can be derived using a template matching-based method. ⅳ. In one example, the intra prediction mode or intra prediction method for obtaining the intra prediction signal can be used for template matching to derive / refine the BV. i. In one example, the BV list can be reordered before being used for CIBCIP. ⅰ. In one example, template matching or bilateral matching cost can be used for reordering. ⅱ. In one example, template matching or bilateral matching can be used during the construction of the BV list for CIBCIP. ⅲ. In one example, the BV list can refer to the IBC Merge list or the IBC AMVP list. ⅳ. In one example, the reordering method for the BV list for CIBCIP can be the same as the reordering method for IBC. ⅴ. Alternatively, the reordering method for the BV list for CIBCIP can be different from the reordering method for IBC. 1) In one example, the number (N 1 ) of BV candidates in the reordered BV list for CIBCIP can be less than or equal to the number (M 1 ) of BV candidates in the reordered BV list for the IBC mode. a) In one example, when the IBC Merge mode is used for CIBCIP, N 1 = 1 or 2 or 3 or 4. b) In one example, when the IBC AMVP is used for CIBCIP, N 1 = 1 or 2 or 3. j. In one example, intra prediction can refer to a conventional intra prediction method (e.g., intra prediction using 35 intra prediction modes in HEVC or intra prediction using 67 intra prediction modes in VVC), or other intra prediction methods that obtain a prediction block from samples in the current strip / slice / sub-picture / picture / other video unit (e.g., CU, PU, TU, CTU, CTU row) except for IBC. i. In one example, the intra prediction signal can be obtained using one or more predefined intra prediction modes. 1) In one example, the predefined intra prediction modes can refer to the planar mode, the DC mode, the horizontal mode, and the vertical mode. ⅱ. In one example, the intra prediction signal can be obtained using one or more of the most probable modes (MPM). ⅲ. In one example, the intra prediction signal can be obtained using an intra prediction mode that is derived using the block vector used to obtain the IBC prediction signal. ⅳ. In one example, the intra prediction signal can be obtained using an intra prediction mode that is derived using a template-based method intra prediction mode such as TIMD. v. In one example, an intra prediction signal may be obtained using an intra prediction mode, and the intra prediction mode is derived using neighboring samples or the gradient of neighboring samples (such as DIMD). vi. In one example, an intra prediction signal may be obtained using ISP. vii. In one example, an intra prediction signal may be obtained using MIP. viii. In one example, an intra prediction signal may be obtained using MRL. ix. In one example, the intra prediction signal may be IntraTMP (Intra Template Matching Prediction). 2. In one example, the weighting parameter for fusing the IBC prediction signal and the intra prediction signal may be signaled or derived. a. In one example, the weighting parameter may be signaled. i. In one example, a set of weighting parameters is constructed, and the index indicating the weighting parameter may be signaled. b. In one example, the weighting parameter may be derived using coding information. i. In one example, the coding information may refer to the coding mode of neighboring units. 1) In one example, the weighting parameter may depend on whether one or more neighboring units are coded using intra prediction or in the IBC mode. ii. In one example, the coding information may refer to the intra prediction mode used to obtain the intra prediction signal. iii. In one example, the coding information may refer to the block size or block dimension of the current video unit and / or neighboring video units. iv. In one example, the weighting parameter may be derived using a template matching method (e.g., with the minimum template matching cost). c. In one example, the weighting parameter may be predefined. 3. In one example, the reference region of CIBCIP may be less than or equal to the reference region of IBC. a. In one example, the reference region of CIBCIP may depend on the coding information of intra prediction. i. In one example, the reference region of CIBCIP may depend on the intra prediction mode. b. Alternatively, the reference region of CIBCIP may be different from the reference region of IBC. 4. Whether and / or how to apply the CIBCIP mode to a video unit may depend on the coding information, and the coding information may refer to: a. Whether IBC or intra prediction method is allowed, b. Block dimension and / or block size, i. In one example, when the block size (W×H) is less than or equal to a threshold (T), the block is allowed to be encoded and decoded using CIBCIP, where W and H represent the block width and block height, respectively. 1) In one example, T = 256, or 512, or 1024, or 2048, or 4096 c. Block depth, d. Slice / picture type and / or split tree type (single or dual tree, or local dual tree), i. In one example, CIBCIP can be applied only to I slices / I pictures. e. Temporal layer identifier, f. Block position, g. Color component. 5. In one example, an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs in the list can be used for intra prediction of CIBCIP. a. In one example, one or more conventional IPMs (e.g., 35 IPMs in HEVC or 67 IPMs in VVC) can be included in the list. i. In one example, one or more IPMs can be associated with whether a specific intra codec tool is used. 1) In one example, the specific intra codec tool can refer to ISP or MRL. b. In one example, one or more MIP modes can be included in the list. c. In one example, some or all of the MPMs can be included in the list. i. In one example, the MPM can be in the main MPM list and / or in the secondary MPM list. d. In one example, one or more derived IPMs can be included in the list. i. In one example, the derived IPM can use a method based on DIMD or TIMD. e. In one example, the IPMs in the list can be reordered. i. In one example, the template of the video unit can be used for reordering. ⅱ. In one example, a method based on TIMD can be used for reordering. ⅲ. In one example, one or more BVs can be used for reordering. f. In one example, one or more IPMs in the list can be replaced by one or more IPMs of the reference video unit located by BV before or after reordering. i. In one example, the Nth IPM in the list can be replaced, such as N = 0, or 1, or 2, or 3, or 4, or 5. 1) In one example, de-duplication is used before replacement, such as if the IPM located by BV is already in the list, the IPM located by BV is not used to replace an existing IPM. ii. Alternatively, one or more IPMs of the reference video unit located by BV can be added to the list. 1) In one example, de-duplication is used when adding IPMs. 6. In one example, the determination of generating an IBC prediction signal and / or generating an intra prediction signal and / or fusing the IBC and intra prediction signals for CIBCIP can be derived. a. In one example, the determination of generating an IBC prediction signal can refer to specific IBC codec tools and / or specific IBC Merge candidate types and / or AMVP / Merge indices. i. In one example, specific IBC codec tools can refer to RR-IBC, or IBC-TM, or IBC- MBVD, or IBC-GPM, or IBC-LIC. ii. In one example, specific IBC Merge candidate types can refer to normal Merge candidates, IBC-TM Merge candidates, or IBC-MBVD candidates, or RR-IBC candidates. b. In one example, the determination of generating an intra prediction signal can refer to specific intra prediction methods, and / or specific IPMs, and / or IPM indices indicating which IPM is used. c. In one example, the determination of fusing IBC and intra prediction signals can refer to weighting parameters, and / or the number of IBCs used for fusion and / or intra prediction signals. d. In one example, a "CIBCIP candidate" list can be constructed, and an index indicating a "CIBCIP candidate" in the list can be derived or signaled. i. In one example, "CIBCIP candidates" can include the determination of generating an IBC prediction signal, and / or the determination of generating an intra prediction signal, and / or the determination of fusing IBC and intra prediction signals. 7. In one example, the codec information used in CIBCIP can be used to codec subsequent video units. a. In one example, the block vector used to generate an IBC prediction signal can be considered as the block vector of normal IBC. i. In one example, the block vector can be added to the IBC HMVP table. ii. In one example, the block vector can be used to construct the IBC AMVP / Merge candidate list for subsequent video units. b. In one example, the IPM used to generate the intra prediction signal can be considered as the IPM of normal intra prediction. i. In one example, the IPM can be used to construct the MPM list of subsequent video units. ii. In one example, the IPM can be used for chrominance prediction. iii. In one example, the IPM can be propagated for video units not intra-coded. c. Alternatively, the coding information may not be used for subsequent video units. i. In one example, instead of the IPM used for the current video unit, a predefined IPM (e.g., DC or planar) can be used for subsequent video units. 8. In one example, CIBCIP may not be allowed to be used with one or more specific coding tools. a. In one example, the specific coding tool may refer to the IBC AMVP mode, or the IBC Merge mode, or the IBC-TM mode, or the IBC-MBVD mode, or the RR-IBC mode, or IBC-LIC, or IBC- GPM, or AMVR for IBC. b. Alternatively, CIBCIP can be used with one or more of the above coding tools. 9. In one example, whether to apply CIBCIP and / or how to apply CIBCIP may depend on the color format and / or color component. a. In one example, CIBCIP can be applied to all color components. b. In one example, when CIBCIP is applied to the chrominance component, the derivation of intra prediction may be different from that of intra prediction for the luminance component. i. In one example, intra prediction can be obtained using CCLM, or MMLM, or CCCM, or chrominance- DIMD, or chrominance-TIMD, or a combination of CCLM / MMLM / CCCM and the angular mode. c. In one example, whether to apply CIBCIP to the first component and / or how to apply CIBCIP to the first component may depend on whether CIBCIP is applied to the second component. i. In one example, the first component may refer to a chrominance component (e.g., Cb and / or Cr), and the second component may refer to a luminance component (e.g., Y). ii. In one example, the way of applying CIBCIP to the first component may be the same as that to the second component. 1) Alternatively, the way of applying CIBCIP to the first component may be different from that to the second component. a) In one example, the weighting parameters may be different. d. In one example, CIBCIP may be applied to the luminance component and not to the chrominance component. i. In one example, the luminance component may refer to Y in the YCbCr color space or G in the RGB color space. ii. In one example, the chrominance component may refer to Cb and / or Cr in the YCbCr color space or R and / or B in the RGB color space. Signaling of combination of intra block copy and intra prediction 10. The indication of the CIBCIP mode can be derived in real time. 11. The indication of the CIBCIP mode can be signaled conditionally, where the conditions may include: a. Whether IBC or the intra prediction method is allowed b. Block dimension and / or block size c. Block depth d. Slice / picture type and / or partition tree type (single or dual tree, or local dual tree) e. Temporal layer identification f. Block position g. Color component. 12. Whether the current block is encoded / decoded using the CIBCIP mode can be signaled using one or more syntax elements. a. In one example, the syntax element can be binarized using fixed-length coding / decoding or rounding-unary coding / decoding or unary coding / decoding or EG coding / decoding or coding / decoding flag. b. In one example, the syntax element can be coded / decoded by bypass coding / decoding or context coding / decoding. i. The context may depend on the coded / decoded information, such as block dimension, and / or block size, and / or slice / picture type, and / or information of neighboring blocks (adjacent or non-adjacent), and / or information of other coding / decoding tools for the current block, and / or information of the temporal layer. c. In one example, when the current video unit is encoded / decoded by IBC, the indication of the CIBCIP mode can be signaled. d. In one example, a syntax element may be signaled before or after an indication of IBC-TM mode, or IBC-MBVD mode, or RR-IBC mode, or IBC-LIC, or IBC-GPM. i. In one example, whether a syntax element is signaled and / or how a syntax element is signaled may depend on whether IBC mode, or IBC-TM mode, or IBC-MBVD mode, or RR-IBC mode, or IBC-LIC, or IBC-GPM is enabled for a video unit. e. In one example, one or more syntax elements may be signaled in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. f. In one example, a syntax element may be coded in a predictive manner. g. For example, a syntax element of a current block may be predicted from syntax elements of neighboring blocks. 13. In one example, for the above items, the RR-IBC or symmetric IBC method may be used in CIBCIP. 14. In one example, for the above items, the RR-IBC or symmetric IBC method may be disabled in CIBCIP. 15. In one example, the flip type of the IBC prediction section may be set to NO_FLIP (e.g., 0). Intra prediction with fused reference lines 16. It is proposed to fuse more than one reference line before performing intra prediction for a video unit. a. In one example, the number of reference lines (N) and the reference lines to be fused may be predefined, signaled in the bitstream, or derived on the fly, where N is an integer greater than 1. i. In one example, N may be predefined, such as N = 2 or N = 3. ⅱ. In one example, N may be signaled in a sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. ⅲ. In one example, N may be determined based on coding information. 1) In one example, the coding information may refer to block size, block dimension, or block position, or coding mode or intra prediction mode. ⅳ. In one example, which reference lines are used in the fusion may be indicated by a reference line index. 1) The reference line index may be predefined, signaled in the bitstream, or derived on the fly. 2) In one example, one of the reference line indices in the reference line index can be predefined, and the remaining N - 1 reference line indices are transmitted via signals. 3) In one example, one of the reference line indices in the reference line index can be transmitted via signals, and the remaining N - 1 reference line indices are predefined or derived on the fly. b. In one example, the N reference lines can be fused using a weighting parameter. i. In one example, L = W 1 *L 1 +W 2 *L 2 +…+W N-1 *L N , where L i and W i represent the i-th reference line and the corresponding weight used in the fusion, and L represents the reference line for the fusion used in intra prediction. 1) In another example, L = (W′ 1 *L 1 +W′ 2 *L 2 +…+W′ N-1 *L N ) >> Shift1, where W′ 1 +W′ 2 +…+W′ N-1 = 2^Shiftl. ⅱ. In one example, the weighting parameter can be predefined, or transmitted via signals in the bitstream, or derived on the fly. ⅲ. In one example, when the reference line L a is closer to the current video unit than the reference line L b , the corresponding weighting parameter W a or W' a can be equal to or greater than the W a of L b or W' b or W' b . ⅳ. In one example, when N = 2, W 1 = 3 / 4 and W 2 = 1 / 4, or W 1 = 5 / 8 and W 2 = 3 / 8, or W 1 = 1 / 2 and W 2 = 1 / 2. ⅴ. In one example, when N = 3, W 1 = 1 / 2 and W 2 = 1 / 4 and W3 = 1 / 4, or W 1 = 5 / 8 and W 2 = 1 / 4 and W 3 = 1 / 8. c. In one example, the number of samples in one reference line can be the same as the number of samples in another reference line. The following shows the fusion of samples: P(x, y) = W 1 * P(x 1 , y 1 ) 1 + W 2 * P(x 2 , y 2 ) 2 + … + W N-1 * P(x N-1 , y N-1 ) N-1 , where P(x i , y i ) i represents a sample in the i-th reference line. i. In one example, when fusing the upper part of the reference lines, samples in different reference lines with the same horizontal position can be fused. 1) In one example, x 1 = x 2 = … = x N-1 . ⅱ. In one example, when fusing the left part of the reference lines, samples in different reference lines with the same vertical position can be fused. 1) In one example, y 1 = y 2 = … = y N-1 . d. In one example, the number of samples in one reference line can be different from the number of samples in another reference line. Denote the number of samples in reference lines L m and L n as S m and S n . An example is shown as Figure 46 . i. In one example, the number of samples in the fused reference line can be the same as the number of samples in the reference line with the smallest number of samples. 1) In one example, to fuse the samples of the fused reference line, more samples can be used in reference line L m than in reference line L n , where the total number of samples in L m is greater than the total number of samples in L n . a) In one example, two or more sample points can be used in L m and one sample point can be used in L n . 2) In one example, the S m sample points in reference line L n can be used in a fused manner with the S n sample points in reference line L n . The example is depicted as Figure 47 . ⅱ. In one example, the number of sample points in the fused reference line can be the same as the number of sample points in the reference line with the largest number of sample points. 1) In one example, when S n is less than S m , the (S m -S n ) sample points can be filled or derived using the sample points in reference line S n and used for fusion. The example is depicted as Figure 48 . e. In one example, fusion of the reference lines can be performed after derivation of each reference line. i. Alternatively, fusion of the reference lines can be performed during derivation of the reference lines. f. In one example, the derivation of the reference sample points in the reference lines used for fusion can be the same as the derivation of the reference sample points not used for fusion of the reference lines. i. Alternatively, the derivation can be different. 1) In one example, how to handle unavailable reference sample points can be different. g. In one example, reference sample point filtering can be performed after fusion of the reference lines. i. Alternatively, reference sample point filtering can be performed before fusion of the reference lines. 1) In one example, reference sample point filtering can be different for different reference lines. 17. Whether and how to use the fused reference lines to derive the intra prediction of the current video unit can depend on the codec information. a. In one example, the codec information can refer to one or more intra prediction methods. ⅰ. In one example, the fused reference lines can be used for conventional intra prediction. ⅱ. In one example, the fused reference lines can be used in MRL / ISP / MIP / DIMD / TIMD. ⅲ. In one example, the fused reference lines can be used for conventional chroma intra prediction. iv. In one example, the fused reference lines can be used in the fusion for chroma LM and angle. v. In one example, the fused reference lines can be used as an additional method or to replace the intra-frame prediction method in the current frame. b. In one example, the codec information can refer to a color component. i. In one example, the fused reference lines can be used for intra-frame prediction of the luminance component. ii. In one example, the fused reference lines can be used for intra-frame prediction of the chroma component. c. In one example, the codec information can refer to an intra-frame prediction mode. i. In one example, when the DC mode is used, the fused reference lines can be used. ii. In one example, when the planar mode is used, the fused reference lines can be used. iii. In one example, when the angular intra-frame prediction mode is used, the fused reference lines can be used. iv. In one example, when the angular intra-frame prediction mode has a non-integer slope, the fused reference lines can be used. v. In one example, the fused reference lines can be used for more than one intra-frame prediction mode. d. In one example, the codec information can refer to the block size / dimension of the current block and / or neighboring blocks. i. In one example, when the block size of the current block is greater than or equal to T 1 then, the fused reference lines can be used. ii. In another example, when the block size of the current block is less than T 2 then, the fused reference lines can be used. e. In one example, the codec information can refer to the slice type, and / or the temporal layer and / or the QP. f. In one example, the fused reference lines may not be allowed for video units in different CTUs. 18. In one example, how the reference lines are fused, and whether and how the fused reference lines are used to derive the intra-frame prediction of the current video unit can be signaled in the bitstream. General aspects 19. In the above example, a video unit may refer to a video unit that may refer to a color component / sub-picture / strip / slice / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transformation unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transformation block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel. 20. Whether and / or how to apply the methods disclosed above may be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / other types of regions containing more than one sample or pixel. 21. Whether to apply the above method and / or how to apply the above method may depend on the following information: a. Messages signaled in DPS / SPS / VPS / PPS / APS / picture header / strip header / slice group header / largest coding unit (LCU) / coding unit (CU) / LCU row / LCU group / TU / PU block / video coding unit; b. The location of CU / PU / TU / block / video coding unit; c. The block size of the current block and / or the block size of neighboring blocks of the current block; d. The block shape of the current block and / or the block shape of neighboring blocks of the current block; e. The coding mode of the block, such as IBC or non-IBC inter-frame mode or non-IBC sub-block mode; f. Indication of color format (e.g., 4:2:0, 4:4:4); g. Coding tree structure; h. Strip / slice group type and / or picture type; i. Color component (e.g., may be applied only to chrominance component or luminance component); j. Temporal layer ID; k. Profile / level / layer of the standard.
[0480] As used herein, the term "video unit" or "video block" may be a sequence, picture, strip, slice, tile, sub-picture, coding tree unit (CTU) / coding tree block (CTB), CTU / CTB row, one or more coding units (CU) / coding blocks (CB), one or more CTU / CTB, one or more virtual pipeline data units (VPDU), a sub-region within a picture / strip / slice / tile. The term "reference line" may refer to reconstructed samples of a line and / or column that are adjacent or non-adjacent to a current block, and the reference line is used to derive the intra prediction of the current video unit via an interpolation filter along a specific direction, and the specific direction is determined by an intra prediction mode (e.g., conventional intra prediction with an intra prediction mode), or to derive the intra prediction of the current video unit by weighting the reference samples of the reference line with a matrix or a vector (e.g., MIP).
[0481] Figure 49 FIG. 4900 is a flow chart of a method 4900 for video processing according to an embodiment of the present disclosure. Method 4900 is implemented during the conversion between a video unit of a video and a bitstream of the video.
[0482] At block 4910, for the conversion between a video unit of a video and a bitstream of the video unit, whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to the video unit is determined based on at least one of the following: coding information of the video unit, color format, color component, or syntax element. In some embodiments, method 4900 further includes: determining a manner of applying CIBCIP to the video unit based on at least one of the following: coding information of the video unit, color format, color component, or syntax element.
[0483] At block 4920, if it is determined that CIBCIP is applied to the video unit, the prediction of the video unit is derived by combining an IBC prediction signal and an intra prediction signal.
[0484] At block 4930, the conversion is performed based on the prediction of the video unit. In some embodiments, the conversion may include encoding the video unit into a bitstream. Alternatively or additionally, the conversion may include decoding the video unit from the bitstream. In this way, the coding efficiency and coding performance can be improved.
[0485] In some embodiments, the coding information includes at least one of the following: whether IBC or an intra prediction mode is allowed, block dimension and / or block size, block depth, strip type, picture type, segmentation tree type, temporal layer identifier, block position, or color component.
[0486] In some embodiments, if the block size represented as WH is less than or equal to a threshold, the video unit is allowed to be encoded and decoded using CIBCIP, where W and H represent the block width and block height, respectively. In some embodiments, the threshold is one of the following: 256, 512, 1024, 2048, or 4096. In some embodiments, CIBCIP is only applied to I slices or I pictures.
[0487] In some embodiments, an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs from the IPM candidate list are used for intra prediction of CIBCIP. In some embodiments, one or more conventional IPMs are included in the IPM candidate list. In some embodiments, one or more conventional IPMs include 35 IPMs or 67 IPMs.
[0488] In some embodiments, one or more IPMs are associated with whether a particular intra codec tool is used. In some embodiments, the particular intra codec tool includes intra sub - partitioning (ISP) or multi - reference line intra prediction (MRL). In some embodiments, one or more matrix - weighted intra prediction (MIP) modes are included in the IPM candidate list.
[0489] In some embodiments, a portion of the most probable mode (MPM) is included in the IPM candidate list. Alternatively, all MPMs are included in the IPM candidate list. In some embodiments, the MPM is in at least one of the following: the primary MPM list or the secondary MPM list.
[0490] In some embodiments, at least one derived IPM is included in the IPM candidate list. In some embodiments, at least one derived IPM uses a method based on decoder - side intra mode derivation (DIMD) or template - based intra mode derivation (TIMD).
[0491] In some embodiments, the IPMs in the IPM candidate list are reordered. In some embodiments, the template of the video unit is used for reordering the IPMs. In some embodiments, a method based on TIMD is used for reordering the IPMs. In some embodiments, at least one block vector (BV) is used for reordering the IPMs.
[0492] In some embodiments, at least one IPM in the IPM candidate list before reordering is replaced by at least one IPM of a reference video unit located by the BV. Alternatively, at least one IPM in the IPM candidate list after reordering is replaced by at least one IPM of a reference video unit located by the BV.
[0493] In some embodiments, the Nth IPM in the IPM candidate list is replaced. In some embodiments, N = 0, or 1, or 2, or 3, or 4, or 5.
[0494] In some embodiments, de-duplication is used before replacement of at least one IPM. For example, if at least one IPM located by BV is already in the IPM candidate list, then at least one IPM located by BV is not used to replace at least one IPM in the IPM candidate list.
[0495] In some embodiments, at least one IPM of the reference video unit located by BV is added to the IPM candidate list. In some embodiments, de-duplication is used when adding at least one IPM of the reference video unit.
[0496] In some embodiments, the determination of at least one of the following is derived: generation of an IBC prediction signal, generation of an intra prediction signal, or a combination of IBC and intra prediction signals for CIBCIP. In some embodiments, the determination of the generation of an IBC prediction signal includes at least one of the following: a specific IBC codec tool, a specific IBC Merge candidate type, an advanced motion vector prediction (AMVP) index, or a Merge index. In some embodiments, the specific IBC codec tool includes at least one of the following: a reconstructed-reordered IBC (RR-IBC) mode, an IBC template matching (IBC-TM) mode, an IBC Merge mode with block vector difference (IBC-MBVD) mode, an IBC geometry partitioning mode (IBC-GPM) mode, or an IBC with local illumination compensation (IBC-LIC) mode. In some embodiments, the specific IBC Merge candidate type includes at least one of the following: a normal Merge candidate, an IBC-TM Merge candidate, an IBC-MBVD candidate, or an RR-IBC candidate.
[0497] In some embodiments, the determination of the generation of an intra prediction signal includes at least one of the following: a specific intra prediction method, a specific IPM, or an IPM index indicating which IPM is used. In some embodiments, the determination of the combination of IBC and intra prediction signals includes at least one of the following: a weighting parameter, the number of IBCs, or the intra prediction signal used for combination.
[0498] In some embodiments, a CIBCIP candidate list is constructed, and an index indicating a CIBCIP candidate in the CIBCIP candidate list is derived or indicated. In some embodiments, the CIBCIP candidate includes at least one of the following: the determination of the generation of an IBC prediction signal, the determination of the generation of an intra prediction signal, or the determination of the combination of IBC and intra prediction signals for CIBCIP.
[0499] In some embodiments, the coding and decoding information is used to code and decode at least one subsequent video unit. In some embodiments, the block vector used to generate the IBC prediction signal is regarded as the block vector of normal IBC. In some embodiments, the block vector is added to the IBC history-based motion vector prediction (HMVP) table.
[0500] In some embodiments, the block vector is used to construct an IBC AMVP candidate list for at least one subsequent video unit. Alternatively, the block vector is used to construct an IBC Merge candidate list for at least one subsequent video unit.
[0501] In some embodiments, the IPM used to generate the intra prediction signal is regarded as the IPM of normal intra prediction. In some embodiments, the IPM is used to construct an MPM list for at least one subsequent video unit. In some embodiments, the IPM is used for chrominance prediction. In some embodiments, the IPM is propagated for video units that are not intra-coded.
[0502] In some embodiments, the coding and decoding information is not used to code and decode at least one subsequent video unit. In some embodiments, a predefined IPM (e.g., DC or planar) is used for at least one subsequent video unit instead of the IPM used for the video unit.
[0503] In some embodiments, CIBCIP is not allowed to be used with at least one specific coding and decoding tool. Alternatively, CIBCIP is used with at least one specific coding and decoding tool. In some embodiments, the at least one specific coding and decoding tool includes at least one of the following: IBC AMVP mode, IBC Merge mode, IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, IBC-GPM, or AMVR for IBC.
[0504] In some embodiments, CIBCIP is applied to all color components. In some embodiments, if CIBCIP is applied to the chrominance component, the derivation of the intra prediction for the chrominance component is different from the derivation of the intra prediction for the luminance component. In some embodiments, the derivation of the intra prediction for the chrominance component uses at least one of the following: cross-component linear mode (CCLM) mode, multi-mode linear mode (MMLM), convolutional cross-component model (CCCM), chrominance decoder-side intra mode derivation (DIMD), template-based intra mode derivation for chrominance (TIMD), combination of CCLM and angular mode, combination of MMLM and angular mode, or combination of CCCM and angular mode.
[0505] In some embodiments, whether to apply CIBCIP to the first component and / or the manner of applying CIBCIP to the first component depends on whether to apply CIBCIP to the second component. In some embodiments, the first component is a chrominance component, and the second component is a luminance component. In some embodiments, the chrominance component includes at least one of Cb or Cr in the YCbCr color space, and the luminance component includes Y in the YCbCr color space.
[0506] In some embodiments, the manner of applying CIBCIP to the first component is the same as that of the second component. In some embodiments, the manner of applying CIBCIP to the first component is different from that of the second component. In some embodiments, the weighting parameters are different.
[0507] In some embodiments, CIBCIP is applied to the luminance component and not to the chrominance component. In some embodiments, the luminance component includes Y in the YCbCr color space. Alternatively, the luminance component includes G in the red-green-blue (RGB) color space.
[0508] In some embodiments, the chrominance component includes at least one of Cb or Cr in the YCbCr color space. Alternatively, the chrominance component includes at least one of R or B in the RGB color space.
[0509] In some embodiments, a syntax element is signaled before an indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-GPM. Alternatively, a syntax element is signaled after an indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-GPM.
[0510] In some embodiments, whether to signal a syntax element and / or the manner of signaling a syntax element depends on whether at least one of the following is enabled for a video unit: IBC mode, IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-GPM.
[0511] In some embodiments, a video unit includes at least one of the following: a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a picture, a sub-picture, a block, a sub-region within a block, or a region including more than one sample or pixel.
[0512] In some embodiments, an indication of whether and / or how to derive a prediction of a video unit by including an IBC prediction signal and an intra prediction signal is indicated at one of the following: sequence level, picture group level, picture level, slice level, or slice group level.
[0513] In some embodiments, an indication of whether and / or how to derive a prediction of a video unit by including an IBC prediction signal and an intra prediction signal is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.
[0514] In some embodiments, method 4900 further includes: determining whether and / or how to derive a prediction of a video unit by including an IBC prediction signal and an intra prediction signal based on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest coding unit (LCU), coding unit (CU), LCU row, LCU group, TU, PU block, video coding unit; a location of one of the following: CU, PU, TU, block, video coding unit; a block size of the current block and / or a block size of neighboring blocks of the current block; a block shape of the current block and / or a block shape of neighboring blocks of the current block; a coding mode of the video unit; an indication of color format; a coding tree structure; slice type; slice group type; picture type; color component; temporal layer identifier; a profile or level or layer of a standard.
[0515] In some embodiments, the transformation includes encoding a video unit into a bitstream.
[0516] In some embodiments, the transformation includes decoding a video unit from a bitstream.
[0517] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing of a video. The method includes: determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of the following: coding information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; and generating a bitstream based on the prediction of the video unit.
[0518] According to still some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. The method includes: determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of the following: codec information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining the IBC prediction signal and the intra prediction signal; generating a bitstream based on the prediction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0519] Embodiments of the present disclosure may be described according to the following items, and the features of these items may be combined in any reasonable manner.
[0520] Item 1. A method for video processing, including: for the conversion between a video unit of a video and a bitstream of the video unit, determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to the video unit based on at least one of the following: codec information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining the IBC prediction signal and the intra prediction signal; and performing the conversion based on the prediction of the video unit.
[0521] Item 2. The method according to Item 1, further including: determining the manner of applying CIBCIP to the video unit based on at least one of the following: codec information of the video unit, color format, color component, or syntax element.
[0522] Item 3. The method according to Item 1 or 2, wherein the codec information includes at least one of the following: whether IBC or the intra prediction mode is allowed, block dimension and / or block size, block depth, slice type, picture type, segmentation tree type, temporal layer identifier, block position, or color component.
[0523] Item 4. The method according to Item 3, wherein if the block size represented as W×H is less than or equal to a threshold, the video unit is allowed to be coded and decoded using CIBCIP, where W and H represent the block width and the block height, respectively.
[0524] Item 5. The method according to Item 4, wherein the threshold is one of the following: 256, 512, 1024, 2048, or 4096.
[0525] Item 6. The method according to Item 3, wherein CIBCIP is only applied to I slices or I pictures.
[0526] Item 7. The method according to Item 1 or 2, wherein an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs in the IPM candidate list are used for intra prediction of CIBCIP.
[0527] Item 8. The method according to Item 7, wherein one or more conventional IPMs are included in the IPM candidate list.
[0528] Item 9. The method according to Item 8, wherein one or more conventional IPMs include 35 IPMs or 67 IPMs.
[0529] Item 10. The method according to Item 7, wherein one or more IPMs are associated with whether a specific intra codec tool is used.
[0530] Item 11. The method according to Item 10, wherein the specific intra codec tool includes intra sub - partitioning (ISP) or multi - reference line intra prediction (MRL).
[0531] Item 12. The method according to Item 7, wherein one or more matrix - weighted intra prediction (MIP) modes are included in the IPM candidate list.
[0532] Item 13. The method according to Item 7, wherein a part of the most probable modes (MPMs) is included in the IPM candidate list, or all MPMs are included in the IPM candidate list.
[0533] Item 14. The method according to Item 13, wherein the MPMs are in at least one of the following: the primary MPM list or the secondary MPM list.
[0534] Item 15. The method according to Item 7, wherein at least one derived IPM is included in the IPM candidate list.
[0535] Item 16. The method according to Item 15, wherein at least one derived IPM uses a method based on decoder - side intra mode derivation (DIMD) or template - based intra mode derivation (TIMD).
[0536] Item 17. The method according to Item 7, wherein the IPMs in the IPM candidate list are reordered.
[0537] Item 18. The method according to Item 17, wherein a template of the video unit is used for reordering of the IPMs.
[0538] Item 19. The method according to Item 17, wherein a method based on TIMD is used for reordering of the IPMs.
[0539] Item 20. The method according to item 17, wherein at least one block vector (BV) is used for reordering of IPM.
[0540] Item 21. The method according to item 7, wherein at least one IPM in the IPM candidate list is replaced by at least one IPM of a reference video unit located by BV before reordering, or wherein at least one IPM in the IPM candidate list is replaced by at least one IPM of a reference video unit located by BV after reordering.
[0541] Item 22. The method according to item 21, wherein the Nth IPM in the IPM candidate list is replaced.
[0542] Item 23. The method according to item 22, wherein N = 0, or 1, or 2, or 3, or 4, or 5.
[0543] Item 24. The method according to item 22, wherein deduplication is used before replacement of at least one IPM.
[0544] Item 25. The method according to item 24, wherein if at least one IPM located by BV is already in the IPM candidate list, then at least one IPM located by BV is not used to replace at least one IPM in the IPM candidate list.
[0545] Item 26. The method according to item 21, wherein at least one IPM of a reference video unit located by BV is added to the IPM candidate list.
[0546] Item 27. The method according to item 26, wherein deduplication is used when adding at least one IPM of a reference video unit.
[0547] Item 28. The method according to item 1 or 2, wherein determination of at least one of the following is derived: generation of an IBC prediction signal, generation of an intra prediction signal, or a combination of IBC and intra prediction signals for CIBCIP.
[0548] Item 29. The method according to item 28, wherein determination of generation of an IBC prediction signal includes at least one of the following: a specific IBC codec tool, a specific IBC Merge candidate type, an advanced motion vector prediction (AMVP) index, or a Merge index.
[0549] Item 30. The method according to Item 29, wherein the specific IBC encoding / decoding tool includes at least one of the following: reconstruction-reordering IBC (RR-IBC) mode, IBC template matching (IBC-TM) mode, IBC Merge mode with block vector difference (IBC-MBVD) mode, IBC geometric partitioning mode (IBC-GPM) mode, or IBC with local illumination compensation (IBC-LIC) mode.
[0550] Item 31. The method according to Item 29, wherein the specific IBC Merge candidate type includes at least one of the following: normal Merge candidate, IBC-TM Merge candidate, IBC-MBVD candidate, or RR-IBC candidate.
[0551] Item 32. The method according to Item 28, wherein the determination of the generation of the intra prediction signal includes at least one of the following: a specific intra prediction method, a specific IPM, or an IPM index indicating which IPM is used.
[0552] Item 33. The method according to Item 28, wherein the determination of the combination of IBC and the intra prediction signal includes at least one of the following: a weighting parameter, the number of IBCs, or the intra prediction signal used for the combination.
[0553] Item 34. The method according to Item 28, wherein a CIBCIP candidate list is constructed, and an index indicating the CIBCIP candidate in the CIBCIP candidate list is derived or indicated.
[0554] Item 35. The method according to Item 34, wherein the CIBCIP candidate includes at least one of the following: the determination of the generation of the IBC prediction signal, the determination of the generation of the intra prediction signal, or the determination of the combination of IBC and the intra prediction signal for CIBCIP.
[0555] Item 36. The method according to Item 1 or 2, wherein the encoding / decoding information is used to encode / decode at least one subsequent video unit.
[0556] Item 37. The method according to Item 36, wherein the block vector used to generate the IBC prediction signal is regarded as the block vector of a normal IBC.
[0557] Item 38. The method according to Item 37, wherein the block vector is added to the IBC history-based motion vector prediction (HMVP) table.
[0558] Item 39. The method according to Item 37, wherein the block vector is used to construct an IBC AMVP candidate list for at least one subsequent video unit, or wherein the block vector is used to construct an IBCMerge candidate list for at least one subsequent video unit.
[0559] Item 40. The method according to Item 36, wherein the IPM used to generate an intra prediction signal is regarded as the IPM for normal intra prediction.
[0560] Item 41. The method according to Item 40, wherein the IPM is used to construct an MPM list for at least one subsequent video unit.
[0561] Item 42. The method according to Item 40, wherein the IPM is used for chrominance prediction.
[0562] Item 43. The method according to Item 40, wherein the IPM is propagated for a video unit that is not intra-coded.
[0563] Item 44. The method according to Item 1 or 2, wherein the coding information is not used to code at least one subsequent video unit.
[0564] Item 45. The method according to Item 44, wherein a predefined IPM is used for at least one subsequent video unit instead of the IPM for the video unit.
[0565] Item 46. The method according to Item 1 or 2, wherein CIBCIP is not allowed to be used with at least one specific coding tool, or wherein CIBCIP is used with at least one specific coding tool.
[0566] Item 47. The method according to Item 46, wherein at least one specific coding tool includes at least one of the following: IBC AMVP mode, IBC Merge mode, IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, IBC-GPM, or AMVR for IBC.
[0567] Item 48. The method according to Item 1 or 2, wherein CIBCIP is applied to all color components.
[0568] Item 49. The method according to Item 48, wherein if CIBCIP is applied to the chrominance component, the derivation of intra prediction for the chrominance component is different from the derivation of intra prediction for the luminance component.
[0569] Item 50. The method according to Item 49, wherein the derivation of intra prediction for the chrominance component uses at least one of the following: cross-component linear mode (CCLM) mode, multi-mode linear mode (MMLM), convolutional cross-component model (CCCM), chrominance decoder-side intra mode derivation (DIMD), template-based intra mode derivation for chrominance (TIMD), a combination of CCLM and angular mode, a combination of MMLM and angular mode, or a combination of CCCM and angular mode.
[0570] Item 51. The method according to Item 1 or 2, wherein whether CIBCIP is applied to the first component and / or the manner in which CIBCIP is applied to the first component depends on whether CIBCIP is applied to the second component.
[0571] Item 52. The method according to Item 51, wherein the first component is a chrominance component and the second component is a luminance component.
[0572] Item 53. The method according to Item 52, wherein the chrominance component includes at least one of Cb or Cr in the YCbCr color space, and wherein the luminance component includes Y in the YCbCr color space.
[0573] Item 54. The method according to Item 52, wherein the manner in which CIBCIP is applied to the first component is the same as that of the second component.
[0574] Item 55. The method according to Item 52, wherein the manner in which CIBCIP is applied to the first component is different from that of the second component.
[0575] Item 56. The method according to Item 55, wherein the weighting parameters are different.
[0576] Item 57. The method according to Item 1 or 2, wherein CIBCIP is applied to the luminance component and not applied to the chrominance component.
[0577] Item 58. The method according to Item 57, wherein the luminance component includes Y in the YCbCr color space, or wherein the luminance component includes G in the red-green-blue (RGB) color space.
[0578] Item 59. The method according to Item 57, wherein the chrominance component includes at least one of Cb or Cr in the YCbCr color space, or wherein the chrominance component includes at least one of R or B in the RGB color space.
[0579] Item 60. The method according to Item 1 or 2, wherein the syntax element is signaled before the indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC or IBC-GPM, or wherein the syntax element is signaled after the indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC or IBC-GPM.
[0580] Item 61. The method according to Item 60, wherein whether the syntax element is signaled and / or the way the syntax element is signaled depends on whether at least one of the following is enabled for the video unit: IBC mode, IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, or IBC-GPM.
[0581] Item 62. The method according to any one of Items 1 to 61, wherein the video unit includes at least one of the following: color component, prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec tree block (CTB), codec unit (CU), codec tree unit (CTU), CTU row, CTU group, slice, picture, sub-picture, block, sub-region within the block, or a region including more than one sample or pixel.
[0582] Item 63. The method according to any one of Items 1 to 61, wherein the indication of whether and / or how to derive the prediction of the video unit by including the IBC prediction signal and the intra prediction signal is indicated at one of the following: sequence level, picture group level, picture level, slice level, or picture group level.
[0583] Item 64. The method according to any one of Items 1 to 61, wherein the indication of whether and / or how to derive the prediction of the video unit by including the IBC prediction signal and the intra prediction signal is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or picture group header.
[0584] Item 65. The method according to any one of Items 1 to 61 further includes: determining whether and / or how to derive a prediction of a video unit by including an IBC prediction signal and an intra prediction signal based on at least one of the following: a message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, slice header, slice group header, largest coding unit (LCU), coding unit (CU), LCU row, LCU group, TU, PU block, video coding unit; the position of one of the following: CU, PU, TU, block, video coding unit; the block size of the current block and / or the block size of neighboring blocks of the current block; the block shape of the current block and / or the block shape of neighboring blocks of the current block; the coding mode of the video unit; an indication of the color format, the coding tree structure, the slice type, the slice group type, the picture type, the color component, the temporal layer identifier, the profile or level or layer of the standard.
[0585] Item 66. The method according to any one of Items 1 to 35, wherein the conversion includes encoding a video unit into a bitstream.
[0586] Item 67. The method according to any one of Items 1 to 35, wherein the conversion includes decoding a video unit from a bitstream.
[0587] Item 68. An apparatus for video processing includes a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 67.
[0588] Item 69. A non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to any one of Items 1 to 67.
[0589] Item 70. A non-transitory computer-readable recording medium stores a bitstream generated by a method executed by an apparatus for video processing for a video, wherein the method includes: determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of the following: coding information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; and generating a bitstream based on the prediction of the video unit.
[0590] Item 71. A method for storing a bitstream of a video, comprising: determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of: codec information of the video unit, color format, color component, or syntax element; if it is determined that CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining an IBC prediction signal and an intra prediction signal; generating a bitstream based on the prediction of the video unit; and storing the bitstream in a non-transitory computer-readable recording medium.
[0591] Example device
[0592] Figure 50 FIG. shows a block diagram of a computing device 5000 in which various embodiments of the present disclosure may be implemented. The computing device 5000 may be implemented as the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300), or may be included in the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).
[0593] It should be understood that Figure 50 the computing device 5000 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0594] As Figure 50 shown, the computing device 5000 includes a general-purpose computing device 5000. The computing device 5000 may include at least one or more processors or processing units 5010, a memory 5020, a storage unit 5030, one or more communication units 5040, one or more input devices 5050, and one or more output devices 5060.
[0595] In some embodiments, the computing device 5000 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 5000 may support any type of interface to the user (such as "wearable" circuitry, etc.).
[0596] The processing unit 5010 may be a physical processor or a virtual processor, and may implement various processes based on programs stored in the memory 5020. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 5000. The processing unit 5010 may also be referred to as a central processing unit (CPU), microprocessor, controller, or microcontroller.
[0597] The computing device 5000 generally includes various computer storage media. Such media may be any media accessible by the computing device 5000, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 5020 may be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 5030 may be any removable or non-removable media, and may include machine-readable media, such as memory, flash drives, magnetic disks, or other media that can be used to store information and / or data and can be accessed in the computing device 5000.
[0598] The computing device 5000 may also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 50 a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk may be provided. In this case, each drive may be connected to a bus (not shown) via one or more data media interfaces.
[0599] The communication unit 5040 communicates with another computing device via a communication medium. Additionally, the functionality of the components in the computing device 5000 can be implemented by a single computing cluster or multiple computing machines, which can communicate via a communication connection. Thus, the computing device 5000 can operate in a networked environment using a logical connection to one or more other servers, networked personal computers (PCs), or other general network nodes.
[0600] The input device 5050 can be one or more of a variety of input devices, such as a mouse, keyboard, trackball, voice input device, and so on. The output device 5060 can be one or more of a variety of output devices, such as a display, speaker, printer, and so on. With the aid of the communication unit 5040, the computing device 5000 can also communicate with one or more external devices (not shown), such as storage devices and display devices, the computing device 5000 can also communicate with one or more devices that enable a user to interact with the computing device 5000, or if needed, the computing device 5000 can also communicate with any device (such as a network card, modem, etc.) that enables the computing device 5000 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0601] In some embodiments, some or all of the components of the computing device 5000 can also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to achieve the functions described in this disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which will not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed via a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0602] In an embodiment of the present disclosure, the computing device 5000 can be used to implement video encoding / decoding. The memory 5020 may include one or more video codec modules 5025 having one or more program instructions. These modules are accessible and executable by the processing unit 5010 to perform the functions of the various embodiments described herein.
[0603] In an example embodiment of performing video encoding, the input device 5050 may receive video data as the input 5070 to be encoded. The video data may be processed, for example, by the video codec module 5025 to generate an encoded bitstream. The encoded bitstream may be provided as the output 5080 via the output device 5060.
[0604] In an example embodiment of performing video decoding, the input device 5050 may receive the encoded bitstream as the input 5070. The encoded bitstream may be processed, for example, by the video codec module 5025 to generate decoded video data. The decoded video data may be provided as the output 5080 via the output device 5060.
[0605] Although the present disclosure has been specifically shown and described with reference to preferred embodiments of the present disclosure, those skilled in the art will understand that various changes in form and detail may be made without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Therefore, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: For the conversion between a video unit of a video and the bitstream of the video unit, determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to the video unit based on at least one of the following: the codec information, color format, color component, or syntax element of the video unit; If it is determined that the CIBCIP is applied to the video unit, deriving the prediction of the video unit by combining the IBC prediction signal and the intra prediction signal; and Performing the conversion based on the prediction of the video unit.
2. The method according to claim 1, further comprising: Determining the manner of applying the CIBCIP to the video unit based on at least one of the following: the codec information, the color format, the color component, or the syntax element of the video unit.
3. The method according to claim 1 or 2, wherein the codec information includes at least one of the following: Whether IBC or intra prediction mode is allowed, Block dimension and / or block size, Block depth, Slice type, Picture type, Partition tree type, Temporal layer identifier, Block position, or Color component.
4. The method according to claim 3, wherein if the block size represented as W×H is less than or equal to a threshold, the video unit is allowed to be coded and decoded using the CIBCIP, where W and H represent the block width and block height, respectively.
5. The method according to claim 4, wherein the threshold is one of the following: 256, 512, 1024, 2048, or 4096.
6. The method according to claim 3, wherein the CIBCIP is only applied to I slices or I pictures.
7. The method according to claim 1 or 2, wherein an intra prediction mode (IPM) candidate list is constructed, and one or more IPMs in the IPM candidate list are used for the intra prediction of the CIBCIP.
8. The method according to claim 7, wherein one or more conventional IPMs are included in the IPM candidate list.
9. The method according to claim 8, wherein the one or more conventional IPMs include 35 IPMs or 67 IPMs.
10. The method according to claim 7, wherein the one or more IPMs are associated with whether a specific intra codec tool is used.
11. The method according to claim 10, wherein the specific intra codec tool includes intra sub-partitioning (ISP) or multi-reference line intra prediction (MRL).
12. The method according to claim 7, wherein one or more matrix weighted intra prediction (MIP) modes are included in the IPM candidate list.
13. The method according to claim 7, wherein a part of the most probable mode (MPM) is included in the IPM candidate list, or wherein all MPMs are included in the IPM candidate list.
14. The method according to claim 13, wherein the MPM is in at least one of the following: the primary MPM list or the secondary MPM list.
15. The method according to claim 7, wherein at least one derived IPM is included in the IPM candidate list.
16. The method according to claim 15, wherein the at least one derived IPM uses a manner based on decoder-side intra mode derivation (DIMD) or template-based intra mode derivation (TIMD).
17. The method according to claim 7, wherein the IPMs in the IPM candidate list are reordered.
18. The method according to claim 17, wherein the template of the video unit is used for the reordering of the IPMs.
19. The method according to claim 17, wherein a TIMD-based manner is used for the reordering of the IPMs.
20. The method according to claim 17, wherein at least one block vector (BV) is used for the reordering of the IPMs.
21. The method according to claim 7, wherein at least one IPM in the IPM candidate list is replaced by at least one IPM of a reference video unit located by the BV before the reordering, or wherein at least one IPM in the IPM candidate list is replaced by at least one IPM of a reference video unit located by the BV after the reordering.
22. The method according to claim 21, wherein the Nth IPM in the IPM candidate list is replaced.
23. The method according to claim 22, wherein N = 0, or 1, or 2, or 3, or 4, or 5.
24. The method according to claim 22, wherein deduplication is used before the replacement of the at least one IPM.
25. The method according to claim 24, wherein if at least one IPM located by the BV is already in the IPM candidate list, at least one IPM located by the BV is not used to replace at least one IPM in the IPM candidate list.
26. The method according to claim 21, wherein at least one IPM of the reference video unit located by the BV is added to the IPM candidate list.
27. The method according to claim 26, wherein deduplication is used when adding at least one IPM of the reference video unit.
28. The method according to claim 1 or 2, wherein the determination of at least one of the following is derived: the generation of the IBC prediction signal, the generation of the intra prediction signal, or the combination of the IBC and the intra prediction signal for the CIBCIP.
29. The method according to claim 28, wherein the determination of the generation of the IBC prediction signal includes at least one of the following: specific IBC codec tools, specific IBC Merge candidate types, advanced motion vector prediction (AMVP) index, or Merge index.
30. The method according to claim 29, wherein the specific IBC codec tools include at least one of the following: reconstruction-reordering IBC (RR-IBC) mode, IBC template matching (IBC-TM) mode, IBC Merge mode with block vector difference (IBC-MBVD) mode, IBC geometric partitioning mode (IBC-GPM) mode, or IBC with local illumination compensation (IBC-LIC) mode.
31. The method according to claim 29, wherein the specific IBC Merge candidate type includes at least one of the following: Normal Merge candidate, IBC-TM Merge candidate, IBC-MBVD candidate, or RR-IBC candidate.
32. The method according to claim 28, wherein the determination of the generation of the intra prediction signal includes at least one of the following: Specific intra prediction method, Specific IPM, or IPM index indicating which IPM is used.
33. The method according to claim 28, wherein the determination of the combination of the IBC and the intra prediction signal includes at least one of the following: Weighting parameter, Number of IBCs, or Intra prediction signal for the combination.
34. The method according to claim 28, wherein a CIBCIP candidate list is constructed, and wherein an index indicating a CIBCIP candidate in the CIBCIP candidate list is derived or indicated.
35. The method according to claim 34, wherein the CIBCIP candidate includes at least one of the following: The determination of the generation of the IBC prediction signal, The determination of the generation of the intra prediction signal, or The determination of the combination of the IBC and the intra prediction signal for the CIBCIP.
36. The method according to claim 1 or 2, wherein the codec information is used to encode and decode at least one subsequent video unit.
37. The method according to claim 36, wherein a block vector used to generate the IBC prediction signal is regarded as a block vector of a normal IBC.
38. The method according to claim 37, wherein the block vector is added to an IBC history-based motion vector prediction (HMVP) table.
39. The method according to claim 37, wherein the block vector is used to construct an IBC AMVP candidate list for the at least one subsequent video unit, or wherein the block vector is used to construct an IBC Merge candidate list for the at least one subsequent video unit.
40. The method according to claim 36, wherein an IPM used to generate the intra prediction signal is regarded as an IPM of normal intra prediction.
41. The method according to claim 40, wherein the IPM is used to construct an MPM list for the at least one subsequent video unit.
42. The method according to claim 40, wherein the IPM is used for chrominance prediction.
43. The method according to claim 40, wherein the IPM is propagated for video units not intra-coded.
44. The method according to claim 1 or 2, wherein the codec information is not used to encode and decode at least one subsequent video unit.
45. The method according to claim 44, wherein a predefined IPM is used for the at least one subsequent video unit instead of the IPM for the video unit.
46. The method according to claim 1 or 2, wherein the CIBCIP is not allowed to be used in combination with at least one specific codec tool, or wherein the CIBCIP is used in combination with the at least one specific codec tool.
47. The method according to claim 46, wherein the at least one specific codec tool includes at least one of the following: IBC AMVP mode, IBC Merge mode, IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC, IBC-GPM, or AMVR for IBC.
48. The method according to claim 1 or 2, wherein the CIBCIP is applied to all color components.
49. The method according to claim 48, wherein if the CIBCIP is applied to the chrominance component, the derivation of intra prediction for the chrominance component is different from the derivation of intra prediction for the luminance component.
50. The method according to claim 49, wherein the derivation of intra prediction for the chrominance component uses at least one of the following: Cross-Component Linear Mode (CCLM) mode, Multi-Mode Linear Mode (MMLM), Convolutional Cross-Component Model (CCCM), Decoder-Side Intra Mode Derivation for Chrominance (DIMD), Template-Based Intra Mode Derivation for Chrominance (TIMD), Combination of CCLM and angular mode, Combination of MMLM and angular mode, or Combination of CCCM and angular mode.
51. The method according to claim 1 or 2, wherein whether the CIBCIP is applied to the first component and / or the way the CIBCIP is applied to the first component depends on whether the CIBCIP is applied to the second component.
52. The method according to claim 51, wherein the first component is a chrominance component and the second component is a luminance component.
53. The method according to claim 52, wherein the chrominance component includes at least one of Cb or Cr in the YCbCr color space, and wherein the luminance component includes Y in the YCbCr color space.
54. The method according to claim 52, wherein the way the CIBCIP is applied to the first component is the same as that of the second component.
55. The method according to claim 52, wherein the way the CIBCIP is applied to the first component is different from that of the second component.
56. The method according to claim 55, wherein the weighting parameters are different.
57. The method according to claim 1 or 2, wherein the CIBCIP is applied to the luminance component and not applied to the chrominance component.
58. The method according to claim 57, wherein the luminance component includes Y in the YCbCr color space, or wherein the luminance component includes G in the Red-Green-Blue (RGB) color space.
59. The method according to claim 57, wherein the chrominance component includes at least one of Cb or Cr in the YCbCr color space, or wherein the chrominance component includes at least one of R or B in the RGB color space.
60. The method according to claim 1 or 2, wherein the syntax element is signaled before an indication of at least one of the following: IBC-TM mode, IBC-MBVD mode, RR-IBC mode, IBC-LIC or IBC-GPM, or wherein the syntax element is signaled after the indication of at least one of the following: the IBC-TM mode, the IBC-MBVD mode, the RR-IBC mode, the IBC-LIC or the IBC-GPM.
61. The method according to claim 60, wherein whether the syntax element is signaled and / or the way the syntax element is signaled depends on whether at least one of the following is enabled for the video unit: the IBC mode, the IBC-TM mode, the IBC-MBVD mode, the RR-IBC mode, the IBC-LIC, or the IBC-GPM.
62. The method according to any one of claims 1 to 61, wherein the video unit includes at least one of the following: a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a picture, a sub-picture, a block, a sub-region within a block, or a region including more than one sample or pixel.
63. The method according to any one of claims 1 to 61, wherein an indication of whether and / or how to derive the prediction of the video unit by including the IBC prediction signal and the intra prediction signal is indicated at one of the following: sequence level, group of pictures level, picture level, slice level, or slice group level.
64. The method according to any one of claims 1 to 61, wherein an indication of whether and / or how to derive the prediction of the video unit by including the IBC prediction signal and the intra prediction signal is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, or slice group header.
65. The method according to any one of claims 1 to 61, further comprising: determining whether and / or how to derive the prediction of the video unit by including the IBC prediction signal and the intra prediction signal based on at least one of the following: The message indicated in one of the following: DPS, SPS, VPS, PPS, APS, picture header, strip header, slice group header, largest coding unit (LCU), coding unit (CU), LCU row, LCU group, TU, PU block, video coding unit, The position of one of the following: CU, PU, TU, block, video coding unit, The block size of the current block and / or the block size of neighboring blocks of the current block, The block shape of the current block and / or the block shape of neighboring blocks of the current block, The coding mode of the video unit; An indication of the color format, Coding tree structure, Strip type, Slice group type, Picture type, Color component, Temporal layer identifier, Profile or level or tier of the standard.
66. The method according to any one of claims 1 to 35, wherein the conversion includes encoding the video unit into the bitstream.
67. The method according to any one of claims 1 to 35, wherein the conversion includes decoding the video unit from the bitstream.
68. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 67.
69. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 67.
70. A non-transitory computer-readable recording medium storing a bitstream generated by a method executed by an apparatus for video processing, wherein the method comprises: Determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of: the coding information, color format, color component, or syntax element of the video unit; If it is determined that the CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining the IBC prediction signal and the intra prediction signal; and Generating the bitstream based on the prediction of the video unit.
71. A method for storing a bitstream of a video, comprises: Determining whether to apply a combination of intra block copy (IBC) and intra prediction (CIBCIP) to a video unit of the video based on at least one of: the coding information, color format, color component, or syntax element of the video unit; If it is determined that the CIBCIP is applied to the video unit, deriving a prediction of the video unit by combining the IBC prediction signal and the intra prediction signal; Generating the bitstream based on the prediction of the video unit; and Storing the bitstream in a non-transitory computer-readable recording medium.