Method and device for video processing and medium
By decoder side refining and sample point refining processing of the motion vector of the video block, the problem of insufficient video encoding and decoding efficiency and quality in the prior art is solved, and a more efficient and higher quality encoding and decoding effect is achieved.
Patent Information
- Application Number
- CN202380074131.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-20
- Filing Date
- 2023-10-19
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video encoding and decoding technologies have shortcomings in improving encoding and decoding efficiency and quality, especially in processing motion vectors and predictions of video blocks.
A video processing method is proposed, including decoder-side motion vector refinement (DMVR) processing on the motion vector of the video block, and sample point refinement on the motion-compensated predicted video block. This method performs deduplication checks by checking the motion information and codec information of multiple Merge candidates to improve the codec efficiency and quality.
Through DMVR and sample point refinement processing, the efficiency and quality of video encoding and decoding are significantly improved, and have better performance than conventional solutions.
Smart Images

Figure CN120077658A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to video encoding and decoding. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the encoding and decoding efficiency and quality of video encoding and decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: for the conversion between a current video block of a video and the bitstream of the video, obtaining a set of motion vectors for the current video block, where the current video block is encoded and decoded using sub-block-based encoding and decoding tools; applying a decoder-side motion vector refinement (DMVR) process to the set of motion vectors; and performing the conversion based on the application.
[0005] According to the method of the first aspect of the present disclosure, the DMVR process is used for video blocks encoded and decoded using sub-block-based encoding and decoding tools. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency and quality.
[0006] In a second aspect, another method for video processing is proposed. The method includes: for the conversion between a current video block of a video and the bitstream of the video, obtaining a motion-compensated prediction of the current video block, where the current video block is encoded and decoded using an intra-template matching mode or an intra-block copy (IBC) mode; applying a sample refinement process to the motion-compensated prediction; and performing the conversion based on the application.
[0007] According to the method of the second aspect of the present disclosure, the sample refinement process is used for video blocks encoded and decoded using the intra-template matching mode or the IBC mode. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency and quality.
[0008] In a third aspect, another method for video processing is proposed. The method includes: obtaining a plurality of Merge candidates for a current video block of a video for conversion between the current video block and a bitstream of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first codec information and motion information of the plurality of Merge candidates, where the first codec information is different from the motion information; and performing the conversion based on the application.
[0009] According to the method of the third aspect of the present disclosure, the duplicate removal check is performed by considering motion information and additional codec information different from the motion information. Compared with conventional solutions, the proposed method can advantageously improve codec efficiency and codec quality.
[0010] In a fourth aspect, an apparatus for video processing is proposed. The apparatus includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0011] In a fifth aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of the present disclosure.
[0012] In a sixth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method includes: obtaining a set of motion vectors for a current video block of the video, where the current video block is decoded using a sub-block-based codec tool; applying a decoder-side motion vector refinement (DMVR) process to the set of motion vectors; and generating a bitstream based on the application.
[0013] In a seventh aspect, a method for storing a bitstream of a video is proposed. The method includes: obtaining a set of motion vectors for a current video block of the video, where the current video block is decoded using a sub-block-based codec tool; applying a decoder-side motion vector refinement (DMVR) process to the set of motion vectors; generating a bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
[0014] In an eighth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method includes: obtaining a motion-compensated prediction of a current video block of the video, where the current video block is decoded using an intra-template matching mode or an intra-block copy (IBC) mode; applying a sample refinement process to the motion-compensated prediction; and generating a bitstream based on the application.
[0015] In a ninth aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining a motion-compensated prediction of a current video block of the video, the current video block being encoded or decoded using an intra-template matching mode or an intra-block copy (IBC) mode; applying a sample refinement process to the motion-compensated prediction; generating a bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
[0016] In a tenth aspect, another non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing. The method includes: obtaining a plurality of Merge candidates for a current video block of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first coding information and motion information of the plurality of Merge candidates, the first coding information being different from the motion information; and generating a bitstream based on the application.
[0017] In an eleventh aspect, a method for storing a bitstream of a video is provided. The method includes: obtaining a plurality of Merge candidates for a current video block of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first coding information and motion information of the plurality of Merge candidates, the first coding information being different from the motion information; generating a bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
[0018] The present invention content is provided to introduce a selection of concepts further described below in the detailed implementation in a simplified form. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The above and other objects, features, and advantages of the example embodiments of the present disclosure will become more apparent through the following detailed description with reference to the accompanying drawings. In the example embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0020] Figure 1 A block diagram showing an example video codec system according to some embodiments of the present disclosure is shown;
[0021] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure is shown;
[0022] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure is shown;
[0023] Figure 4 The positions of spatial Merge candidates are shown;
[0024] Figure 5 shows candidate pairs considered for redundancy check of spatial merge candidates;
[0025] Figure 6 The motion vector scaling for the temporal Merge candidate is shown;
[0026] Figure 7 The candidate position C for the time domain Merge candidate is shown 0 and C 1 ;
[0027] Figure 8 The MMVD search point is shown;
[0028] Figure 9 shows the extended CU area used in BDOF;
[0029] Figure 10 A symmetric MVD pattern is shown;
[0030] Figure 11 An affine motion model based on control points is shown;
[0031] Figure 12 The affine MVF of each sub-block is shown;
[0032] Figure 13 The positions of the inherited affine motion prediction values are shown;
[0033] Figure 14 Control point motion vector inheritance is shown;
[0034] Figure 15 The positions of candidate positions for constructing the affine Merge pattern are shown;
[0035] Figure 16 is a diagram of the use of motion vectors for the proposed combination method;
[0036] Figure 17 Sub-block MV VSB and pixel Δv(i, j) are shown;
[0037] Figure 18A shows the spatial neighboring blocks used by ATVMP;
[0038] Figure 18B The sub-CU motion field is derived by applying motion displacements from spatial neighbors and scaling motion information from corresponding co-located sub-CUs;
[0039] Figure 19 shows the extended CU area used in BDOF;
[0040] Figure 20 Shows motion vector refinement on the decoding side;
[0041] Figure 21 Shows the top neighboring block and the left neighboring block used in CIIP weight derivation;
[0042] Figure 22 Shows an example of GPM partitioning grouped at the same angle;
[0043] Figure 23 Shows the unidirectional prediction MV selection for the geometric partitioning mode;
[0044] Figure 24 Shows the generation of the bending weight w 0 using the geometric partitioning mode;
[0045] Figure 25 Shows the spatial neighboring blocks for deriving spatial Merge candidates;
[0046] Figure 26 Shows performing template matching on the search area around the initial MV;
[0047] Figure 27 Shows the diamond area in the search area;
[0048] Figure 28 Shows the frequency responses of the interpolation filter and the VVC interpolation filter at the half - pixel phase;
[0049] Figure 29 Shows the template and the reference sample points of the template in the reference picture;
[0050] Figure 30 Shows the reference sample points of the template and the template for the block with sub - block motion using the motion information of the sub - blocks of the current block;
[0051] Figure 31 Shows the filling candidates for replacing the zero vectors in the IBC list.
[0052] Figure 32 Shows the IBC reference region depending on the current CU position;
[0053] Figure 33 Shows the reference region for IBC when CTU(m, n) is encoded / decoded. The blue block represents the current CTU; the green block represents the reference region; and the white block represents the invalid reference region;
[0054] Figure 34 Shows the first HPT and the second HPT;
[0055] Figure 35Shows the spatial neighbors for deriving affine Merge candidates / AMVP candidates;
[0056] Figure 36 Shows the constructed affine Merge candidates / AMVP candidates of the first type from non - adjacent neighbors;
[0057] Figure 37 Shows the low - frequency non - separable transform (LFNST) process;
[0058] Figure 38 Shows the SBT position, type, and transform type;
[0059] Figure 39 Shows the ROI for LFNST 16;
[0060] Figure 40 Shows the ROI for LFNST 8;
[0061] Figure 41 Shows the discontinuity measurement;
[0062] Figure 42 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure;
[0063] Figure 43 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure;
[0064] Figure 44 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure; and
[0065] Figure 45 Shows a block diagram of a computing device in which various embodiments of the present disclosure may be implemented.
[0066] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0067] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, without implying any limitation to the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than those described below.
[0068] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0069] As used in this disclosure, the terms "one embodiment", "an embodiment", "an exemplary embodiment", etc. indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment must include the specific feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in connection with an exemplary embodiment, it is contended that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0070] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0071] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Exemplary Environment
[0072] Figure 1 is a block diagram showing an exemplary video codec system 100 that can utilize the techniques of this disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0073] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.
[0074] Video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded pictures and associated data. The encoded pictures are the encoded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 through the network 130A. The encoded video data may also be stored on the storage medium / server 130B for access by the destination device 120.
[0075] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.
[0076] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0077] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.
[0078] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 the example, the video encoder 200 includes a plurality of functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0079] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0080] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0081] In addition, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for the purpose of explanation, these components are shown separately in Figure 2 the examples.
[0082] The segmentation unit 201 may segment a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0083] The mode selection unit 203 may select, for example, one coding / decoding mode (intra coding / decoding or inter coding / decoding) among multiple coding / decoding modes based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block to be used as a reference picture. In some examples, the mode selection unit 203 may select a Combined Intra and Inter Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0084] To perform inter prediction on the current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.
[0085] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a part of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to parts of a picture composed of macroblocks independent of macroblocks in the same picture.
[0086] In some examples, the motion estimation unit 204 can perform uni-directional prediction on the current video block, and the motion estimation unit 204 can search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 can then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial displacement between the current video block and the reference video block. The motion estimation unit 204 can output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block based on the reference video block indicated by the motion information of the current video block.
[0087] Alternatively, in other examples, the motion estimation unit 204 can perform bi-directional prediction on the current video block. The motion estimation unit 204 can search the reference pictures in list 0 to find one reference video block for the current video block, and can also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 can then generate a plurality of reference indices and a plurality of motion vectors, where the plurality of reference indices indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial displacements between the plurality of reference video blocks and the current video block. The motion estimation unit 204 can output the plurality of reference indices and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 can generate a predicted video block of the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0088] In some examples, the motion estimation unit 204 can output a complete set of motion information for use in the decoding process of the decoder. Alternatively, in some embodiments, the motion estimation unit 204 can signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 can determine that the motion information of the current video block is similar enough to the motion information of a neighboring video block.
[0089] In one example, the motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0090] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0091] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0092] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.
[0093] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0094] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0095] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0096] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0097] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transformed coefficient video block respectively to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0098] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the blocking artifacts in the video block.
[0099] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 can perform one or more entropy coding operations to generate the entropy-coded data and output a bitstream including the entropy-coded data.
[0100] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0101] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3 an example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0102] In Figure 3 an example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.
[0103] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, which includes motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and the Merge mode. AMVP is used, including deriving several most likely candidates based on data from adjacent PBs and reference pictures. Motion information generally includes horizontal motion vector displacement values and vertical motion vector displacement values, one or two reference picture indices, and, in the case of a prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, the "Merge mode" can refer to deriving motion information from spatially adjacent blocks or temporally adjacent blocks.
[0104] The motion compensation unit 302 can generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.
[0105] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of a video block to calculate interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0106] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of an encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be the entire picture or can also be a region of the picture.
[0107] The intra prediction unit 303 can use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially adjacent blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0108] The reconstruction unit 306 can obtain the decoded block by, for example, adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the buffer 307 also generates the decoded video for presentation on a display device.
[0109] Some exemplary embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. In addition, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. In addition, although some embodiments describe the video coding steps in detail, it should be understood that the corresponding decoding steps of the decoding will be implemented by the decoder. In addition, the term video processing includes video coding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another compression format or at different compression bit rates. 1. Brief Overview Embodiments of the present disclosure relate to video coding techniques. Specifically, it is about coding and decoding techniques for transformation, screen content coding and decoding, and local illumination compensation in image / video coding. It can be applied to existing video coding standards such as HEVC, VVC, ECM, etc. It is also applicable to future video coding and decoding standards or video codecs. 2. Introduction Video coding standards have mainly evolved through the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, ISO / IEC developed MPEG-1 and MPEG-4 Visual, and the two organizations jointly developed H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) as well as H.265 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure, which utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. JVET meetings are held quarterly simultaneously, and in the April 2018 JVET meeting, the new video coding standard was officially named Versatile Video Coding (VVC), and the first version of the VVC Test Model (VTM) was released at this time. Then the VVC working draft and the test model VTM are updated after each meeting. The VVC project achieved Feature Complete (FDIS) at the July 2020 meeting. In January 2021, JVET established the Exploration Experiments (EE), aiming at enhanced compression efficiency beyond VVC capabilities using novel conventional algorithms. Soon after, ECM was built as a common software library for the long-term exploration work towards the next-generation video coding standard. 2.1. Inter-frame prediction coding tools For each inter-frame prediction CU, the motion parameters include the motion vector, the reference picture index and the reference picture list use index, and additional information required for the new decoding features of VVC that will be used for inter-frame prediction sample generation. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded using the skip mode, the CU is associated with a PU and has no significant residual coefficients, no coded motion vector difference or reference picture index. A Merge mode is specified, whereby the motion parameters for the current CU are obtained from neighboring CUs (including spatial candidates and temporal candidates, and additional lists introduced in VVC). The Merge mode can be applied to any inter-frame prediction CU, not just the skip mode. An alternative to the Merge mode is the explicit transmission of the motion parameters, where the motion vector, the corresponding reference picture index and the reference picture list use flag for each reference picture list, and other required information are explicitly signaled for each CU. In addition to the inter-frame coding features in HEVC, VVC includes multiple new and refined inter-frame prediction coding tools listed as follows: – Extended Merge prediction; – Merge Mode with MVD (MMVD); – Symmetric MVD (SMVD) signaling; – Affine motion compensation prediction; – Sub-block based temporal motion vector prediction (SbTMVP); – Adaptive motion vector resolution (AMVR); – Motion field storage: 1 / 16 luma sample MV storage and 8×8 motion field compression; – Bi-directional prediction with CU-level weights (BCW); – Bidirectional optical flow (BDOF); – Decoder-side motion vector refinement (DMVR); – Geometric partition mode (GPM); – Combined inter and intra prediction (CIIP). The following text provides details on those inter prediction methods specified in VVC. 2.1.1. Extended Merge prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: 1) Spatial MVP from spatially neighboring CUs; 2) Temporal MVP from co-located CUs; 3) History-based MVP from the FIFO table; 4) Pairwise-averaged MVP; 5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coded in the Merge mode, the index of the best Merge candidate is encoded using truncated unary binary (TU). The first binary bit of the Merge index is coded using context, and bypass coding is used for the other binary bits. The derivation process for each type of Merge candidate is provided in this session. Similar to what was done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain sized region. 2.1.1.1. Spatial candidate derivation The derivation of spatial Merge candidates in VVC is the same as in HEVC, except for swapping the positions of the first two Merge candidates. Among the candidates at the positions depicted in Figure 4 select up to four Merge candidates. The derivation order is B 0 , A 0 , B 1 , A 1 and B 2 . Only when position B 0, A 0 , B 1 , A 1 When one or more CUs of A (e.g., because it belongs to another strip or slice) are unavailable or are intra-coded, position B 2 is considered. After a candidate at position A 1 is added, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, so as to improve the coding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Instead, only the pairs linked by the arrows in Figure 5 are considered, and a candidate is added to the list only if the corresponding candidates for the redundancy check do not have the same motion information. 2.1.1.2. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, when deriving the temporal Merge candidate at this time, the scaled motion vector is derived based on the co-located CUs belonging to the co-located reference picture. The list of reference pictures to be used for deriving the co-located CUs is explicitly signaled in the slice header. As shown by the dashed line in Figure 6 , the scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. As depicted in Figure 7 , position C 0 for the temporal candidate is selected between the candidates 1 and C. If the CU at position C 0 is unavailable, intra-coded, or outside the current row of the CTU, position C 1 is used. Otherwise, position C 0 is used to derive the temporal Merge candidate. Merge candidate derivation based on history After the spatial MVP and TMVP, the Merge candidates of the history-based MVP (HMVP) are added to the Merge list. In this method, the motion information of the previously coded blocks is stored in a table and used as the MVP for the current CU. The table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a CU that is non-sub-block inter-coded, the associated motion information is added to the last entry of the table as a new HMVP candidate. The size S of the HMVP table is set to 6, which indicates that at most 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if the same HMVP exists in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are shifted forward. HMVP candidates can be used in the Merge candidate list construction process. Check the latest several HMVP candidates in the table in order and insert them into the candidate list after the TMVP candidates. Apply a redundancy check to the HMVP candidates for both spatial and temporal Merge candidates. To reduce the number of redundancy check operations, the following simplifications are introduced: 1. The number of HMPV candidates for Merge list generation is set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list, and M indicates the number of available HMVP candidates in the table. 2. Once the total number of available Merge candidates reaches one less than the maximum allowed Merge candidates, the Merge candidate list construction process from HMVP is terminated. 2.1.1.3. Pairwise-average Merge candidate derivation Pairwise-average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices into the Merge candidate list. The average motion vectors are calculated separately for each reference list. If both motion vectors are available in a list, the two motion vectors are averaged even if they point to different reference images; if only one motion vector is available, that one vector is used directly; if no motion vector is available, the list is kept invalid. When the Merge list is not full after adding pairwise-average Merge candidates, zero MVPs are inserted at the end until the maximum Merge candidate number is reached. 2.1.1.4. Merge estimation region The Merge Estimation Region (MER) allows for the independent derivation of the Merge candidate list for a CU within the same Merge Estimation Region (MER). Candidate blocks within the same MER as the current CU are not included for the generation of the Merge candidate list for the current CU. Additionally, the update process for the history-based motion vector prediction candidate list is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2ParMrgLevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected on the encoder side and signaled in the sequence parameter set as log2_parallel_merge_level_minus2. 2.1.2. Merge Mode with MVD (MMVD) In addition to the Merge mode, in the case where implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge Mode with Motion Vector Difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and the Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after selecting the Merge candidate, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, an index for specifying the motion size, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The Merge candidate flag is signaled to specify which one is used. The distance index specifies the motion size information and indicates a predefined offset from the starting point. As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1. Table 1 - Relationship between Distance Index and Predefined Offset The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate the four directions shown in Table 2. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where two of the lists point to the same side of the current picture (i.e., both of the two referenced POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 2 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV with two MVs pointing to different sides of the current picture (i.e., one referenced POC is greater than the POC of the current picture and the other referenced POC is less than the POC of the current picture), the symbol in Table 2 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value. Table 2 - Signs of MV Offsets Specified by the Direction Index Direction Index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.1.2.1. Bi - directional Prediction with CU - level Weights (BCW) In HEVC, a bi - directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi - directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (2 - 1) Five weights are allowed in weighted - average bi - directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi - directional prediction CU, the weight w is determined in one of two ways: 1) For non - Merge CUs, the weight index is signaled after the motion vector difference; 2) For Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (w ∈ {3, 4, 5}) are used. – At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low - latency picture, unequal weights are only conditionally checked for 1 - pixel and 4 - pixel motion vector precisions. – When combined with affine, unequal weights are only conditionally checked for 1 - pixel and 4 - pixel motion vector precisions. – When the two reference pictures in bi - directional prediction are the same, the unequal weights are only conditionally checked. – When specific conditions are met, the unequal weights are not searched, which depends on the POC distance between the current picture and its reference pictures, the coding - decoding QP, and the temporal level. The BCW weight index is coded using a context - coded binary bit, followed by a bypass - coded binary bit. The first context - coded binary bit indicates whether equal weights are used; and if unequal weights are used, the bypass - coded binary bits are used to signal additional binary bits to indicate the use of unequal weights. Weighted prediction (WP) is a coding - decoding tool supported by the H.264 / AVC and HEVC standards for efficient coding - decoding of video content in fading scenarios. Support for WP is also added in the VVC standard. WP allows signaling of weighted parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled, and w is presumed to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control - point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded using the CIIP mode, the BCW index of the current CU is set to 2, for example, equal weights. 2.1.2.2. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. The BDOF, previously called BIO, is included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version, which requires less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bi - directional prediction signal of a CU at the 4×4 sub - block level. BDOF is applied to a CU if the CU meets all of the following conditions: – The CU is coded using the "true" bi - directional prediction mode, i.e., one of the two reference pictures is before the current picture in display order, and the other reference picture is after the current picture in display order. – The distances from two reference pictures to the current picture (i.e., the POC differences) are the same. – Both of the two reference pictures are short-term reference pictures. – The CU is coded or decoded without using the affine mode or the ATMVP Merge mode. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current CU. – The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of the object is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bi-predictive sample values in the 4×4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the differences between two neighboring samples, and i.e., where I (k) (i,j) is the sample value at the coordinate (i,j) of the prediction signal in list k (k = 0,1), and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luma bit depth (bitDepth). Then, the auto-correlations and cross-correlations of the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated as: where where Ω is a 6×6 window around the 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then the motion refinement (v x ,v y ) is derived using the cross-correlation terms and auto-correlation terms with the following equation: where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on motion refinement and gradient, the following adjustment is calculated for each sample point in a 4×4 sub-block: Finally, the BDOF sample points of the CU are calculated by adjusting the bi-predicted sample points as follows: pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + ο offset ) >> shift(2 - 7) These values are selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient values, some predicted sample points I (k) (i,j) in list k (k = 0,1) outside the current CU boundary need to be generated. As Figure 9 depicted, BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating the predicted sample points outside the boundary, the predicted sample points in the extended region (white positions) are generated by directly obtaining the reference sample points at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the predicted sample points within the CU (gray positions). These extended sample point values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample points and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of the CU is greater than 16 luma sample points, it is divided into sub-blocks with a width and / or height equal to 16 luma sample points, and the sub-block boundaries are considered as the CU boundaries during the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 predicted sample points and the L1 predicted sample points is less than the threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H >> 1), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD calculated during the DVMR process between the initial L0 predicted sample points and the L1 predicted sample points is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., luma_weight_lx_flag is 1 for either of the two reference pictures, then BDOF is also disabled. When a CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, BDOF is also disabled. 2.1.2.3. Symmetric MVD Coding (SMVD) In VVC, in addition to normal uni-directional prediction and bi-directional prediction mode MVD signaling, a symmetric MVD mode for bi-directional prediction MVD signaling is applied. In the symmetric MVD mode, the motion information including the reference picture indices for both list 0 and list 1 and the MVD for list 1 is not signaled, but is derived. The decoding process of the symmetric MVD mode is as follows: 1) At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0. – Otherwise, if the nearest reference picture in list -0 and the nearest reference picture in list -1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, then BiDirPredFlag is set to 1, and both the list -0 reference picture and the list -1 reference picture are short-term reference pictures. Otherwise, BiDirPred - Flag is set to 0. 2) At the CU level, if the CU is bi-directionally predicted and encoded and BiDirPredFlag is equal to 1, then the symmetric mode flag indicating whether to use the symmetric mode is signaled explicitly. When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are signaled explicitly. The reference indices for list 0 and list 1 are set to be equal to the reference picture pair respectively. MVD1 is set to be equal to (-MVD0). The final motion vector is as follows. Figure 10 The symmetric MVD mode is shown. In the encoder, the symmetric MVD motion estimation starts with an initial MV evaluation. A set of initial MV candidates includes the MVs obtained from the uni-directional prediction search, the MVs obtained from the bi-directional prediction search, and the MVs from the AMVP list. The MV with the lowest rate-distortion cost is selected as the initial MV for the symmetric MVD motion search. 2.1.3. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensated prediction (MCP). However, in the real world, there are many types of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensated prediction is applied. As Figure 11 shown, the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters). For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in a block is derived as: For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in a block is derived as: where (mv 0x , mv 0y ) is the motion vector of the upper left control point, (mv 1x , mv 1y ) is the motion vector of the upper right control point, and (mv 2x , mv 2y ) is the motion vector of the lower left control point. To simplify motion compensated prediction, block-based affine transform prediction is applied. To derive the motion vector of each 4×4 luma sub-block, the motion vector of the central sample of each sub-block is calculated according to the above equations (as Figure 12 shown), and rounded to 1 / 16 fractional precision. Then a motion compensated interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size of the chrominance component is also set to 4×4. The MV of a 4×4 chrominance sub-block is calculated as the average of the MVs of the upper left luma sub-block and the lower right luma sub-block in the co-located 8×8 luma region. Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.1.3.1. Affine Merge Prediction The AF_MERGE mode can be applied to CUs whose width and height are both greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMVP candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPVM candidates are used to form the affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMVs of neighboring CUs; – Constructed affine Merge candidate CPMVPs derived using the translational MVs of neighboring CUs; – Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as Figure 13 shown. For the prediction values on the left side, the scanning order is A0 -> A1, and for the prediction values on the upper side, the scanning order is B0 -> B1 -> B2. Only the first inherited candidate from each side is selected. No deduplication check is performed between the two inherited candidates. When the neighboring affine CU is identified, its control point motion vectors are used to derive the CPMV candidates in the affine Merge list of the current CU. As Figure 30 shown, if the neighboring lower-left block A is coded using the affine mode, the motion vectors v 2 , v 3 and v 4 of the upper-left, upper-right, and lower-left corners of the CU containing block A are obtained. When block A is coded using the 4-parameter affine model, two CPMVs of the current CU are calculated based on v 2 and v 3 . In the case where block A is coded using the 6-parameter affine model, three CPMVs of the current CU are calculated based on v 2 , v 3 and v 4 . Figure 14 The control point motion vector inheritance is shown. The constructed affine candidates refer to constructing candidates by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the specified spatial neighbors and temporal neighbors shown in Figure 15 , and CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV 1 , the B2 -> B3 -> A2 blocks are checked, and the MV of the first available block is used. For CPMV 2 , the B1 -> B0 blocks are checked, and for CPMV 3 , the A1 -> A0 blocks are checked. TMVP is used as CPMV 4 (if available). After obtaining the MVs of the four control points, affine Merge candidates are constructed based on those motion information. The following combinations of control point MVs are used to construct in sequence: {CPMV 1 , CPMV 2 , CPMV 3}, {CPMV 1 , CPMV 2 , CPMV 4}, {CPMV 1 , CPMV 3,CPMV 4 ,{CPMV 2 ,CPMV 3 ,CPMV 4 ,{CPMV 1 ,CPMV 2 ,{CPMV 1 ,CPMV 3}} Combinations of 3 CPMVs construct 6-parameter affine Merge candidates, and combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of the control point MVs are discarded. After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list. 2.1.3.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with both width and height greater than or equal to 16. In the bitstream, a CU-level affine flag is signaled to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The affine AVMP candidate list size is 2, and it is generated by sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMVs of neighboring CUs; – Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs; – Translational MVs from neighboring CUs. – Zero MVs. The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture as in the current block are considered. When inserting the inherited affine motion prediction values into the candidate list, the deduplication process is not applied. The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 15 . The same checking order as in the construction of affine Merge candidates is used. In addition, the reference picture indices of neighboring blocks are also checked. Use the first block in the checking order that is inter-coded and has the same reference picture as the current CU. There is only one. When the current CU is coded using the 4-parameter affine mode, and mv 0 and mv 1When all of them are available, they are added as a candidate in the affine AMVP list. When the current CU is coded / decoded using the 6-parameter affine mode and all three CPMVs are available, they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable. If the affine AMVP list candidates are still less than 2 after inserting valid inherited affine AMVP candidates and the constructed AMVP candidates, then mv 0 , mv 1 and mv 2 will be added in sequence as translational MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list. 2.1.3.3. Affine Motion Information Storage In VVC, the CPMVs of affine CUs are stored in a separate cache. The stored CPMVs are only used to generate the inherited CPMVs in the affine Merge mode and the inherited CPMVs in the affine AMVP mode for the most recently coded CU. The sub-block MVs derived from the CPMVs are used for motion compensation, MV derivation of the Merge / AMVP list of translational MVs, and deblocking. To avoid picture line caching for additional CPMVs, the inheritance from the affine motion data of the CU above the CTU is processed differently from the inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the row above the CTU, the left-bottom sub-block MV and the right-bottom sub-block MV in the line cache are used instead of the CPMV for affine MVP derivation. In this way, the CPMVs are only stored in the local cache. If the candidate CU is 6-parameter affine coded / decoded, the affine model is degraded to a 4-parameter model. As Figure 16 shown, along the top boundary of the CTU, the left-bottom sub-block motion vector and the right-bottom sub-block motion vector of the CU are used for affine inheritance of the CU in the bottom of the CTU. 2.1.3.4. Prediction Refinement using Optical Flow (PROF) for Affine Mode Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a more refined motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the differences derived from the optical flow equation. PROF is described as the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate sub-block prediction I(i,j). Step 2) Using a 3-tap filter [-1, 0, 1], the spatial gradient g of sub-block prediction x (i, j) and g y (i, j) are calculated at each sample position. The gradient calculation is exactly the same as that in BDOF. g x (i, j) = (I(i + 1, j) >> shift1) - (I(i - 1, j) >> shift1) (2 - 11) g y (i, j) = (I(i, j + 1) >> shift1) - (I(i, j - 1) >> shift1) (2 - 12) shift1 is used to control the accuracy of the gradient. The sub-block (i.e., 4×4) prediction extends one sample on each side of the gradient calculation. To avoid additional memory bandwidth and additional interpolation calculations, those extended samples on the extended boundaries are copied from the nearest integer pixel positions in the reference picture. Step 3) Luminance prediction refinement is calculated through the following optical flow equation. ΔI(i, j) = g x (i, j) * Δv x (i, j) + g y (i, j) * Δv y (i, j) (2 - 13) where as Figure 17 shown, Δv(i, j) is the sample MV calculated for the sample position (i, j), denoted as v(i, j), and is the difference from the sub-block MV of the sub-block to which the sample (i, j) belongs. Δv(i, j) is quantized in units of 1 / 32 luminance sample accuracy. Since the affine model parameters and the sample position relative to the sub-block center do not change from sub-block to sub-block, Δv(i, j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let dx(i, j) and dy(i, j) be the horizontal and vertical offsets of the sample position (i, j) to the sub-block center (x SB , y SB ), and Δv(x, y) can be derived through the following equation. To maintain accuracy, the input of the sub-block (x SB , y SB ) is calculated as ((W SB – 1) / 2, (H SB – 1) / 2), where W SB and H SB are the width and height of the sub-block respectively. For the 4-parameter affine model, For the 6-parameter affine model, where (v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y ) are the motion vectors of the control points in the upper-left, upper-right, and lower-left, and w and h are the width and height of the CU. Step 4) Finally, the luminance prediction refinement ΔI(i, j) is added to the sub-block prediction I(i, j). The final prediction I’ is generated by the following equation. I′(i, j) = I(i, j) + ΔI(i, j) PROF is not applicable to affine-coded CUs in two cases: 1) all control point MVs are the same, which indicates that the CU only has translational motion; 2) the affine motion parameters are greater than the specified limit, because the sub-block-based affine MC is degraded to CU-based MC to avoid large memory access bandwidth requirements. Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. In the following two cases, PROF is not applied to the affine motion estimation stage: a) if the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low-delay picture, then PROF is not applied because the improvement introduced by PROF for this case is small. In this way, the affine motion estimation using PROF can be accelerated. 2.1.4. Sub-block-based Temporal Motion Vector Prediction (SbTMVP) VVC supports the sub-block-based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and Merge mode of the CU in the current picture. The same co-located picture used by TMVP is used for SbTMVP. SbTMVP differs from TMVP in the following two main aspects: – TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; – While TMVP prefetches the temporal motion vector from the co-located block in the co-located picture (the co-located block is the bottom-right block or the central block relative to the current CU), SbTMVP applies a motion displacement before prefetching the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process is shown in FIG. 18. SbTMVP predicts the motion vector of the sub-CU within the current CU in two steps. In the first step, the spatial neighbor A1 in FIG. 18(a) is examined. If A1 has a motion vector using the co-located picture as its reference picture, then that motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the co-located picture as shown in FIG. 18(b). The example in FIG. 18(b) assumes that the motion displacement is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample) in the co-located picture is used to derive the motion information of the sub-CU. After identifying the motion information of the co-located sub-CU, it is converted to the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub-block based Merge list containing both SbTMVP candidates and affine Merge candidates is used to signal the sub-block based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry of the list of sub-block based Merge candidates, followed by the affine Merge candidates. The size of the sub-block based Merge list is signaled in the SPS, and the maximum allowed size of the sub-block based Merge list in VVC is 5. The sub-CU size used in SbTMVP is fixed at 8×8, and like the affine Merge mode, the SbTMVP mode only applies to CUs whose width and height are both greater than or equal to 8. The encoding and decoding logic for the additional SbTMVP Merge candidate is the same as that of the other Merge candidates, i.e., for each CU in the P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.1.5. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in quarter-luminance samples. In VVC, a CU-level Adaptive Motion Vector Resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be coded and decoded with different precisions. Depending on the current CU's mode (Normal AMVP mode or Affine AVMP mode), the MVD of the current CU can be adaptively selected as follows: – Normal AMVP mode: quarter-luminance samples, half-luminance samples, integer-luminance samples, or four-luminance samples. – Affine AMVP mode: quarter-luminance samples, integer-luminance samples, or 1 / 16-luminance samples. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is signaled conditionally. If all MVD components (i.e., both the horizontal and vertical MVDs of reference list L0 and reference list L1) are zero, the quarter-luminance sample MVD resolution is assumed. For a CU with at least one non-zero MVD component, the first flag is signaled to indicate whether quarter-luminance sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter-luminance sample MVD precision is used for the current CU. Otherwise, the second flag is signaled to indicate whether half-luminance sample or other MVD precision (integer or four-luminance samples) is used for Normal AMVP CUs. In the case of half-luminance samples, the half-luminance sample positions use a 6-tap interpolation filter instead of the default 8-tap interpolation filter. Otherwise, the third flag is signaled to indicate whether integer-luminance sample or four-luminance sample MVD precision is used for Normal AMVP CUs. In the case of Affine AMVP CUs, the second flag is used to indicate whether integer-luminance sample MVD precision or 1 / 16-luminance sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luminance samples, half-luminance samples, integer-luminance samples, or four-luminance samples), the motion vector prediction value of the CU will be rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded to zero (i.e., a negative motion vector prediction value is rounded to positive infinity, and a positive motion vector prediction value is rounded to negative infinity). The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM13, the RD check for MVD accuracy beyond quarter-luma samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for quarter-luma sample MVD accuracy and the RD cost for full-luma sample MV accuracy are calculated. Then, the RD cost for full-luma sample MVD accuracy is compared with the RD cost for quarter-luma sample MVD accuracy to decide whether it is necessary to further check the RD cost for four-luma sample MVD accuracy. When the RD cost for quarter-luma sample MVD accuracy is much smaller than the RD cost for full-luma sample MVD accuracy, the RD check for four-luma sample MVD accuracy is skipped. Then, if the RD cost for full-luma sample MVD accuracy is significantly greater than the best RD cost of the previously tested MVD accuracy, the check for half-luma sample MVD accuracy is skipped. For the affine AMVP mode, if the inter prediction affine mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, Merge / skip mode, quarter-luma sample MVD accuracy normal AMVP mode, and quarter-luma sample MVD accuracy affine AMVP mode, the 1 / 16-luma sample MV accuracy and 1-pixel MV accuracy affine inter prediction modes are not checked. Additionally, in the 1 / 16-luma sample and quarter-luma sample MV accuracy affine inter prediction modes, the affine parameters obtained in the quarter-luma sample MV accuracy affine inter prediction mode are used as the starting search points. 2.1.6. Bi-directional prediction with CU-level weights (BCW) In HEVC, a bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (2 - 18) Five weights are allowed in weighted-average bi-directional prediction, w ∈ {-2, 3, 4, 5, 10}. For each bi-directional prediction CU, the weight w is determined in one of two ways: 1) for non-Merge CUs, the weight index is signaled after the motion vector difference; 2) for Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5}) are used. – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-latency picture, unequal weights for 1-pixel and 4-pixel motion vector precisions are only conditionally checked. – When combined with affine, affine ME for unequal weights is performed if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bidirectional prediction are the same, unequal weights are only conditionally checked. – Unequal weights are not searched when certain conditions are met, depending on the POC distance between the current picture and its reference pictures, the coding / decoding QP, and the temporal level. The BCW weight index is decoded using a context-coded bit followed by a bypass-coded bit. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, additional bits are signaled using the bypass coding to indicate which unequal weight is used. Weighted prediction (WP) is a coding / decoding tool supported by the H.264 / AVC and HEVC standards for efficient coding / decoding of video content in fading situations. Support for WP is also added in the VVC standard. WP allows signaling of weighting parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weights and offsets of the corresponding reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For a Merge CU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded / decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.1.7. Bidirectional Optical Flow (BDOF) The Bidirectional Optical Flow (BDOF) tool is included in VVC. The BDOF, previously known as BIO, is included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bidirectional prediction signal of a CU at the 4×4 sub-block level. BDOF is applied to a CU if all of the following conditions are met for the CU: – The CU is encoded and decoded using the "true" bidirectional prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other reference picture is after the current picture in display order. – The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same. – Both reference pictures are short-term reference pictures. – The CU is not encoded and decoded using the affine mode or the SbTMVP Merge mode. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current CU. – The CIIP mode is not used for the current CU. BDOF is only applied to the luma component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4×4 sub-block, the motion refinement (v x ,v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bidirectional prediction sample values in the 4×4 sub-block. The following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals are calculated by directly computing the differences between two neighboring samples, and i.e., where I (k) (i,j) is the sample value at the coordinate (i,j) of the prediction signal in list k (k = 0,1), and shift1 is calculated as shift1 = max(6, bitDepth - 6) based on the luma bit depth (bitDepth). Then, the gradients S 1 , S 2 , S 3 , S 5 and S 6The autocorrelation and cross-correlation are calculated as follows: where where Ω is a 6×6 window around a 4×4 sub-block, and the values of n a and n b are set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8), respectively. Then, the motion refinement (v x , v y ) is derived using the cross-correlation term and the autocorrelation term with the following equation: where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on the motion refinement and the gradient, the following adjustment is calculated for each sample point in the 4×4 sub-block: Finally, the BDOF sample points of the CU are calculated by adjusting the bi-predicted sample points as follows: pred BDOF (x,y) = (I (0) (x,y) + I (1) (x,y) + b(x,y) + ο offset ) >> shift (2 - 24) These values are selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process remains within 32 bits. To derive the gradient values, some predicted sample points I (k) (i,j) in list k (k = 0, 1) outside the current CU boundary need to be generated. As Figure 9 depicted, the BDOF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating the predicted sample points outside the boundary, the predicted sample points in the extended region (white positions) are generated by directly obtaining the reference sample points at nearby integer positions (using the floor() operation on the coordinates) without interpolation, and the normal 8-tap motion compensation interpolation filter is used to generate the predicted sample points within the CU (gray positions). These extended sample point values are only used for gradient calculation. For the remaining steps in the BDOF process, if any sample points and gradient values outside the CU boundary are needed, they are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of a CU is greater than 16 luma samples, it is divided into sub-blocks with a width and / or height equal to 16 luma samples, and the sub-block boundaries are considered CU boundaries during the BDOF process. The maximum unit size for the BDOF process is limited to 16×16. For each sub-block, the BDOF process can be skipped. When the SAD between the initial L0 prediction samples and the L1 prediction samples is less than a threshold, the BDOF process is not applied to the sub-block. The threshold is set to be equal to (8*W*(H>>1), where W indicates the sub-block width and H indicates the sub-block height. To avoid additional complexity in SAD calculation, the SAD between the initial L0 prediction samples and the L1 prediction samples calculated during the DVMR process is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, the bidirectional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., for either of the two reference pictures, luma_weight_lx_flag is 1, the BDOF is also disabled. When the CU is encoded or decoded using the symmetric MVD mode or the CIIP mode, the BDOF is also disabled. 2.1.8. Decoder-Side Motion Vector Refinement (DMVR) To improve the accuracy of the Merge mode MV, decoder-side motion vector refinement based on bilateral matching is applied in VVC. During the bidirectional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and the list L1. As Figure 20 shown, the SAD between the red blocks of each MV candidate around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, DMVR can be applied to CUs encoded or decoded using the following modes and features: – CU-level Merge mode with bidirectional prediction MV. – For the current picture, one reference picture is past and the other reference picture is future. – The distance (i.e., POC difference) from the two reference pictures to the current picture is the same. – Both reference pictures are short-term reference pictures. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current block. – The CIIP mode is not used for the current block. The refined MVs derived through the DMVR process are used to generate inter-prediction samples and are also used for temporal motion vector prediction in future picture coding. The original MVs are used for the deblocking process and are also used for spatial motion vector prediction in future CU coding. The additional functions of DMVR are mentioned in the following subclauses. 2.1.8.1. Search Scheme In DVMR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point checked by DMVR represented by the candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0 + MV_offset (2-25) MV1′ = MV1 - MV_offset (2-26) where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luminance samples starting from the initial MV. The search includes an integer sample offset search stage and a fractional sample refinement stage. The integer sample offset search uses a 25-point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample stage of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV in the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates reduces the SAD value by 1 / 4. After the integer sample search is the fractional sample refinement. To save computational complexity, the fractional sample refinement is derived using the parametric error surface equation instead of using SAD comparison for additional search. The fractional sample refinement is conditionally invoked based on the output of the integer sample search stage. When the integer sample search stage ends at the center with the minimum SAD in the first iteration or the second iteration search, the fractional sample refinement is further applied. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at the four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C (2-27) where (x min , y min) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as: x min = (E(-1, 0) - E(1, 0)) / (2(E(-1, 0) + E(1, 0) - 2E(0, 0))) (2 - 28) y min = (E(0, -1) - E(0, 1)) / (2(9E(0, -1) + E(0, 1) - 2E(0, 0))) (2 - 29) x min and y min values are automatically limited between -8 and 8 because all cost values are positive and the minimum value is E(0, 0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined delta MV. 2.1.8.2. Bilinear Interpolation and Sample Padding In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate the samples at the fractional positions. In DMVR, the search points are around the initial fractional pixel MV with integer sample offsets, so the samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples during the DMVR search process. Another important effect is that by using the bilinear filter, within the 2-sample search range, compared to the normal motion compensation process, DVMR does not access more reference samples. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples of the normal MC process, samples will be padded from those available samples that are not required for the interpolation process based on the original MV but are required for the interpolation process based on the refined MV. 2.1.8.3. Maximum DMVR Processing Unit When the width and / or height of the CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16×16. 2.1.9. Combined Inter - frame and Intra - frame Prediction (CIIP) In VVC, when a CU is encoded / decoded using the Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width times the CU height is equal to or greater than 64), and if both the CU width and CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in the CIIP mode inter is derived using the same inter prediction process applied to the regular Merge mode; and the intra prediction signal P intra is derived after the regular intra prediction process with the planar mode. Then, a weighted average is used to combine the intra prediction signal and the inter prediction signal, where the weight value depends on the coding modes of the top neighboring block and the left neighboring block (depicted in Figure 21 ) and is calculated as follows: – If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) equals 2, set wt to 3; – Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2; – Otherwise, set wt to 1. The CIIP prediction is formed as follows: P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2 (2 - 30) 2.1.10. Geometric Partitioning Mode (GPM) In VVC, geometric partitioning mode is supported for inter prediction. The geometric partitioning mode is signaled using a CU-level flag as a type of Merge mode, where other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub-block Merge mode. Among a total of 64 partitions, for each possible CU size w×h = 2 m ×2 n , m,n ∈ {3…6}, the partition is supported by the geometric partitioning mode. When using this mode, the CU is divided into two parts by a geometrically positioned line ( Figure 22)。The position of the dividing line is derived mathematically from the angular parameter and offset parameter of a specific partition. Each part of the geometric partition in the CU uses its own motion for inter-frame prediction; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. The unidirectional prediction motion constraint is applied to ensure that, the same as in the case of conventional bidirectional prediction, only two motion-compensated predictions are required for each CU. If the current CU uses the geometric partition mode, the geometric partition index indicating the partition mode of the geometric partition (angle and offset) and two Merge indices (one Merge index for each partition) are further signaled. The number of the maximum GPM candidate sizes is explicitly signaled in the SPS and the syntax binarization of the GPM Merge index is specified. After predicting each part in the geometric partition, a hybrid process with adaptive weights is used to adjust the sample values along the geometric partition edge. This is the prediction signal of the entire CU, and the transform and quantization processes will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partition mode is stored. 2.1.10.1. Unidirectional Prediction Candidate List Construction The unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector of the nth extended Merge candidate (where X is equal to the parity of n) is used as the nth unidirectional prediction motion vector for the geometric partition mode. These motion vectors are marked with "x" in Figure 23 . In the case where there is no corresponding LX motion vector of the nth extended Merge candidate, the L(1-X) motion vector of the same candidate is used instead of the unidirectional prediction motion vector for the geometric partition mode. 2.1.10.2. Hybrid Along the Geometric Partition Edge After predicting each part of the geometric partition using its own motion, a hybrid is applied to the two prediction signals to derive the samples around the geometric partition edge. The hybrid weight for each position of the CU is derived based on the distance between the respective position and the partition edge. The distance from the position (x,y) to the partition edge is derived as: where i,j are the indices of the angle and offset of the geometric partition, which depend on the signaled geometric partition index. The symbols ρ x,j and ρ y,j depend on the angle index. The weight for each part of the geometric partition is derived as follows: wIdxL(x,y) = partIdx? 32 + d(x,y) : 32 - d(x,y) w 1 (x,y) = 1 - w 0 (x,y) The partIdx depends on the angle index i. The weight w 0 An example of Figure 24 is shown in 2.1.10.3. Storage of motion field for geometric partitioning pattern Mv1 from the first part of the geometric partitioning, Mv2 from the second part of the geometric partitioning, and the combination Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded and decoded in the geometric partitioning pattern. The type of motion vector stored for each individual position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx <= 0? (1 – partIdx) : partIdx) where motionIdx is equal to d(4x + 2, 4y + 2). The partIdx depends on the angle index i. If sType is equal to 0 or 1, then Mv0 or Mv1 is stored in the corresponding motion field, otherwise if sType is equal to 2, then the combination Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then simply combine Mv1 and Mv2 to form a bi - directional predicted motion vector. 2) Otherwise, if Mv1 and Mv2 are from the same list, then only store the unidirectional predicted motion Mv2. 2.1.11. Local Illumination Compensation (LIC) LIC is an inter - frame prediction technique that models the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template and the reference block template. The parameters of the function can be represented by a scale α and an offset β, which form a linear equation, i.e., α*p[x] + β to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. Since α and β can be derived based on the current block template and the reference block template, no signaling overhead is required for them, except for signaling the LIC flag for the AMVP mode to indicate the use of LIC. Local illumination compensation is used for unidirectional predicted inter - frame CUs with the following modifications. · Intra - frame neighboring samples can be used for LIC parameter derivation; · LIC is disabled for blocks with fewer than 32 luma samples; ·For both non-sub-block and affine modes, perform LIC parameter derivation based on the sample points of the modulo-block corresponding to the current CU rather than the partial modulo-block sample points corresponding to the first top-left 16×16 unit. ·Generate the sample points of the reference block template by using MC with block MV without rounding it to integer pixel precision. 2.1.12. Non-adjacent spatial candidates Insert non-adjacent spatial Merge candidates after TMVP in the regular Merge candidate list. The modes of the spatial Merge candidates are shown in Figure 25 . The distance between the non-adjacent spatial candidates and the current codec block is based on the width and height of the current codec block. The row buffer limit is not applied. 2.1.13. Template matching I Template matching I (TM) is a decoder-side MV derivation method to refine the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top neighboring block and / or the left neighboring block of the current CU) and a block in the reference picture (i.e., the same size as the template). As shown in Figure 26 , search for a better MV around the initial motion of the current CU within the [–8, +8] pixel search range. The template matching method is used for the following modifications: determine the search step size based on the AMVR mode, and TM can be cascaded with the bilateral matching process in the Merge mode. In the AMVP mode, determine the MVP candidates based on the template matching error to select the MVP candidate that achieves the minimum difference between the current block template and the reference block template, and then perform TM only for this specific MVP candidate for MV refinement. TM refines this MVP candidate by starting from the full pixel MVD precision (or 4 pixels for the 4-pixel AMVR mode) within the [–8, +8] pixel search range using iterative diamond search. The AMVP candidates can be further refined by using a cross-shaped search with full pixel MVD precision (or 4 pixels for the 4-pixel AMVR mode), and then followed by sequential searches of half pixels and quarter pixels depending on the AMVR mode specified in Table 3. This search process ensures that the MVP candidates still maintain the same MV precision as the MV precision indicated by the AMVR mode after the TM process. Table 3. Search patterns for AMVR and Merge modes with AMVR. In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 3, depending on the merged motion information and the alternative interpolation filter (the interpolation filter used when AMVR is in the half-pixel mode), the TM can be performed in all ways up to 1 / 8 pixel MVD accuracy or skipped for accuracies beyond the half-pixel MVD accuracy. In addition, when the TM mode is enabled, template matching can be used as an independent process or an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether the BM can be enabled according to its enable condition check. 2.1.14. Multiple-pass decoder-side motion vector refinement (mpDMVR) Multiple-pass decoder-side motion vector refinement is applied. In the first pass, bilateral matching (BM) is applied to the coded / decoded block. In the second pass, BM is applied to each 16×16 sub-block within the coded / decoded block. In the third pass, the MV in each 8×8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial motion vector prediction and temporal motion vector prediction. 2.1.14.1. First pass - Block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying BM to the coded / decoded block. Similar to decoder-side motion vector refinement (DMVR), in the bi-prediction operation, the refined MV is searched around the two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1. BM performs a local search to derive the integer sample accuracy intDeltaMV. The local search applies a 3×3 square search pattern to loop within the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW*cbH is greater than 64, the MRSAD cost function is applied to remove the distorted DC effect between the reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV local search terminates. Otherwise, the current minimum cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. Further apply the existing fractional sample refinement to derive the final deltaMV. Then, the refined MV after the first pass is derived as: ·MV0_pass1 = MV0 + deltaMV; ·MV1_pass1 = MV1 - deltaMV. 2.1.14.2. Second pass - sub-block based bilateral matching MV refinement In the second pass, the refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM performs a full search to derive the integer sample accuracy intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between the two reference sub-blocks, as: bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into 5 diamond search areas, as Figure 27 shown. Each search area is assigned a cost factor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order from the upper left corner of the area to the lower right corner. When the minimum bilCost within the current search area is less than a threshold equal to sbW * sbH, the full search for integer pixels is terminated; otherwise, the full search for integer pixels continues to the next search area until all search points have been checked. Further apply the existing VVC DMVR fractional sample refinement to derive the final deltaMV(sbIdx2). Then, the refined MV in the second pass is derived as: ·MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2); ·MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2). 2.1.14.3. Third pass - Sub - block - based bidirectional optical flow MV refinement In the third pass, refined MVs are derived by applying BDOF to 8×8 grid sub - blocks. For each 8×8 sub - block, starting from the refined MV of the parent - child block in the second pass, BDOF refinement is applied to derive scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 sample precision and clipped between - 32 and 32. The refined MVs in the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) are derived as follows: ·MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv; ·MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) - bioMv. 2.1.15. OBMC When applying OBMC, the motion information of neighboring blocks with weighted prediction is used to refine the top - boundary pixels and left - boundary pixels of the CU. The conditions for not applying OBMC are as follows: ·When OBMC is disabled at the SPS level. ·When the current block has an intra mode or an IBC mode. ·When LIC is applied to the current block. ·When the current luma block area is less than or equal to 32. Sub - block boundary OBMC is performed by applying the same blending to the top - sub - block boundary pixels, left - sub - block boundary pixels, bottom - sub - block boundary pixels, and right - sub - block boundary pixels using neighboring sub - blocks. It enables the following sub - block - based coding / decoding tools: ·Affine AMVP mode; ·Affine Merge mode and sub - block - based temporal motion vector prediction (SbTMVP); ·Sub - block - based bilateral matching. 2.1.16. Sample - based BDOF In sample - based BDOF, instead of deriving motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample. The coding / decoding block is divided into 8×8 sub-blocks. For each sub-block, it is determined whether to apply BDOF by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, for each sample point in the sub-block, a sliding 5×5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bidirectional prediction sample value for the central sample point of the window. 2.1.17. Interpolation The 8-tap interpolation filter used in VVC is replaced by a 12-tap filter. The interpolation filter is derived from a sine function, where the frequency response is truncated at the Nyquist frequency and clipped by a cosine window function. Table 4 gives the filter coefficients for all 16 phases. Figure 28 The frequency response of the interpolation filter is compared with the VVC interpolation filter, all at half-pixel phases. Table 4. Filter Coefficients of 12-Tap Interpolation Filter 2.1.18. Multiple Hypothesis Prediction (MHP) In the multiple hypothesis inter-frame prediction mode, in addition to the regular bi-prediction signal, one or more additional motion-compensated prediction signals are signaled. The resulting overall prediction signal is obtained by weighted sample-by-sample superposition. Using the bi-prediction signal p bi and the first additional inter-frame prediction signal / hypothesis h 3 , the resulting prediction signal p 3 is obtained as follows. p 3 =(1 - α)p bi +αh 3 According to the following mapping, the weighting factor α is specified by the new syntax element add_hyp_weight_idx. Similarly, more than one additional prediction signal can be used. The resulting overall prediction signal is cumulatively obtained iteratively using each additional prediction signal. p n+1 =(1 - α n+1 )p n +α n+1 h n+1 The resulting overall prediction signal is obtained as the last p n (i.e., with the largest index). Within this EE, up to two additional prediction signals can be used (i.e., n is limited to 2). The motion parameters of each additional prediction hypothesis can be signaled explicitly by specifying a reference index, a motion vector predictor index, and a motion vector difference or implicitly by specifying a Merge index. A separate multi-hypothesis Merge flag differentiates between these two signaling modes. For the inter-frame AMVP mode, if unequal weights in the BCW are selected in the bi-predictive mode, only the MHP is applied. A combination of MHP and BDOF is possible; however, BDOF is only applied to the bi-predictive signal part of the prediction signal (i.e., the normal first two hypotheses). 2.1.19. Adaptive Reordering of Merge Candidates Using Template Matching (ARMC-TM) Merge candidates are adaptively reordered using template matching (TM). The reordering method is applied to the regular Merge mode, the template matching (TM) Merge mode, and the affine Merge mode (excluding SbTMVP candidates). For the TM Merge mode, the Merge candidates are reordered before the refinement process. After constructing the Merge candidate list, the Merge candidates are divided into several subgroups. The subgroup size is set to 5 for the regular Merge mode and the TM Merge mode. The subgroup size is set to 3 for the affine Merge mode. The Merge candidates in each subgroup are reordered in ascending order according to the template matching-based cost value. For simplicity, the Merge candidates in the last rather than the first subgroup are not reordered. The template matching cost of a Merge candidate is measured by the sum of absolute differences (SAD) between the samples of the template of the current block and its corresponding reference samples. The template includes a set of reconstructed samples adjacent to the current block. The reference samples of the template are located by the motion information of the Merge candidate. When a Merge candidate uses bi-prediction, the reference samples of the template of the Merge candidate are also generated by bi-prediction as Figure 29 shown. For sub-block-based Merge candidates with a sub-block size equal to Wsub×Hsub, the above template includes several sub-templates of size Wsub×1, and the left template includes several sub-templates of size 1×Hsub. As Figure 30 shown, the motion information of the sub-blocks in the first row and the first column of the current block is used to derive the reference samples of each sub-template. 2.1.20. Geometric Partitioning Mode (GPM) with Merged Motion Vector Difference (MMVD) The GPM in VVC is extended by applying motion vector refinement to the top of the existing GPM unidirectional MVs. First, the flag of the GPM CU is signaled to specify whether to use this mode. If the mode is used, each geometric partition of the GPM CU can further decide whether to signal the MVD. If the MVD is signaled for a geometric partition, after selecting the GPMMerge candidate, the motion of the partition is further refined by the signaled MVD information. All other processes remain the same as GPM. The MVD is signaled as a pair of distance and direction, similar to in MMVD. There are nine candidate distances (1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 3 pixels, 4 pixels, 6 pixels, 8 pixels, 16 pixels) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) involved in GPM-MMVD in GPM. Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD is left-shifted by 2 in MMVD. 2.1.21. Geometric Partitioning Mode (GPM) Using Template Matching (TM) Template matching is applied to GPM. When the GPM mode is enabled for a CU, a CU-level flag is signaled to indicate whether TM is applied to the two geometric partitions. TM is used to refine the motion information of each geometric partition. When TM is selected, depending on the partition angle, left, above, or both left and above neighboring samples are used to construct the template, as shown in Table 5. Then the motion is refined by minimizing the difference between the current template and the template in the reference picture using the same search pattern as the Merge mode with the half-pixel interpolation filter disabled. Table 5 Templates for the first geometric partition and the second geometric partition, where A indicates using above samples, L indicates using left samples, and L+A indicates using both left and above samples. Partition Angle 0 2 3 4 5 8 11 12 13 14 First Partition A A A A L + A L + A L + A L + A A A Second Partition L + A L + A L + A L L L L L + A L + A L + A Partition Angle 16 18 19 20 21 24 27 28 29 30 First Partition A A A A L + A L + A L + A L + A A A Second Partition L + A L + A L + A L L L L L + A L + A L + A The GPM candidate list is constructed as follows: 1. The interleaved list 0 MV candidates and list 1 MV candidates are directly derived from the regular Merge candidate list, where the list 0 MV candidates have a higher priority than the list 1 MV candidates. A deduplication method with an adaptive threshold based on the current CU size is applied to remove redundant MV candidates. 2. The interleaved list 1 MV candidates and list 0 MV candidates are further directly derived from the regular Merge candidate list, where the list 1 MV candidates have a higher priority than the list 0 MV candidates. The same deduplication method with an adaptive threshold is also applied to remove redundant MV candidates. 3. Zero MV candidates are filled until the GPM candidate list is full. GPM-MMVD and GPM-TM are specifically enabled for one GPM CU. This is done by first signaling the GPM-MMVD syntax. When both GPM-MMVD control flags are equal to false (i.e., GPM-MMVD is disabled for both GPM partitions), the GPM-TM flag is signaled to indicate whether template matching is applied to both GPM partitions. Otherwise (at least one GPM-MMVD flag is equal to true), the value of the GPM-TM flag is presumed to be false. 2.1.22. GPM with Inter-Frame and Intra-Frame Prediction (GPM inter-intra) With GPM inter-intra, in addition to the Merge candidates for each non-rectangular partition region in the CU to which GPM is applied, a predetermined intra-frame prediction mode for the geometric segmentation line can also be selected. In the proposed method, the intra-frame prediction mode or the inter-frame prediction mode is determined for each GPM separable region with a flag from the encoder. When in the inter-frame prediction mode, a unidirectional prediction signal is generated from the MV in the Merge candidate list. On the other hand, when in the intra-frame prediction mode, a unidirectional prediction signal is generated from neighboring pixels for the intra-frame prediction mode specified by the index from the encoder. The variation of possible intra-frame prediction modes is restricted by the geometry. Finally, the two unidirectional prediction signals are blended in the same way as in ordinary GPM. 2.1.23. Adaptive Decoder-Side Motion Vector Refinement (Adaptive DMVR) The adaptive decoder-side motion vector refinement method consists of two new Merge modes, which are introduced to refine the MV only in one direction (L0 or L1) of the bi-directional prediction of the Merge candidates that meet the DMVR conditions. The selected Merge candidates are applied with a multi-pass DMVR process to refine the motion vector. However, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is zero. Similar to the conventional Merge mode, the Merge candidates for the proposed Merge mode are derived from spatially neighboring coded / decoded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that meet the DMVR conditions are added to the candidate list. The same Merge candidate list is used by the two proposed Merge modes, and the Merge index is coded / decoded in the conventional Merge mode. 2.1.24. Bilateral Matching AMVP-MERGE Mode (AMVP-MERGE) In the AMVP Merge mode, the bi-directional prediction value consists of the AMVP prediction value in one direction and the Merge prediction value in the other direction. The AMVP part of the proposed mode is signaled as regular unidirectional AMVP, i.e., the reference index and MVD are signaled, and it has a derived MVP index (TM_AMVP) if template matching is used, or the MVP index is signaled when template matching is disabled. The Merge index is not signaled, and the Merge prediction value is selected from the candidate list with the minimum template or bilateral matching cost. When the selected Merge prediction value and AMVP prediction value meet the DMVR condition (which is at least one reference picture from the past and one reference picture from the future relative to the current picture) and the distance from the two reference pictures to the current picture is the same, bilateral matching MV refinement is applied to the Merge MV candidate and AMVP MVP as a starting point. Otherwise, if the template matching function is enabled, template matching MV refinement is applied to the Merge prediction value or AMVP prediction value with a higher template matching cost. The third pass of 8×8 sub-PU BDOF refinement as multi-pass DMVR is enabled as AMVP Merge mode codec block. 2.1.25.IBC Merge / AMVP List Construction IBC Merge / AMVP list construction is modified as follows: An IBC Merge / AMVP candidate can be inserted into the IBC Merge / AMVP candidate list only if it is valid. The upper right spatial domain candidate, the lower left spatial domain candidate, the upper left spatial domain candidate and a pairwise average candidate may be added to the IBC Merge / AMVP candidate list. Adaptive Reordering Based on Template (ARMC-TM) is applied to the IBC Merge list. The HMVP table size for IBC is increased to 25. After deriving up to 20 IBC Merge candidates with full deduplication, they are re-ranked together. After re-ranking, the top 6 candidates with the lowest template matching cost are selected as the final candidates in the IBC Merge list. Candidates for zero vectors used to populate the IBC Merge / AMVP list are replaced with a set of BVP candidates located in the IBC reference region. Zero vectors are invalid as block vectors in IBC Merge mode, and therefore, are discarded as BVPs in the IBC candidate list. Three candidates are located at the nearest corners of the reference region, and three additional candidates are determined in the middle of the three sub-regions (A, B, and C), whose coordinates are determined by the width and height of the current block and the ΔX and ΔY parameters, as Figure 31 Depicted in. 2.1.26. IBC Using Template Matching Template matching is used in both the IBC Merge mode and the IBC AMVP mode in IBC. The IBC-TM Merge list is modified compared to the list used by the regular IBC Merge mode such that candidates are selected according to a deduplication method using the motion distance between candidates as in the regular TM Merge mode. End-zero motion is fulfilled by motion vectors of the left (-W, 0), top (0, -H), and top-left (-W, -H), where W is the width of the current CU and H is the height of the current CU. In the IBC-TM Merge mode, the selected candidates are refined using a template matching method before the RDO or decoding process. The IBC-TM Merge mode has been competing with the regular IBC Merge mode and is signaled by the TM-Merge flag. In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC-TM Merge list. A template matching method is used to refine each of these 3 selected candidates and they are sorted according to their resulting template matching cost. Then usually only the first two candidates are considered during the motion estimation process. The template matching refinement for both the IBC-TM Merge mode and the AMVP mode is very simple because the IBC motion vectors are constrained (i) to be integers and (ii) within the reference region as shown in Figure 32 Therefore, in the IBC-TM Merge mode, all refinements are performed with integer precision, and in the IBC-TM AMVP mode, it is performed with integer or 4-pixel precision depending on the AMVR value. Such refinements only access samples without interpolation. In both cases, the refined motion vectors and the templates used in each refinement step must comply with the constraints of the reference region. 2.1.27. IBC Reference Region The reference region of IBC extends to two CTU rows above. Figure 33Shows the reference region for encoding and decoding CTU(m, n). Specifically, for the CTU(m, n) to be encoded and decoded, the reference region includes CTUs with indices (m-2, n-2)…(W, n-2), (0, n-1), (W, n-1), (0, n)…(m, n), where W represents the maximum horizontal index within the current slice, strip, or picture. This setting ensures that for a CTU size of 128, IBC does not require additional memory in the current ETM platform. The per-sample block vector search (or local search) range is horizontally limited to [-(C<<1), C>>2] and vertically limited to [-C, C>>2] to accommodate reference region expansion, where C represents the CTU size. 2.1.28. MVD Symbol Prediction In this method, possible MVD symbol combinations are sorted according to the template matching cost, and the index corresponding to the true MVD symbol is derived and context decoded. On the decoder side, the MVD symbol is derived as follows: 1. Parse the size of the MVD component; 2. Parse the context decoded MVD symbol prediction index; 3. Construct MV candidates by creating combinations between possible symbols and absolute MVD values and adding them to the MV prediction value; 4. Derive the MVD symbol prediction cost for each derived MV based on the template matching cost and sorting; 5. Use the MVD symbol prediction index to pick the true MVD symbol. MVD symbol prediction is applied to the inter-frame AMVP mode, affine AMVP mode, MMVD mode, and affine MMVD mode. 2.1.29. Enhanced Bidirectional Motion Compensation In bidirectional motion compensation, out-of-bounds (OOB) predicted samples are discarded, and only non-OOB predicted values are used to generate the final predicted value. Specifically, assume Pos_x i,j and Pos_y i,j represent the position of a predicted sample in a current block, and represent the MV of the current block; Pos LeftBdry 、Pos RightBdry 、Pos TopBdry and Pos BottomBdry are the positions of the four boundaries of the picture. A predicted sample is considered OOB when at least one of the following conditions is met: where half_pixel is equal to 8, which represents the half-pixel sample distance in 1 / 16 pixel sample precision. After checking the OOB condition for each sample point, a final predicted sample point of a bi - directional block is generated as follows: If is OOB and is non - OOB Otherwise, if is non - OOB and is OOB. Otherwise When BCW is enabled, the OOB checking process also applies. 2.1.30. Block - level reference picture list re - ordering Use a block - level reference picture re - ordering method based on template matching. For the unidirectional prediction AMVP mode, the reference pictures in list 0 and list 1 are interleaved to generate a combined list. For each hypothesis of the reference pictures in the combined list, a match is performed to calculate the cost. The combined list is re - ordered in ascending order of the template - matching cost. The index of the selected reference picture in the re - ordered combined list is signaled in the bitstream. For the bi - directional prediction AMVP mode, a list of reference picture pairs from list 0 and list 1 is generated and similarly re - ordered based on the template - matching cost. The index of the selected pair is signaled. 2.1.31. Affine model inheritance based on historical parameters and non - adjacent affine mode Affine model inheritance based on historical parameters (HAMI) allows an affine model to inherit from a previously affine - encoded block that may not be adjacent to the current block. Similar to the enhanced regular Merge mode, a non - adjacent affine mode (NA - AFF) is introduced. A first historical parameter table (HPT) is established. The entries of the first HPT store sets of affine parameters: a, b, c, and d, each of which is represented by a 16 - bit signed integer. The entries in the HPT are classified by reference list and reference index. Five reference indices are supported for each reference list in the HPT. In a formulaic way, the category of the HPT (denoted as HPTCat) is calculated as HPTCat(RefList, RefIdx) = 5×RefList+min(RefIdx, 4), where RefList and RefIdx represent the reference picture list (0 or 1) and the reference index respectively. For each category, up to seven entries can be stored, resulting in a total of 70 entries in the HPT. At the start of each CTU row, the number of entries for each category is initialized to zero. When decoding a block with reference list RefList curand RefIdx cur After the CU with affine encoding and decoding, the affine parameters are utilized to update the entries in the category HPTCat(RefList cur , RefIdx cur ) in a manner similar to the update of the HMVP table. Candidate based on historical affine parameters (HAPC) is derived from one of the seven neighboring 4×4 blocks represented as A0, A1, A2, B0, B1, B2, or B3 in Figure 4 and the set of affine parameters in the corresponding entry stored in the first HPT. The MV of the neighboring 4×4 block serves as the base MV. In a formulated way, the MV of the current block at position (x, y) is calculated as: where (mv h base , mv v base ) represents the MV of the neighboring 4×4 block, and (x base , y base ) represents the center position of the neighboring 4×4 block. (x, y) can be the upper-left, upper-right, and lower-left corners of the current block to obtain the corner position MV (CPMV) of the current block, or it can be the center of the current block to obtain the regular MV of the current block. A second historical parameter table (HPT) with base MV information is also appended. There are nine entries in the second HPT, where the entries include the base MV, reference indices for each reference list, four affine parameters, and the base position. Additional Merge HAPC can be generated from the second HPT with base MV information, and the corresponding affine models are stored in the entries. The difference between the first HPT and the second HPT is shown in Figure 34 . In addition, paired affine Merge candidates are generated from two affine Merge candidates, which can be historically derived or non-historically derived. The paired affine Merge candidates are generated by averaging the CPMVs of the existing affine Merge candidates in the list. In response to the introduction of new HAPC, the size of the Merge candidate list based on sub-blocks is increased from 5 to 15, all of which are involved in the ARMC
[10] process. In NA-AFF, the mode of obtaining non-adjacent spatial neighbors is shown in Figure 6 . Similar to the existing non-adjacent regular Merge candidates [8], the distance between the non-adjacent spatial neighbors and the current coding block in NA-AFF is also defined based on the width and height of the current CU. Utilize Figure 6Motion information of non - adjacent spatial neighbors in [ [ ] ] is used to generate additional inherited and constructed affine Merge / AMVP candidates. Specifically, for inherited candidates, except that CPMV is inherited from non - adjacent spatial neighbors, the same derivation process of inherited affine Merge / AMVP candidates in VVC remains unchanged. Non - adjacent spatial neighbors are checked based on their distance from the current block (i.e., from near to far). At a specific distance, only the first available neighbors (coded / decoded using affine mode) from each side of the current block (e.g., left and above) are included for inherited candidate derivation. Figure 35 shows the spatial neighbors for deriving affine Merge / AMVP candidates. Additionally, Figure 35 sub - picture (a) of [ [ ] ] shows the spatial neighbors for deriving inherited candidates, and Figure 35 sub - picture (b) of [ [ ] ] shows the spatial neighbors for deriving the first type of constructed candidates. As Figure 35 indicated by the dashed arrows in sub - picture (a) of [ [ ] ], the checking order of the left and above neighbors is from bottom to top and from right to left, respectively. For the first type of constructed candidates, as Figure 35 shown in sub - picture (b) of [ [ ] ], first, the positions of a left and an above non - adjacent spatial neighbor are independently determined; then, the position of the upper - left neighbor can be determined accordingly, which can enclose a rectangular virtual block with the left and above non - adjacent neighbors. Then, as Figure 36 shown, the motion information of the three non - adjacent neighbors is used to form CPMVs at the upper - left (A), upper - right (B), and lower - left (C) of the virtual block, and finally projected onto the current CU to generate the corresponding constructed candidates. NA - AFF candidates are inserted into the existing affine Merge candidate list and affine AMVP candidate list according to the following order: Affine Merge Mode: 1. SbTMVP candidates, if available. 2. Inheritance from adjacent neighbors. 3. Inheritance from non - adjacent neighbors. 4. Construction from adjacent neighbors. 5. The first type of constructed affine candidates from non - adjacent neighbors. 6. Zero MV. Affine AMVP Mode: 1. Inheritance from adjacent neighbors. 2. Construction from adjacent neighbors. 3. Translational MV from adjacent neighbors. 4. Translational MV from temporal neighbors. 5. Inheritance from non - adjacent neighbors. 6. First type of constructing affine candidates from non - adjacent neighbors. 7. Zero MV. Due to including additional candidates generated by NA - AFF, the size of the affine Merge candidate list increases from 5 to 15. The subgroup size of ARMC for the affine Merge mode increases from 3 to 15. In NA - AFF: 1. Regions from non - adjacent neighbors are restricted within the current CTU (i.e., there is no additional storage requirement for line buffering). 2. The storage granularity of affine motion information including CPMV and reference index is reduced from 8×8 to 16×16 (i.e., only the affine motion from the top - left 8×8 block is saved). Additionally, the saved CPMV is projected onto each 16×16 block before storage, such that no position and size information is required. 3. Only the top - left CPMV and the top - right CPMV are stored (i.e., always use the 4 - parameter affine model for NA - AFF). 2.1.32. Regression - based affine candidate derivation method A regression - based affine candidate derivation method is proposed. The sub - block motion fields from previously decoded affine CUs and the motion vectors from adjacent sub - blocks of the current CU are used as inputs to the regression process. The predicted CPMV instead of the sub - block motion field of the current block is derived as the output. The derived CPMV can be added to the sub - block Merge candidate list or the affine AMVP list. The scan pattern of previously decoded affine CUs is the same as the non - adjacent scan pattern used in the construction of the regular Merge candidate list. 2.2. Transform and coefficient coding 2.2.1. Large - block - size transform with high - frequency zeroing In VVC, large - block - size transforms with sizes up to 64×64 are enabled, which are mainly used for higher - resolution videos such as 1080p and 4K sequences. For transform blocks with a size (width or height, or both width and height) equal to 64, the high - frequency transform coefficients are zeroed, such that only the lower - frequency coefficients are retained. For example, for an M×N transform block, where M is the block width and N is the block height, when M equals 64, only the left 32 - column transform coefficients are retained. Similarly, when N equals 64, only the first 32 - row transform coefficients are retained. When the transform skip mode is used for large blocks, the whole block is used without zeroing any values. Additionally, the transform displacement is removed in the transform skip mode. VTM also supports a configurable maximum transform size in the SPS, such that the encoder has the flexibility to select a transform size of up to 32 - length or 64 - length according to the needs of a specific implementation. 2.2.2. Multiple transform selection (MTS) for kernel transform In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of inter- and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 6 shows the basis functions of the selected DST / DCT. Table 6 - Transform basis functions of DCT-II / VIII and DSTVII for N-point input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10-bit after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter frames respectively. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applicable to luminance. The MTS signaling is skipped when one of the following conditions is met: – The position of the last significant coefficient of the luminance TB is less than 1 (i.e., only DC). – The last significant coefficient of the luminance TB is within the MTS zeroing region. If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 7. A unified transform selection for ISP and implicit MTS is used by eliminating the intra mode and block shape dependencies. If the current block is in the ISP mode or if the current block is an intra block and both intra and inter explicit MTS are on, only DST7 is used for horizontal and vertical transform kernels. In terms of transform matrix precision, an 8-bit primary transform kernel is used. Thus, all transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8) all use an 8-bit primary transform kernel. Table 7 - Transform and signaling mapping table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are zeroed. Only the coefficients within the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded or decoded using the transform skip mode. To avoid redundancy in syntax coding, the transform skip flag is not signaled when the MTS_CU_flag at the CU level is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, the implicit MTS can still be enabled. 2.2.3. Low-Frequency Non-Separable Transform (LFNST) In VVC, as Figure 37 shown, LFNST is applied between the forward main transform and quantization (at the encoder) and between the de-quantization and the inverse main transform (at the decoder side). In LFNST, a 4×4 non-separable transform or an 8×8 non-separable transform is applied according to the block size. For example, the 4×4 LFNST is applied to small blocks (i.e., min(width, height) < 8), and the 8×8 LFNST is applied to larger blocks (i.e., min(width, height) > 4). The following uses the input as an example to describe the application of the non-separable transform used in LFNST. To apply the 4×4 LFNST, the 4×4 input block X is first represented as a vector The non-separable transform is calculated as where indicates the transform coefficient vector, and T is a 16×16 transform matrix. Subsequently, the 16×1 coefficient vector is reorganized into a 4×4 block using the scan order (horizontal, vertical, or diagonal) for the block. The coefficients with smaller indices will be placed in the 4×4 coefficient block together with the smaller scan indices. 2.2.3.1. Reduced Non-Separable Transform The LFNST (Low-Frequency Non-Separable Transform) applies the non-separable transform based on the direct matrix multiplication method, such that it is implemented in a single pass without multiple iterations. However, it is necessary to reduce the non-separable transform matrix size to minimize the computational complexity and the memory spatial domain to store the transform coefficients. Therefore, the reduced non-separable transform (or RST) method is used in LFNST. The main idea of the reduced non-separable transform is to map an N-dimensional vector (where N is usually equal to 64 for 8×8 NSST) to an R-dimensional vector in a different space, where N / R (R < N) is the reduction factor. Thus, instead of an N×N matrix, the RST matrix becomes an R×N matrix as follows: The transformed R rows are the R bases of the N-dimensional space. The inverse transform matrix for RT is the transpose of its forward transform. For the 8×8 LFNST, a reduction factor of 4 is applied, and the 64×64 direct matrix (which is the conventional 8×8 non-separable transform matrix size) is reduced to a 16×48 direct matrix. Thus, a 48×16 inverse RST matrix is used on the decoder side to generate the kernel (primary) transform coefficients in the top-left 8×8 region. When applying the 16×48 matrix instead of 16×64 with the same transform set configuration, each of them takes 48 input data from three 4×4 blocks in the top-left 8×8 block except the bottom-right 4×4 block. With the reduced size, the memory usage for storing all LFNST matrices is reduced from 10 KB to 8 KB with a reasonable performance degradation. To reduce complexity, it is applicable to limit the LFNST only when all coefficients outside the first coefficient subgroup are not significant. Thus, when applying the LFNST, all only the primary transform coefficients must be zero. This allows adjusting the LFNST index signaling at the last valid position and thus avoids the additional coefficient scanning in the current LFNST design, which requires checking the valid coefficients only at specific positions. The worst-case processing of the LFNST (in terms of multiplications per pixel) limits the non-separable transforms for 4×4 blocks and 8×8 blocks to 8×16 transforms and 8×48 transforms respectively. In these cases, when applying the LFNST, the last valid scan position must be less than 8 for other sizes less than 16. For blocks with shapes of 4×N and N×4 and N>8, the proposed limitation means that the LFNST is now applied only once and only to the top-left 4×4 region. Since all only the primary coefficients are zero when applying the LFNST, the number of operations required for the primary transform is reduced in this case. From the encoder's perspective, when testing the LFNST transform, the quantization of the coefficients is significantly simplified. Rate-distortion optimized quantization must be maximally done for the first 16 coefficients (in scan order), and the remaining coefficients are forced to zero. 2.2.3.2. LFNST Transform Selection There are a total of 4 transform sets and 2 non-separable transform matrices (kernels) used in the LFNST. As shown in Table 8, the mapping from the intra prediction mode to the transform set is predefined. If one of the three CCLM modes (INTRA_LT_CCLM, INTRA_T_CCLM, or INTRA_L_CCLM) is used for the current block (81 <= predModeIntra <= 83), then transform set 0 is selected for the current chrominance block. For each transform set, the selected non-separable quadratic transform candidate is further specified by the explicitly signaled LFNST index. The index is signaled once per intra CU in the bitstream after the transform coefficients. Table 8 - Transform Selection Table IntraPredMode Transform Set Index IntraPredMode < 0 1 0 <= IntraPredMode <= 1 0 2 <= IntraPredMode <= 12 1 13 <= IntraPredMode <= 23 2 24 <= IntraPredMode <= 44 3 45 <= IntraPredMode <= 55 2 56 <= IntraPredMode <= 80 1 81 <= IntraPredMode <= 83 0 2.2.3.3. LFNST Index Signaling and Interaction with Other Tools Since LFNST is restricted to apply only when all coefficients outside the first coefficient subgroup are not significant, the LFNST index encoding and decoding depends on the position of the last significant coefficient. Additionally, the LFNST index is context - decoded, but does not depend on the intra - prediction mode, and only the first binary bit is context - decoded. Moreover, LFNST is applied to intra CUs in both intra - slices and inter - slices as well as for both luminance and chrominance. If dual - tree is enabled, the LFNST indices for luminance and chrominance are signaled separately. For inter - slices (dual - tree disabled), a single LFNST index is signaled and used for both luminance and chrominance. Considering that due to the existing maximum transform size limit (64×64), large CUs larger than 64×64 are implicitly partitioned (TU slicing), the LFNST index search can increase the data buffer up to four times the number of specific decoding pipeline stages. Therefore, the maximum size allowing LFNST is restricted to 64×64. Note that LFNST only enables DCT2. The LFNST index signaling is placed before the MTS index signaling. The use of the scaling matrix for perceptual quantization is not obvious, and the scaling matrix specified for the main matrix can be used for LFNST coefficients. Therefore, the use of the scaling matrix for LFNST coefficients is not allowed. For the single - tree partition mode, chrominance LFNST is not applied. 2.2.4. Sub - block Transform (SBT) In VTM, a sub - block transform is introduced for CUs in inter - prediction. In this transform mode, for a CU, only a sub - part of the residual block is encoded and decoded. When the cu_cbf of an inter - predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub - part of the residual block is encoded and decoded. For the former case, the inter - frame MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is encoded and decoded by the presumptive adaptive transform while the other part of the residual block is zeroed. When SBT is used for an inter - decoded CU, the SBT type and SBT position information are signaled in the bit - stream. There are two SBT types and two SBT positions, as Figure 38As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to the binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to the asymmetric binary tree (ABT) partition. In the ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 (in terms of luma samples), a 1:3 / 3:1 partition along that dimension is not allowed. A CU can have at most 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chrominance TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal transform and the vertical transform for each SBT position are specified in Figure 38 . For example, the horizontal transform and the vertical transform for SBT-V position 0 are DCT-8 and DST-7 respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Therefore, the sub-block transform jointly specifies the TU slicing, cbf, and the horizontal and vertical kernel transform types of the residual block. SBT is not applied to CUs encoded and decoded using a combined inter-intra mode. 2.2.5. Maximum Transform Size and Zeroing of Transform Coefficients Both the CTU size and the maximum transform size (i.e., all MTS transform kernels) are extended to 256, where the maximum intra-coded block can have a size of 128×128. For UHD sequences, the maximum CTU size is set to 256, otherwise it is set to 128. During the main transform process, there is no standardized zeroing operation applied to the transform coefficients. However, if LFNST is applied, the main transform coefficients outside the LFNST region are standardized to zero. 2.2.6. Enhanced MTS for Intra Coding In the current VVC design, for MTS, only the DST7 transform kernel and the DCT8 transform kernel are utilized, which are used for both intra and inter coding. Additional main transforms including DCT5, DST4, DST1, and the identity transform (IDT) are adopted. The MTS set also depends on the TU size and the intra mode information. 16 different TU sizes are considered, and for each TU size 5, different categories are considered according to the intra mode information. For each category, 1, 4, or 6 different transform pairs are considered. Multiple intra MTS candidates (among 1, 4, and 6 MTS candidates) are adaptively selected according to the sum of the absolute values of the transform coefficients. The sum is compared with two fixed thresholds to determine the total number of allowed MTS candidates: 1 candidate: sum <= th0. 4 candidates: th0 < sum <= th1. 6 candidates: sum > th1. Note that although 80 different categories are considered in total, some of these different categories usually share exactly the same set of transforms. Thus, there are 58 (less than 80) unique entries in the resulting LUT. For the angular mode, joint symmetry on the TU shape and intra prediction is considered. Thus, a mode i (i > 34) with a TU shape A×B will be mapped to the same category corresponding to a mode j = (68 - i) with a TU shape B×A. However, for each transform pair, the order of the horizontal transform kernel and the vertical transform kernel is swapped. For example, a 16×4 block with mode 18 (horizontal prediction) and a 4×16 block with mode 50 (vertical prediction) are mapped to the same category. However, the vertical transform kernel and the horizontal transform kernel are swapped. For the wide-angle mode, the closest regular angular mode is used for transform set determination. For example, mode 2 is used for all modes between -2 and -14. Similarly, mode 66 is used for modes 67 to 80. 2.2.7. Quadratic Transform: LFNST Extension with Large Kernels The LFNST in VVC is extended as follows: · The number of the LFNST set (S) and candidates (C) is extended to S = 35 and C = 3, and the LFNST set (lfnstTrSetIdx) for a given intra mode (predModeIntra) is derived according to the following formula: o For predModeIntra < 2, lfnstTrSetIdx is equal to 2; o lfnstTrSetIdx = predModeIntra, for predModeIntra in [0, 34]; o lfnstTrSetIdx = 68 - predModeIntra, for predModeIntra in [35, 66]. · Three different kernels LFNST 4, LFNST 8, and LFNST 16 are defined to indicate the sets of LFNST kernels applied to 4×N / N×4 (N≥4), 8×N / N×8 (N≥8), and M×N (M, N≥16), respectively. The kernel size is specified as follows: (LFSNT4, LLFNST8*, LFNST16*) = (16×16, 32×64, 32×96) The forward LFNST is applied to the upper-left low-frequency region called the region of interest (ROI). When applying the LFNST, the main transform coefficients existing in the region other than the ROI are cleared and are not changed from the VVC standard. The ROI of LFNST16 is in Figure 39 shown. It consists of six 4×4 sub-blocks which are consecutive in scan order. Since the number of input samples is 96, the transform matrix for the forward LFNST16 can be R×96. In this contribution, R is chosen as 32, and accordingly 32 coefficients (two 4×4 sub-blocks) are generated from the forward LFNST16, which are placed after the coefficient scan order. The ROI of LFNST8 is in Figure 40 shown. The forward LFNST8 matrix can be R×64, and R is chosen as 32. The generated coefficients are positioned in the same way as LFNST 16. The mapping from the intra prediction mode to these sets is shown in Table 9, Table 9. Mapping of Intra Prediction Mode to LFNST Set Index Intra Prediction Mode -14 -13 -12 -11 -10 -9 -8 -7 -6 -5 -4 -3 -2 -1 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 LFNsT Set Index 2 2 2 2 2 2 2 2 2 2 2 2 2 2 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 Intra Prediction Mode 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 LFNsT Set Index 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 33 32 31 30 29 28 27 26 25 24 23 22 21 20 19 Intra Prediction Mode 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 LFNsT Set Index 18 17 16 15 14 13 12 11 10 9 8 7 6 5 4 3 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2 2.2.8. Sign Prediction The basic idea of the coefficient sign prediction method is to calculate the reconstruction residuals for the negative and positive sign combinations for applicable transform coefficients and select the hypothesis that minimizes the cost function. To derive the optimal sign, the cost function is defined as the measurement of the discontinuity across Figure 41 the block boundaries shown above. All hypotheses are measured, and the hypothesis with the minimum cost is selected as the predicted value of the coefficient sign. The cost function is defined as the sum of the absolute second derivatives in the residual domain of the upper row and the left column as follows: where R is the reconstructed neighbor, P is the prediction of the current block, and r is the residual hypothesis. The term (-R - 1 +2R 0 -P1) can be calculated only once per block and only the residual hypotheses are subtracted. The transform coefficients with the maximum KqIdx value in the upper-left 4×4 region are selected. After compensating for the effects of multiple quantizers in DQ, the qIdx value is at the transform system level. A larger qIdx value will produce a larger dequantized transform coefficient level. The qIdx is derived as follows: qIdx = (abs(level) << 1) - (state & 1); where level is the transform coefficient level parsed from the bitstream, and state is a variable maintained by the encoder and decoder in DQ. The symbol prediction region is extended to a maximum of 32×32. Predict the symbols of the top-left M×N block. The values of M and N are calculated as follows: οM = min(w, maxW) οN = min(h, maxH) where w and h are the width and height of the transformed block. The maximum region for symbol prediction is not always set to 32×32. The encoder sets the maximum region (maxW, maxH) based on the configuration, sequence category, and QP, and signals the region in the SPS. The maximum number of predicted symbols remains unchanged. Symbol prediction is also applied to the LFNST block. And for the LFNST block, symbol prediction is allowed for the 4 largest coefficients in the top-left 4×4 region. 3. Problems / Issues There are several problems with existing video coding and decoding technologies, which can be further improved for higher coding and decoding gains. 1. In ECM-6.0, affine candidates can be derived from adjacent affine-based candidates, history-based affine candidates, non-adjacent affine candidates, and regression-based affine candidates. Similarity checks are performed for affine candidate derivation. However, differential similarity check rules are used for affine candidate derivation. This may not be optimal. 2. In ECM-6.0, hybrid modes such as CIIP and OBMC are applied to both videos captured by cameras and screen content videos, which may not be efficient. 3. In ECM-6.0, KLT is allowed for the explicit inter-frame MTS mode. Specifically, if the TU for inter-frame coding is less than or equal to 16×16, two KLT options (i.e., KLT0 and KLT1) of the inter-frame MTS kernel are used to replace DST7 and DCT8. This design can be changed for higher coding and decoding efficiency. 4. In ECM-6.0, KLT is allowed for the inter-frame MTS mode, but the use of KLT does not depend on which inter-frame prediction technique is used for the video unit, which can be further improved. 5. In ECM-6.0, the following intra-frame mode derivation / mapping for intra-frame MTS and LFNST indices can be improved. 1) In the dual-tree case, LFNST is applied to the chrominance component. For the CCLM mode, the co-located luma mode is used for the LFNST transform set and transpose flag index. 2) For the MIP mode, the planar mode is used for the LFNST transform set and transpose flag index. 3) MIP is regarded as a special mode for the intra-frame MTS transform class and intra-frame MTS transform pair index. 4) Allow IntraTMP to use implicit MTS (e.g., DST7) and LFNST (considered as planar mode). 5) In hybrid modes such as TIMD hybrid mode and DIMD hybrid mode, only consider the first intra mode for MTS / LFNST indexing. 6. For codec tools such as IntraTMP and IBC, there is no sample refinement process for motion-compensated prediction, which can be designed in the future. 7. For sbTMVP codec, DMVR is not allowed in the current codec, which can be changed for higher codec gain. 8. For codec tools such as LIC, OBMC, LM, derive uniform parameters for the samples to be processed in the block. However, non-uniform mixing can be performed using a multi-model-based method. 4. Embodiments of the present disclosure The following detailed embodiments should be considered as examples for explaining general concepts. These embodiments should not be interpreted in a narrow way. In addition, these embodiments can be combined in any way. The term "video unit" or "codec unit" or "block" may represent a coding tree block (CTB), a coding tree unit (CTU), a coding block (CB), a CU, a PU, a TU, a PB, a TB. The term "KLT" may refer to a type of transform. For example, it may refer to the Karhunen–Loève transform. For example, it may refer to any transform type that is not DCT or DST or Hadmard. The coefficient matrix associated with a specific KLT can be trained online or predefined (e.g., offline training) according to some prior knowledge (e.g., residuals / coefficients from already decoded neighboring blocks). In the present disclosure, regarding "a block coded using mode N", here "mode N" can be a specific prediction mode (e.g., MODE_INTRA, MODE_INTER, MODE_PLT, MODE_IBC, etc.), or a specific prediction technique (e.g., AMVP, Merge, SMVD, BDOF, PROF, DMVR, AMVR, TM, affine, CIIP, GPM, GPM intra, MHP, OBMC, LIC, GEO, TPM, MMVD, BCW, HMVP, SbTMVP, sub-block coding, hypothesis coding, etc.), or a specific transform process (IDTX, DCT-X, DST-Y, KLT-Z, where X / Y / Z are constants) or a specific filter process (deblocking, SAO, bilateral filter, adaptive loop filter, CCSAO, CC-ALF, etc.). Note that the terms mentioned below are not limited to the specific terms defined in existing standards. Any changes to the codec tools are also applicable. 4.1. Regarding the first problem of the derivation of affine candidates and other motion candidate deduplication processes, the following method is proposed: a. The same logic / rules / process for similarity / consistency / deduplication checking can be used for the derivation of all affine candidates. a. For example, it can refer to the derivation of affine Merge candidates. b. For example, it can refer to the derivation of affine AMVP candidates. c. For example, it can refer to both the derivation of affine Merge candidates and affine AMVP candidates. d. For example, it can refer to the derivation of history-based affine candidates, non-adjacent affine candidates, and regression-based affine candidates. b. The logic / rules / process for similarity / consistency / deduplication checking can refer to comparing one or more of the following elements associated with the first affine candidate with these elements associated with the second affine candidate: a. Inter-frame direction (prediction direction); b. Affine type (e.g., 6-parameter affine or 4-parameter affine); c. Sub-block Merge type (e.g., sbTMVP or affine); d. Bcw index; e. LIC flag; f. Reference index; g. Motion vector (e.g., horizontal component and / or vertical component); h. Control point motion vector (CPMV); i. First CPMV (e.g., upper left CPMV) and / or second CPMV (e.g., upper right CPMV) and / or third CPMV (e.g., lower left CPMV); j. Horizontal displacement and / or vertical displacement between the first CPMV and the second CPMV (e.g., absolute difference between the horizontal components and / or vertical components of the first CPMV and the second CPMV); k. Horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV (e.g., absolute difference between the horizontal components and / or vertical components of the first CPMV and the third CPMV). c. A consistency check can be performed on the comparison of the elements listed in item b. a. For example, if the elements associated with the second candidate are the same as those associated with the first candidate, the second candidate is not added to the affine candidate list. d. Similarity checks can be performed on the comparisons listed in item b of the project. a. For example, if the elements associated with the second candidate are similar to the elements associated with the first candidate, the second candidate is not added to the affine candidate list. b. For example, "similar" can refer to a threshold-based comparison. i. For example, the absolute difference is less than the threshold. ii. For example, the absolute difference is not greater than the threshold. c. For example, the threshold can depend on the block size, such as the width and / or height. i. For example, the threshold is determined adaptively according to the block width / height. ii. For example, a smaller threshold can be set for a smaller block size, while a larger threshold can be set for a larger block size. d. For example, the threshold can be a predefined fixed value (such as 0 or 1). e. For example, similarity checks can be performed separately for the horizontal and vertical components of the motion vector. f. For example, similarity checks can be performed separately for the horizontal and vertical components of the control point motion vector. e. Whether to apply similarity / consistency / deduplication checks to the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV can depend on the affine type (e.g., 6-parameter affine or 4-parameter affine). a. For example, similarity / consistency / deduplication checks on the horizontal displacement and / or vertical displacement between the first CPMV and the third CPMV can be performed only when the affine types of the first affine candidate and the second affine candidate are 6-parameter. f. For example, for Merge list / candidate deduplication, coding information other than motion similarity can be checked (e.g., BCW index, LIC flag, etc.). a. For example, if the second candidate has the same / similar motion vector as the first candidate but a different BCW index value, the second candidate can be considered non-redundant / dissimilar to the first candidate. b. For example, if the second candidate has the same / similar motion vector as the first candidate but a different LIC flag value, the second candidate can be considered non-redundant / dissimilar to the first candidate. c. For example, Merge list / candidate deduplication can refer to a regular Merge list. d. For example, Merge list / candidate deduplication can refer to an MMVD-based Merge list. e. For example, Merge list / candidate deduplication can refer to a TM-based Merge list. f. For example, the Merge list / candidate deduplication may refer to a BM-based Merge list. g. For example, the Merge list / candidate deduplication may refer to a DMVR-based Merge list. h. For example, the Merge list / candidate deduplication may refer to an affine DMVR Merge list. i. For example, the Merge list / candidate deduplication may refer to a CIIP Merge list (with or without using TM). j. For example, the Merge list / candidate deduplication may refer to a GPM Merge list (with or without using TM). k. For example, the Merge list / candidate deduplication may refer to a sbTMVP Merge list (with or without using TM). 4.2. Regarding the second issue of the codec tools for screen content tools, the following methods are proposed: a. For a video unit, at least one of the following codec tools may be disallowed or constrained or prohibited. a) CIIP and / or its variants (e.g., CIIP PDPC, CIIP TM, CIIP TIMD, etc.). b) OBMC and / or its variants (e.g., OBMC TM, etc.). c) TIMD and / or its variants. d) DIMD and / or its variants. e) MHP and / or its variants. f) DMVR and / or its variants. g) interTM and / or its variants. h) CCALF and / or its variants. i) CCSAO and / or its variants. b. Whether to disallow or constrain or prohibit the codec tool for a video unit may depend on the profile / level / layer. c. Whether to disallow or constrain or prohibit the codec tool for a video unit may depend on whether the video unit belongs to a specific video type. a) The specific video type may refer to a screen content video. d. In addition, for a video sequence or a group of pictures or a picture or a slice, the codec tool may be disallowed or constrained or prohibited. a) Such a constraint or disallowance may be reflected by a bitstream constraint. b) Such a constraint or disallowance or allowance may be reflected by a syntax element (e.g., a flag) signaled in the bitstream. e. Such a constraint can be imposed, or not allowed or allowed, on the specific coding / decoding tools listed in item a by signaling a syntax element (e.g., a flag) in the bitstream. a) A syntax element can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. f. Such a constraint can be imposed, or not allowed or allowed, on more than one coding / decoding tool listed in item a by signaling a single syntax element (e.g., a flag). a) A syntax element can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, such as in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. g. Different weight values / factors / tables / sets can be allowed to mix multiple prediction hypotheses of video units coded with coding / decoding tool X. a) For example, which weight value / factor / table / set is allowed to mix multiple prediction hypotheses of a video unit can depend on the content type. b) For example, a first weight value / factor / table / set can be allowed to mix multiple prediction hypotheses of a first video unit, while a second weight value / factor / table / set can be allowed to mix multiple prediction hypotheses of a second video unit. c) For example, the first video unit can belong to a video sequence captured by a camera. d) For example, the second video unit can belong to a screen content video sequence. e) For example, assume that two prediction hypotheses are mixed into the final prediction of a video unit: i. For the first type of video unit, the final prediction block of the video unit coded with coding / decoding tool X can be "exactly equal to the first prediction hypothesis" or "exactly equal to the second prediction hypothesis". 1. For example, the allowed weight value / factor / table / set of the video unit can be equal to 0 or 1. ii. For the second type of video unit, the final mixed prediction block of the video unit coded with coding / decoding tool X can be "a fusion of both the first prediction hypothesis and the second prediction hypothesis". 1. For example, the allowed weight value / factor / table / set for the video unit can be a fraction between 0 and 1 (e.g., the fractional value of the weight of a specific prediction hypothesis can be quantized to an integer in the codec). f) For example, the final prediction (e.g., the mixed) of a video unit with coding / decoding tool X can be further mixed / weighted / fused with another video unit coded with another coding / decoding tool. g) For example, the specific codec tool X can be CIIP and / or its variants. h) For example, the specific codec tool X can be GPM and / or its variants. i) For example, the specific codec tool X can be MHP and / or its variants. j) For example, the specific codec tool X can be OBMC and / or its variants. k) For example, the specific codec tool X can be TIMD hybrid mode and / or its variants. l) For example, the specific codec tool X can be DIMD hybrid mode and / or its variants. 4.3. Regarding the third problem of the general use of KLT for the transformation process, the following methods are proposed: a. More than two KLT kernels can be allowed in the codec. b. The KLT kernel can be allowed for the primary transformation and / or the secondary transformation. c. The KLT kernel can be allowed for the chrominance component. a. For example, different KLT kernels can be used for the luminance component and the chrominance component of the video unit. i. Alternatively, all color components of the video unit can share the same KLT kernel. b. For example, different KLT kernels can be used for the chrominance Cb component and the chrominance Cr component of the video unit. i. Alternatively, the chrominance Cb component and the chrominance Cr component of the video unit can share the same KLT kernel. d. A pair {KLT, flipped-KLT} can be allowed / used for the {horizontal, vertical} transformation or {vertical, horizontal} transformation of the video unit. a. For example, {KLT, flipped-KLT} represents a pair of transformation kernels including a first KLT and a second KLT, where the transformation coefficient matrix of the second KLT (i.e., the flipped-KLT) can be the transposed matrix of the transformation coefficient matrix of the first KLT. e. The horizontal or vertical transformation type of the video unit can be selected from {KLT-X, flipped-KLT-X, DCT-Y, DST-Z}, where X / Y / Z are constants (e.g., Y = 8, Z = 7). a. For example, if more than one KLT kernel is defined in the codec, the more than one KLT kernel can be represented as KLT-X, such as X being an integer value, such as 1, 2, 3,..., n, i.e., KLT-1, KLT-2, KLT-3,..., KLT-n. b. For example, for the transformation block, KLT can be used for the horizontal (or vertical) transformation, while the flipped-KLT can be used for the vertical (or horizontal) transformation. c. For example, for a transform block, KLT can be used for horizontal (or vertical) transform, and non-KLT can be used for vertical (or horizontal) transform. d. For example, DCT2-KLT and KLT-DCT2 can be allowed for horizontal-vertical or vertical-horizontal transform of a block. e. For example, DST7-DCT2, DCT2-DST7, DCT8-DCT2, and DCT2-DCT8 can be allowed for horizontal-vertical or vertical-horizontal transform of a block. f. For example, a video unit can be encoded / decoded using specific prediction / transform / filter mode / technique. f. For a video unit encoded / decoded in a specific mode, KLT can be the only transform type. a. In one example, the specific mode can be SBT. i. In one example, for different SBT modes (SBT_horizontal_split, SBT_vertical_split, SBT_half_split, SBT_quad_split), KLT can be used for both the horizontal and vertical dimensions. b. In one example, the specific mode can be MIP. c. In one example, the same KLT can be used for both the horizontal and vertical dimensions of a video unit encoded / decoded in a specific mode. d. In one example, different KLTs can be used for the horizontal and vertical dimensions of a video unit encoded / decoded in a specific mode. g. A KLT-based transform type can additionally be allowed for a video unit. a. For example, in addition to existing MTS options (e.g., MTS indices from 0 to 5), a KLT-based transform type can be explicitly signaled. b. For example, a first syntax element (e.g., a flag) can be signaled to indicate whether KLT is used for a video unit. i. Additionally, alternatively, if KLT is used for a video unit, a second syntax element (e.g., a flag or an index) can be signaled to indicate whether and / or which KLT is used for horizontal transform and / or vertical transform. c. For example, a syntax element (e.g., an index) can be signaled to indicate which KLT is used for a video unit. i. For example, a syntax element can be signaled to indicate which pair of KLTs is used for horizontal and vertical transforms for a video unit. ii. For example, a syntax element can be signaled to indicate which KLT is used for a specific size (e.g., width or height) of a video unit. iii. For example, signal transmission syntax elements can be used for both non-KLT transforms and KLT transforms. 1. For example, indices 0 to 1 indicate DCT2-DCT2 pairs and transform skip; while indices 2 to N indicate non-DCT2-DCT2 pairs and non-transform skip pairs including combinations of non-KLT and KLT. iv. For example, signal transmission syntax elements can be used when KLT is applied to a video unit. 1. For example, a signal transmission index (possible values starting from 0) can be used to indicate which KLT pair is used (i.e., at least one of the horizontal or vertical directions is using KLT). 2. For example, in this case, the intra (and / or inter) MTS index may not be signaled (e.g., disabled) to the video unit. d. For example, which KLT-based transform type is used for a video unit can be determined implicitly. i. For example, the implicit determination can be based on the dimension / shape / size of the video unit. ii. For example, for a specific length of the block size (e.g., width or height), a specific KLT type is used in this direction without signaling. h. A KLT-based transform type can be applied to replace a specific existing transform type of a video unit. a. For example, the existing transform type to be replaced can be DST7, or DCT8, or DCT2. b. For example, a separable transform can be replaced by KLT. c. For example, the primary transform and / or the secondary transform can be replaced by KLT. i. For example, the KLT can be an inseparable KLT. j. For example, the KLT can be a separable KLT. k. For example, the video unit to which KLT is applied can be intra-coded. l. For example, the video unit to which KLT is applied can be inter-coded. m. For example, KLT can be used as the primary transform. n. For example, KLT can be used as the secondary transform. 4.4. Regarding the fourth issue of mode-dependent KLT for the transform process, the following method is proposed: a. Which KLT kernel is used for a video unit can depend on a combination of at least one of the following coding information types: a. The coding mode of the video unit. b. The size of the video unit. c. The motion vector of the video unit. d. Quantization parameters of the video unit. e. Temporal layer of the video unit. b. Which KLT kernel is used for the video unit may depend on the coding / decoding mode of the video unit. a. For example, it may be based on whether SBT is used for the video unit. b. For example, it may be based on whether implicit MTS is used for the video unit. c. For example, it may be based on whether explicit MTS is used for the video unit. d. For example, it may be based on whether intra MTS is used for the video unit. e. For example, it may be based on whether inter MTS is used for the video unit. f. For example, it may be based on whether LFNST is used for the video unit. g. For example, it may be based on whether IBC and / or its variant modes are used for the video unit. h. For example, it may be based on whether PLT and / or its variant modes are used for the video unit. i. For example, it may be based on whether the intra prediction mode is used for the video unit. j. For example, it may be based on whether ISP and / or its variant modes are used for the video unit. k. For example, it may be based on whether MIP and / or its variant modes are used for the video unit. l. For example, it may be based on whether DIMD and / or its variant modes are used for the video unit. m. For example, it may be based on whether TIMD and / or its variant modes are used for the video unit. n. For example, it may be based on whether LM / CCLM / CCCM / GLM and / or their variant modes are used for the video unit. o. For example, it may be based on whether the inter prediction mode is used for the video unit. p. For example, it may be based on whether the AMVP mode is used for the video unit. q. For example, it may be based on whether the Merge mode is used for the video unit. r. For example, it may be based on whether inter / intra / IBC template matching and / or their variant modes are used for the video unit. s. For example, it may be based on whether DMVR and / or its variant modes are used for the video unit. t. For example, it may be based on whether the sub-block prediction mode is used for the video unit. u. For example, it may be based on whether affine and / or its variant modes are used for the video unit. v. For example, it can be based on whether the sbTMVP and / or its variant mode is used for a video unit. w. For example, it can be based on whether the hybrid / fusion / multi-hypothesis mode and / or its variant mode is used for a video unit. i. In one example, it can be based on whether the hybrid / fusion / multi-hypothesis mode includes an intra coding part, such as GPM inter-intra, GPM intra, CIIP, MHP using intra, partitioned GPM, etc. x. For example, it can be based on whether the GPM and / or its variant mode is used for a video unit. y. For example, it can be based on whether the CIIP and / or its variant mode is used for a video unit. z. For example, it can be based on whether the MHP and / or its variant mode is used for a video unit. aa. For example, it can be based on whether the OBMC and / or its variant mode is used for a video unit. bb. For example, it can be based on whether the LIC and / or its variant mode is used for a video unit. c. Whether to use the KLT and / or which KLT kernel is used for a video unit can depend on the size of the video unit. a. Different KLTs can be applied to blocks with different sizes. b. For example, it can be based on whether the width (W) and / or height (H) of the video unit satisfies predefined conditions, such as one or more combinations of the following: i. W < T1 or W <= T1, where T1 can be 8 or 16 or 32 or 64. ii. W > T2 or W >= T2, where T2 can be 2 or 4 or 8. iii. H < T3 or H <= T3, where T3 can be 8 or 16 or 32 or 64. iv. H > T4 or H >= T4, where T4 can be 2 or 4 or 8. v. W / H < T5 or W / H <= T5, where T5 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vi. W / H > T6 or W / H >= T6, where T6 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. vii. H / W < T7 or H / W <= T7, where T7 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. viii. H / W > T8 or H / W >= T8, where T8 can be 1 / 8 or 1 / 4 or 1 / 2 or 1 or 2 or 4 or 8 or 16. ix. W == T9, where T9 can be 8 or 16 or 32 or 64. x. H == T10, where T10 can be 8 or 16 or 32 or 64. d. Which KLT kernel is used for a video unit can depend on the motion vector of the video unit. a. In one example, it depends on the magnitude of the motion vector. e. Which KLT kernel is used for a video unit can depend on the quantization parameter of the video unit. a. In one example, it depends on the base QP derived / signaled at a syntax level higher than the slice level (e.g., PPS or SPS). b. In one example, it depends on the slice QP. c. In one example, it depends on the QP of the coding unit. f. Which KLT kernel is used for a video unit can depend on the temporal layer of the video unit. a. In one example, it depends on whether it is in temporal layer 0 or in temporal layer 1, or in temporal layer 2, or.... g. Depending on the coding mode and / or size of the video unit, different KLT kernels can be allowed for different video units. a. In one example, a first KLT set can be used for a first mode set, while a second KLT set can be used for a second mode set. i. For example, the first KLT set includes at least one type of KLT kernel. ii. For example, the second KLT set includes at least one type of KLT kernel. iii. For example, the first mode set includes at least one type of prediction / transformation / filter mode. iv. For example, the second mode set includes at least one prediction / transformation / filter mode. b. In one example, a video unit coded with more than one different type of prediction / transformation / filter mode can use the same KLT kernel. c. Alternatively, video units coded with different types of prediction / transformation / filter modes can use different KLT kernels. d. For example, one type of prediction / transformation / filter mode can be a sub-block based prediction mode (e.g., affine, sbTMVP, etc.). e. For example, one type of prediction / transformation / filter mode can be an affine based prediction mode (e.g., affine AMVP, affine Merge, etc.). f. For example, one type of prediction / transformation / filter mode can be a hybrid / fusion / multi-hypothesis based prediction mode (e.g., GPM inter-intra, GPM intra, CIIP, etc.). g. For example, one type of prediction / transformation / filter mode can be SBT and its variants. h. For example, one type of prediction / transformation / filter mode can be ISP and its variants. i. For example, one type of prediction / transformation / filter mode can be IBC and its variants. 4.5. Regarding the fifth problem of intra-mode derivation / mapping for intra MTS and LFNST indices, the following methods are proposed: a. The intra-mode derived from the information of neighboring samples can be used for the chrominance transformation process. a. For example, in the case of single-tree and / or double-tree, the chrominance transformation process can refer to the primary transformation of the chrominance component (e.g., intra MTS, inter MTS, KLT, DCT2,...). b. For example, in the case of single-tree and / or double-tree, the chrominance transformation process can refer to the secondary transformation of the chrominance component (e.g., LFNST). c. For example, the template constructed from the upper neighboring samples and / or the left neighboring samples can be used to derive the intra-mode. i. The gradient of the template samples can be used. ii. The maximum histogram amplitude value is constructed from the gradient (e.g., gradient histogram). iii. The DIMD-based method can be used. iv. The TIMD-based method can be used. v. The intra-mode can be determined from the template samples (e.g., according to the measurement based on SAD / SATD cost) based on the preset intra-mode candidates. vi. How many rows / columns and / or which neighboring samples can be based on the multi-reference row index. vii. How many rows / columns and / or which neighboring samples can be predefined (e.g., one row / column, or four rows / columns, etc.). viii. The neighboring samples in different rows / columns can have different weights for cost calculation. d. The derived intra-mode can be generated for video units encoded / decoded using the following modes: i. CCLM mode and / or its variants. ii. CCCM mode and / or its variants. iii. GLM mode and / or its variants. iv. TIMD chrominance mode and / or its variants. v. DIMD chroma mode and / or its variants. vi. Planar mode and / or its variants. vii. Chroma intra mode with a mode index larger than a specific number, where the specific number represents an angular intra mode (such as, the number equal to 80). e. The derived intra mode can be used for MTS transform class index derivation. f. The derived intra mode can be used for MTS transform pair index derivation. g. The derived intra mode can be used for MTS transform index derivation. h. The derived intra mode can be used for LFNST transform set index derivation. i. The derived intra mode can be used for LFNST transpose flag derivation. b. The intra mode derived from the information of neighboring samples can be used for the luminance transformation process. a. For example, the luminance transformation process can refer to the primary transformation of the luminance component (e.g., intra MTS, inter MTS, KLT, DCT2,...). b. For example, the luminance transformation process can refer to the secondary transformation of the luminance component (e.g., LFNST). c. For example, a template constructed from upper neighboring samples and / or left neighboring samples can be used to derive the intra mode. i. The gradient of the template samples can be used. ii. The maximum histogram amplitude value is constructed from the gradient (e.g., gradient histogram). iii. The DIMD-based method can be used. iv. The TIMD-based method can be used. v. The intra mode can be determined from the template samples (e.g., according to the measurement based on SAD / SATD cost) based on the preset intra mode candidates. vi. How many rows / columns and / or which neighboring samples can be based on the multi-reference row index. vii. How many rows / columns and / or which neighboring samples can be predefined (e.g., one row / column, or four rows / columns, etc.). viii. The neighboring samples in different rows / columns can have different weights for cost calculation. d. The derived intra mode can be generated for video units encoded and decoded using the following modes: i. Intra template matching and / or its variants. ii. MIP mode and / or its variants. iii. ISP mode and / or its variants. iv. Prediction obtained by mixing at least one intra mode and another mode. v. TIMD mixing mode and / or its variants. vi. DIMD mixing mode and / or its variants. vii. GPM intra mode and / or its variants. viii. Partitioned GPM mode and / or its variants. ix. CIIP mode and / or its variants. x. MHP with intra mode mixing. xi. Screen content coding / decoding tool with intra mode mixing. xii. Planar mode. xiii. Planar horizontal mode. xiv. Planar vertical mode. e. The derived intra mode can be used for MTS transform class index derivation. f. The derived intra mode can be used for MTS transform pair index derivation. g. The derived intra mode can be used for MTS transform index derivation. h. The derived intra mode can be used for LFNST transform set index derivation. i. The derived intra mode can be used for LFNST transpose flag derivation. c. The intra mode derived from the predicted samples of the current block can be used for the transform process. a. The predicted samples can refer to the prediction of the current video unit. b. The predicted samples can refer to the prediction of the template samples of the current video unit. c. The transform process can refer to the primary transform (e.g., intra MTS, inter MTS, KLT, DCT2,...). d. The transform process can refer to the secondary transform (e.g., LFNST). e. The derived intra mode can be generated for video units coded / decoded using the following modes: i. Intra template matching and / or its variants. ii. MIP mode and / or its variants. iii. ISP mode and / or its variants. iv. Prediction obtained by mixing at least one intra mode and another mode. v. TIMD mixing mode and / or its variants. vi. DIMD mixing mode and / or its variants. vii. GPM intra mode and / or its variants. viii. Partitioned GPM mode and / or its variants. ix. CIIP mode and / or its variants. x. MHP with intra mode mixing. xi. Screen content codec tool with intra mode mixing. xii. Planar mode. xiii. Planar horizontal mode. xiv. Planar vertical mode. xv. CCLM mode and / or its variants. xvi. CCCM mode and / or its variants. xvii. GLM mode and / or its variants. xviii. Chrominance intra mode with mode index larger than a specific number, where the specific number represents an angular intra mode (such as, the number equal to 80). f. The derived intra mode can be used for MTS transform class index derivation. g. The derived intra mode can be used for MTS transform pair index derivation. h. The derived intra mode can be used for MTS transform index derivation. i. The derived intra mode can be used for LFNST transform set index derivation. j. The derived intra mode can be used for LFNST transpose flag derivation. d. For example, the derived intra mode can be used to index the intra MTS transform class and / or transform pair and / or transform set for the block coded by MIP. a. The derived intra mode can be based on the predicted samples of the MIP block before MIP prediction upsampling. b. The derived intra mode can be based on the predicted samples of the MIP block after MIP matrix vector multiplication. c. The derived intra mode can be based on the horizontal / vertical gradient of the predicted samples of the MIP block. d. The derived intra mode can be based on the maximum histogram amplitude value constructed from the gradient of the MIP block (such as, gradient histogram). e. The derived intra mode can be based on the neighboring sample values of the MIP block. f. Alternatively, the derived intra mode can be a fixed / predefined mode (such as, other than the planar mode), regardless of the current predicted samples and neighboring samples of the current block. e. For example, the derived intra mode can be used to index the LFNST transform set and / or LFNST transpose flag for the block coded by intra template matching (such as, intraTMP, intraTM, etc.). a. The derived intra mode can be based on the predicted samples of the block coded with intraTMP. b. The derived intra mode can be based on the horizontal / vertical gradients of the predicted samples of the block coded with intraTMP. c. The derived intra mode can be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the block coded with intraTMP. d. The derived intra mode can be based on the neighboring sample values of the block coded with intraTMP. e. The derived intra mode can be based on the intra mode of the reference block of the block coded with intraTMP. f. The derived intra mode can be based on the gradients of the reference block of the block coded with intraTMP. g. The derived intra mode can be based on the angle of the block vector (or motion vector) of the block coded with intraTMP. h. The derived intra mode can be based on the intra mode information stored in the history-based intra mode cache for the block coded with intraTMP. i. Alternatively, the derived intra mode can be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current predicted samples and neighboring samples of the current block. f. For example, the MTS index can be signaled for the blocks that can be coded for intra template matching (e.g., intraTMP, intraTM, etc.). a. For example, the intra MTS index can be signaled for the blocks coded with intraTMP. b. For example, the inter MTS index can be signaled for the blocks coded with intraTMP. c. The derived intra mode can be used to index the MTS transform set, MTS transform class, MTS transform pair for the block coded with intraTMP. i. The derived intra mode can be based on the predicted samples of the block coded with intraTMP. ii. The derived intra mode can be based on the horizontal / vertical gradients of the predicted samples of the block coded with intraTMP. iii. The derived intra mode can be based on the maximum histogram magnitude value constructed from the gradients (e.g., gradient histogram) of the block coded with intraTMP. iv. The derived intra mode can be based on the neighboring sample values of the block coded with intraTMP. v. The derived intra mode can be based on the intra mode of the reference block of the block coded / decoded by intraTMP. vi. The derived intra mode can be based on the gradient of the reference block of the block coded / decoded by intraTMP. vii. The derived intra mode can be based on the angle of the block vector (or motion vector) of the block coded / decoded by intraTMP. viii. The derived intra mode can be based on the intra mode information stored in the history-based intra mode cache for the block coded / decoded by intraTMP. ix. Alternatively, the derived intra mode can be a fixed / predefined mode (e.g., other than the planar mode), regardless of the current prediction samples and neighboring samples of the current block. d. Alternatively, the implicit MTS core can be applied to the block coded / decoded by intraTMP. i. For example, it can be determined based on the derived intra mode. ii. For example, it can be determined based on the block shape / size / dimension (e.g., width / height). iii. For example, it can be determined based on the value of the transform coefficients. iv. For example, for such an MTS core, it is not signaled by syntax elements (e.g., index). e. In addition, KLT can be used for the block coded / decoded by intraTMP. f. Alternatively, the main transform cores other than DCT2 can be restricted for the block coded / decoded by intraTMP. g. For example, the intra MTS transform class and / or the intra MTS transform pair and / or the LFNST transform set and / or the LFNST transpose flag of the TIMD hybrid mode can be derived based on the final prediction samples of the TIMD block. a. Alternatively, it can be derived based on neighboring sample information (such as gradient, TIMD information, DIMD information, etc.). b. Alternatively, it can be based on a fixed / predefined mode (e.g., other than the planar mode), regardless of the current prediction samples and neighboring samples of the current block. h. For example, the intra MTS transform class and / or the intra MTS transform pair and / or the LFNST transform set and / or the LFNST transpose flag of the DIMD hybrid mode can be derived based on the final prediction samples of the DIMD block. a. Alternatively, it can be derived based on neighboring sample information (such as, gradient, etc.). b. Alternatively, it can be based on a fixed / predefined pattern (e.g., other than the planar pattern), regardless of the current predicted samples and neighboring samples of the current block. i. For example, the intra MTS transform class and / or intra MTS transform pair and / or LFNST transform set and / or LFNST transpose flag of the planar (and / or planar horizontal and / or planar vertical) pattern can be derived based on the final predicted samples of the block. a. Alternatively, it can be derived based on neighboring sample information such as gradient, TIMD information, DIMD information, etc. b. Alternatively, it can be based on a fixed / predefined pattern (e.g., other than the planar pattern), regardless of the current predicted samples and neighboring samples of the current block. 4.6. In one example, how to encode and decode the residual block can depend on the selected transform. 4.7. In one example, the KLT can be separable or non-separable. 4.8. In one example, whether or how to apply the KLT can depend on the sum of the absolute values of the transform coefficients. a. For example, whether to allow the KLT for the block encoded / decoded intra (and / or inter) can be based on the sum of the absolute values of the transform coefficients of the block. b. For example, whether to allow the KLT for a specific block size / dimension / shape / orientation can be based on the sum of the absolute values of the transform coefficients of the block. c. For example, whether to apply a specific KLT-based transform kernel / set / pair to the block encoded / decoded intra (and / or inter) can be based on the sum of the absolute values of the transform coefficients of the block. d. For example, the number of allowed KLT transform kernels / sets / pairs can be determined based on the sum of the absolute values of the transform coefficients of the block. e. For example, the sum of the absolute values of the transform coefficients of the block can be compared with at least one threshold. a. The threshold can be predefined. b. The threshold can be equal to a fixed value. c. The threshold can be adaptively determined based on predefined rules (e.g., block size / dimension, sequence resolution, prediction mode, whether it is screen content, etc.). 4.9. Regarding the sixth issue of sample refinement for a specific coding tool (e.g., intraTMP, IBC, etc.), the following methods are proposed: a. For example, the sample refinement process can be applied to the motion-compensated (or block-vector-compensated) prediction of a specific video block. a. For example, the specific video block can be encoded / decoded by intraTMP. b. For example, a specific video block may be encoded / decoded by IBC. c. For example, the block may be a screen content video block. d. For example, the block may be a camera-captured content video block. e. For example, the sample refinement process may refer to local illumination compensation (i.e., LIC). f. For example, the sample refinement process may refer to overlapping block motion compensation (i.e., OBMC). g. For example, the sample refinement process may be applied to all samples within a specific block. h. For example, the sample refinement process may be applied to some samples within a specific block. i. For example, only the boundary samples at the left and / or top boundaries of a specific block may be refined. i. For example, uniform / consistent refinement parameters (e.g., weights, biases, scaling factors, etc.) may be used for all samples to be refined. j. For example, at least one sample may use refinement parameters (e.g., weights, biases, scaling factors, etc.) different from those of another sample. k. For example, sample-based (e.g., position-dependent) refinement parameters (e.g., weights, biases, scaling factors, etc.) may be used. i. For example, refinement parameters may be assigned based on the position of each sample relative to the left and / or upper template (or block boundary). l. For example, refinement parameters (e.g., weights, biases, scaling factors, etc.) may be derived based on the difference / error / cost / SAD / MRSAD / SATD between the samples adjacent to the current block and the samples adjacent to the second block. i. For example, the second block may be pointed to by the motion vector (block vector) of the current block. ii. For example, a template-based method may be used. b. For example, the difference / error / cost measurement based on mean removal may be used as a criterion for searching for the motion vector (block vector) of a specific video block. a. For example, a specific video block may be encoded / decoded by intraTMP. b. For example, a specific video block may be encoded / decoded by intraTMP and LIC. c. For example, a specific video block may be encoded / decoded by intraTMP and OBMC. d. For example, a specific video block may be encoded / decoded by IBC. e. For example, a specific video block may be encoded / decoded by IBC and LIC. f. For example, a specific video block can be encoded and decoded by IBC and OBMC. g. For example, the block can be a screen content video block. h. For example, the block can be a camera-captured content video block. i. For example, a cost function based on mean removal can be used to measure the difference / error / cost / SAD / MRSAD / SATD between two templates / blocks to be compared. i. For example, the first template can be constructed from samples adjacent to a specific block, and the second template can be constructed from samples adjacent to the search candidates of the specific block. ii. For example, a method based on mean removal can first calculate the average / mean value between two templates. 1. For example, first accumulate all the sample differences between the first template and the second template, and then the accumulated value is further divided by the total number of samples in one template to obtain the mean (for example, the total number of samples in the first template is equal to the total number of samples in the second template). iii. For example, when calculating the difference / error / cost / SAD / SATD, each sample difference in the sample differences can be further subtracted by the mean, and then the subtracted value is used to calculate the final difference / error / cost / SAD / MRSAD / SATD between the two templates. 4.10. Regarding the seventh question about DMVR for extending codec tools, the following methods are proposed: a. For example, DMVR refinement can be used for blocks in sub-block encoding and decoding (such as sbTMVP, affine, etc.). a. For example, DMVR can be a PU / CU-level DMVR process that outputs motion offsets based on PUs / CUs. b. For example, DMVR can be a sub-PU / sub-CU (such as 16×16, 8×8, 4×4, etc.)-level DMVR process that outputs motion offsets based on sub-PUs / sub-CUs (such as 16×16, 8×8, 4×4, etc.). c. For example, DMVR can be a multi-pass DMVR process that includes both a PU / CU-level DMVR process and a sub-PU / sub-CU-level DMVR process. b. For example, DMVR refinement can be applied to the motion of blocks encoded and decoded by sbTMVP. a. For example, the motion can be bi-directionally encoded and decoded. b. For example, the motion can meet DMVR conditions, such as pointing to a forward reference picture and a backward reference picture with the same POC distance. c. For example, the motion can be used to find a first CU-level reference block in the forward reference picture and a second CU-level reference block in the second reference picture. d. For example, the motion may be a sub-block based motion vector to derive the prediction value of the block coded / decoded by sub-TMVP. e. For example, to calculate the bilateral matching cost for the block coded / decoded by sbTMVP, sub-block based motion compensation may be performed to generate prediction values in two prediction directions. Then, the bilateral matching cost is calculated as the distortion between the two prediction values. f. For example, a motion offset may be added to each motion vector of each sub-block in the sub-block. i. For example, after the reference block at the PU / CU level is retrieved, then the sub-block motion vector is derived, and it is assumed that a motion offset is added to each motion vector of each sub-block in the sub-block to perform sub-block based motion compensation. ii. For example, the same motion offset (e.g., delta) may be added to the motion vectors of all sub-blocks in one prediction direction. iii. For example, the opposite motion offset (e.g., -delta) may be added to the motion vectors of all sub-blocks in the other prediction direction. iv. For example, the motion offset may be based on the steps of DMVR refinement. g. For example, a motion offset may be added to the motion candidates for finding two reference blocks at the CU level. i. For example, before the reference block at the PU / CU level is retrieved, a motion offset (e.g., delta) may be added to one direction of the motion candidates, and the opposite offset (e.g., -delta) may be added to the other direction of the motion candidates, and then the reference block at the PU / CU level is retrieved based on the refined motion candidates. ii. For example, the motion offset may be based on the DMVR refinement steps. 4.11. Regarding the eighth issue of the coding / decoding tools for multi-model / hypothesis mixing, the following methods are proposed: a. For example, the prediction may be generated based on mixing the prediction of the LM-T mode and the prediction of the LM-L mode. b. For example, the prediction of the LM-TL mode may be generated by mixing the first prediction derived based on the upper neighbor and the second prediction derived based on the left neighbor. c. For example, the prediction of the LIC mode may be generated by mixing the first LIC prediction derived based on the upper neighbor / template and the second LIC prediction derived based on the left neighbor / template. d. For example, it may determine which direction of the neighbor / template (e.g., upper or left) has a major impact on the final prediction. a. For example, it may determine whether the final prediction mainly comes from the upper neighbor / template or the left neighbor / template. b. For example, if the difference / error / cost / SAD / MRSAD / SATD of the neighbors / templates in one direction (e.g., upward, left) is less than (or greater than) that in another direction, the first direction can be considered the main contributor. e. For example, the blending weights of the two predictions can be uniform. a. For example, weight A can be assigned to all samples of the first prediction, and weight B can be assigned to all samples of the second prediction. b. For example, the value of weight A can be equal to the value of weight B. c. For example, the value of weight A can be not equal to the value of weight B. i. For example, the weights can be determined based on which prediction makes more contributions. ii. For example, a greater weight can be assigned to the prediction that makes more contributions. f. For example, the blending weights of the two predictions can be sample-based. a. For example, for different samples, the blending weights can be different. b. For example, weights A0, A1, A2,... can be assigned to samples 0, 1, 2,... of the first prediction, and weights B0, B1, B2,... can be assigned to samples 0, 1, 2,... of the second prediction. c. For example, the value of the blending weights can depend on the position. i. For example, if the position of the sample is closer to the direction determined to be the main contributor (e.g., upward or left), the blending weight of the sample can be greater. 4.12. Whether and / or how to apply the methods disclosed above can be signaled at the sequence level / picture group level / picture level / strip level / slice group level, e.g., in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / strip header / slice group header. 4.13. Whether and / or how to apply the methods disclosed above can be signaled at the PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / strip / slice / sub-picture / other types of regions containing more than one sample or pixel. 4.14. The binary bits of the syntax elements signaled in the disclosed methods can be context-coded or bypass-coded. 4.15. Whether and / or how to apply the methods disclosed above can depend on the coded information, such as block size, color format, single / double tree segmentation, color component, strip / picture type.
[0110] More details of embodiments of the present disclosure related to transforms in image / video coding / decoding, screen content coding / decoding (SCC), affine, deduplication checking, prediction refinement with motion compensation (MC), and weighted-based schemes will be described below. Embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be construed in a narrow manner. Additionally, these embodiments can be applied individually or in combination in any way.
[0111] As used herein, the term "block" may represent a color component, sub-picture, picture, stripe, slice, coding tree unit (CTU), CTU row, CTU group, coding unit (CU), prediction unit (PU), transform unit (TU), coding tree block (CTB), coding block (CB), prediction block (PB), transform block (TB), sub-blocks of a video block, sub-regions within a video block, a video processing unit including a plurality of samples / pixels, etc. A block can be rectangular or non-rectangular.
[0112] Figure 42 A flowchart of a method 4200 for video processing according to some embodiments of the present disclosure is shown. Method 4200 can be implemented during the conversion between a current video block of a video and a bitstream of the video. As Figure 42 shown, method 4200 starts at 4202, where a set of motion vectors for the current video block is obtained. The current video block is coded / decoded using sub-block-based coding tools. By way of example and not limitation, sub-block-based coding tools can include sub-block-based temporal motion vector prediction (SbTMVP) mode, affine mode, etc.
[0113] At 4204, a decoder-side motion vector refinement (DMVR) process is applied to the set of motion vectors. In some embodiments, the DMVR process can include: a prediction unit (PU)-level DMVR process that outputs a PU-based motion offset; a coding unit (CU)-level DMVR process that outputs a CU-based motion offset; a sub-PU-level DMVR process that outputs a sub-PU-based motion offset; a sub-CU-level DMVR process that outputs a sub-CU-based motion offset; a multi-pass DMVR process that can include a PU-level DMVR process and a sub-PU-level DMVR process; or a multi-pass DMVR process that can include a CU-level DMVR process and a sub-CU-level DMVR process. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0114] At 4206, a conversion is performed based on the application. In some embodiments, the conversion can include encoding the current video block into the bitstream. Alternatively or additionally, the conversion can include decoding the current video block from the bitstream.
[0115] In view of the above, DMVR processing is used for video blocks encoded and decoded by sub-block based encoding and decoding tools. Compared with conventional solutions, the proposed method can advantageously improve the encoding and decoding efficiency and quality.
[0116] In some embodiments, a set of motion vectors may be bi-directionally encoded and decoded. In some embodiments, a set of motion vectors may satisfy DMVR conditions for applying DMVR processing. By way of example and not limitation, the DMVR conditions may include motion vectors pointing to a forward reference picture and a backward reference picture having the same POC distance. For example, a first motion vector in a set of motion vectors points to a forward reference picture for a current picture including a current video block, a second motion vector in the set of motion vectors points to a backward reference picture for the current picture, and the picture order count (POC) distance between the forward reference picture and the current picture is the same as the POC distance between the current picture and the backward reference picture. In such a case, the DMVR process is applied to the first motion vector and the second motion vector.
[0117] In some embodiments, a set of motion vectors may be used to obtain a first CU-level reference block in a forward reference picture for a current picture and a second CU-level reference block in a backward reference picture for the current picture.
[0118] In some embodiments, a set of motion vectors may include sub-block based motion vectors for determining the prediction of a current video block. In additional embodiments, sub-block based motion compensation may be performed to generate two predictions for the current video block in two prediction directions, and a bilateral matching cost may be determined as the distortion between the two predictions.
[0119] In some embodiments, a motion offset may be added to each motion vector of each sub-block in a sub-block of a current video block. In one example, after a PU-level reference block or a CU-level reference block is obtained, sub-block motion vectors may be determined. Sub-block based motion compensation may be performed by adding a motion offset to each motion vector of each sub-block in a sub-block of a current video block. For example, a first motion offset may be added to a motion vector in a first prediction direction of a sub-block of a current video block. Additionally, a second motion offset may be added to a motion vector in a second prediction direction different from the first prediction direction of a sub-block of a current video block. The second motion offset may be opposite to the first motion offset. By way of example and not limitation, if the first motion offset is an increment, the second motion offset may be an increment. Alternatively, the motion offset may depend on the steps of the DMVR processing.
[0120] In some embodiments, the motion vectors for obtaining two PU or CU level reference blocks for a current video block can be refined by adding a motion offset to the motion vectors. For example, before the two PU or CU level reference blocks are obtained, a first motion offset can be added to the motion vectors in a first prediction direction for all sub-blocks of the current video block. Additionally, a second motion offset can be added to the motion vectors in a second prediction direction different from the first prediction direction for all sub-blocks of the current video block. Similarly, the second motion offset can be opposite to the first motion offset. Alternatively, the motion offset can depend on the steps of the DMVR process. Additionally, two PU or CU level reference blocks can be obtained based on the refined motion vectors.
[0121] According to further embodiments of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing. In the method, a set of motion vectors for a current video block of the video is obtained. The current video block is encoded and decoded using sub-block-based encoding and decoding tools. A decoder-side motion vector refinement (DMVR) process is applied to the set of motion vectors. Additionally, a bitstream is generated based on the application.
[0122] According to still further embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, a set of motion vectors for a current video block of the video is obtained. The current video block is encoded and decoded using sub-block-based encoding and decoding tools. A decoder-side motion vector refinement (DMVR) process is applied to the set of motion vectors. Additionally, a bitstream is generated based on the application, and the bitstream is stored in a non-transitory computer-readable recording medium.
[0123] Figure 43 A flowchart of a method 4300 for video processing according to some embodiments of the present disclosure is shown. Method 4300 can be implemented during the conversion between a current video block of a video and the bitstream of the video. As Figure 43 shown, method 4300 starts at 4302, where a motion-compensated prediction of the current video block is obtained. The current video block is encoded and decoded using an intra-frame template matching mode or an intra-block copy (IBC) mode.
[0124] At 4304, a sample refinement process is applied to the motion-compensated prediction. In some embodiments, the sample refinement process can be LIC. Alternatively, the sample refinement process can be OBMC. It should be understood that the above examples are described only for purposes of illustration. The scope of the present disclosure is not limited in this regard.
[0125] At 4306, a transformation is performed based on the application. In some embodiments, the transformation may include encoding the current video block into a bitstream. Alternatively or additionally, the transformation may include decoding the current video block from the bitstream.
[0126] In view of the above, the sample refinement process is used for video blocks encoded or decoded using the intra-template matching mode or the IBC mode. Compared with conventional solutions, the proposed method can advantageously improve the encoding / decoding efficiency and the encoding / decoding quality.
[0127] In some embodiments, a cost metric based on mean removal can be used as a criterion for determining the motion vector or the block vector for the current video block. In one example, the current video block may be encoded or decoded using the intra-template matching mode and the local illumination compensation (LIC) mode. In another example, the current video block may be encoded or decoded using the intra-template matching mode and the overlapped block motion compensation (OBMC) mode. In another example, the current video block may be encoded or decoded using the IBC mode and the LIC mode. In yet another example, the current video block may be encoded or decoded using the IBC mode and the OBMC mode. Additionally or alternatively, the current video block may be a screen content video block or a camera-captured content video block.
[0128] In some embodiments, a cost function based on mean removal can be used to measure: the difference between two templates to be compared, the error between two templates to be compared, the cost between two templates to be compared, the sum of absolute differences (SAD) between two templates to be compared, the mean-removed SAD (MRSAD) between two templates to be compared, or the sum of absolute transform differences (SATD) between two templates to be compared, etc.
[0129] In some embodiments, the first template may be determined based on samples adjacent to the current video block. Additionally, the second template may be determined based on samples adjacent to the block pointed to by the candidate motion vector or the candidate block vector for the current video block. Further, the average value between the first template and the second template may be determined. For example, the accumulated difference may be obtained by accumulating all the sample differences between the first template and the second template, and the average value may be obtained based on the result of dividing the accumulated difference by the number of samples in the first template. In this case, the number of samples in the first template is equal to the number of samples in the second template. It should be understood that the above description is only for the purpose of description. The scope of the present disclosure is not limited in this regard.
[0130] In some embodiments, each sample difference in the sample differences between the first template and the second template may be subtracted by an average value to obtain a mean-removed difference, and a first cost metric may be determined based on the mean-removed difference to obtain a cost metric based on mean removal. By way of example and not limitation, the first cost metric may be a difference, an error, SAD, MRSAD, SATD, etc.
[0131] In some embodiments, a sample refinement process may be applied to all samples within the current video block. Alternatively, the sample refinement process may be applied to some samples within the current video block. For example, some samples may include samples at the left boundary of the current video block. Additionally or alternatively, some samples may include samples at the top boundary of the current video block.
[0132] In some embodiments, the same parameters for the sample refinement process may be used for all samples to be refined. Alternatively, different parameters for the sample refinement process may be used for different samples to be refined. In some embodiments, the parameters for the sample refinement process may depend on the position of the sample to be refined. By way of example and not limitation, the parameters may depend on: the relative position between the sample and the left template of the current video block, the relative position between the sample and the upper template of the current video block, and / or the relative position between the sample and the boundary of the current video block.
[0133] In some embodiments, the parameters for the sample refinement process may depend on the cost metric between samples adjacent to the current video block and samples adjacent to another video block different from the current video block. For example, the cost metric may include a difference, an error, SAD, MRSAD, SATD, etc. Additionally, the other video block may be pointed to by a motion vector or a block vector for the current video block. In some embodiments, the cost metric may be determined based on a template-based scheme.
[0134] In some embodiments, the prediction of the current video block may be generated by mixing the prediction of the current video block generated based on the linear model top (LM-T) mode and the prediction of the current video block generated based on the linear model left (LM-L) mode.
[0135] In some additional embodiments, the prediction of the current video block for the LM-TL mode may be generated by mixing the prediction of the current video block generated based on the upper neighboring samples of the current video block and the prediction of the current video block generated based on the left neighboring samples of the current video block.
[0136] In some additional embodiments, the prediction for the current video block in the LIC mode may be generated by mixing the LIC prediction of the current video block generated based on the upper neighboring samples of the current video block and the LIC prediction of the current video block generated based on the left neighboring samples of the current video block.
[0137] In some embodiments, the target prediction for the current video block may be generated by mixing a first prediction of the current video block generated based on neighboring samples in a first direction (such as, above or left) and a second prediction of the current video block generated based on neighboring samples in a second direction (such as, left or above). It should be understood that the specific directions described herein are intended to be exemplary and not to limit the scope of the present disclosure.
[0138] In some embodiments, the direction that has a major impact on the target prediction may be determined. In other words, information about which prediction contributes more to the target prediction may be determined. For example, whether the target prediction mainly comes from the first prediction or the second prediction may be determined. In one example, if the cost metric of the neighboring samples in the first direction is less than the cost metric of the neighboring samples in the second direction, the target prediction may mainly come from the first prediction. Alternatively, if the cost metric of the neighboring samples in the first direction is greater than the cost metric of the neighboring samples in the second direction, the target prediction may mainly come from the first prediction. By way of example and not limitation, the cost metric may include difference, error, SAD, MRSAD, SATD, etc.
[0139] In some embodiments, the mixing weights of the first prediction and the second prediction may be uniform. For example, a first weight may be assigned to all samples of the first prediction, and a second weight may be assigned to all samples of the second prediction. In one example, the first weight may be equal to the second weight. Alternatively, the first weight may be different from the second weight. For example, the first weight and the second weight may be determined based on information about which prediction contributes more to the target prediction.
[0140] By way of example and not limitation, if the first prediction contributes more to the target prediction, the first weight may be greater than the second weight. If the second prediction contributes more to the target prediction, the second weight may be greater than the first weight.
[0141] In some additional embodiments, the mixing weights of the first prediction and the second prediction may be sample-based. For example, different mixing weights may be assigned to different samples. The mixing weight of a sample may depend on the position of the sample. By way of example, if the position of a first sample is closer to the direction determined to be the main contributor to the target prediction than the position of a second sample, the mixing weight of the first sample may be greater than the mixing weight of the second sample.
[0142] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a device for video processing. In the method, a motion-compensated prediction of a current video block of the video is obtained. The current video block is encoded and decoded using an intra-template matching mode or an intra-block copy (IBC) mode. A sample refinement process is applied to the motion-compensated prediction. Further, a bitstream is generated based on the application.
[0143] According to still other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In the method, a motion-compensated prediction of a current video block of the video is obtained. The current video block is encoded and decoded using an intra-template matching mode or an intra-block copy (IBC) mode. A sample refinement process is applied to the motion-compensated prediction. Further, a bitstream is generated based on the application, and the bitstream is stored in a non-transitory computer-readable recording medium.
[0144] Figure 44 A flowchart of a method 4400 for video processing according to some embodiments of the present disclosure is shown. The method 4400 may be implemented during the conversion between a current video block of a video and a bitstream of the video. As Figure 44 shown, the method 4400 starts at 4402, where a plurality of Merge candidates for the current video block are obtained. By way of example and not limitation, the plurality of Merge candidates may be included in one of the following: a regular Merge list, a Merge list based on a Merge mode with motion vector difference (MMVD), a Merge list based on template matching (TM), a Merge list based on bilateral matching (BM), a Merge list based on DMVR, an affine DMVR Merge list, a combined inter and intra prediction (CIIP) Merge list, a CIIP TM Merge list, a geometric partition mode (GPM) Merge list, a GPM TM Merge list, an SbTMVP Merge list, or an SbTMVP TM Merge list. It should be understood that the above examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0145] At 4404, a duplicate removal check is applied to the plurality of Merge candidates by checking first coding information and motion information of the plurality of Merge candidates. The first coding information is different from the motion information. In some embodiments, the first coding information may include a bi-directional prediction (BCW) index having coding unit-level weights. Additionally or alternatively, the first coding information may include a LIC flag. It should be understood that the possible implementations of the first coding information described herein are merely illustrative and should not be construed as limiting the present disclosure in any way.
[0146] At 4406, a transformation is performed based on the application. In some embodiments, the transformation may include encoding the current video block into a bitstream. Alternatively or additionally, the transformation may include decoding the current video block from the bitstream.
[0147] In view of the above, a deduplication check is performed by considering motion information and additional codec information different from the motion information. Compared with conventional solutions, the proposed method can advantageously improve codec efficiency and codec quality.
[0148] In some embodiments, the plurality of Merge candidates may include a first Merge candidate and a second Merge candidate. If the motion vector of the first Merge candidate is the same as the motion vector of the second Merge candidate and the first codec information of the first Merge candidate is the same as the codec information of the second Merge candidate, then the first Merge candidate may be determined to be different from the second Merge candidate during the deduplication check.
[0149] In some additional embodiments, if the motion vector of the first Merge candidate is similar to the motion vector of the second Merge candidate and the first codec information of the first Merge candidate is not similar to the codec information of the second Merge candidate, then the first Merge candidate may be determined to be dissimilar to the second Merge candidate during the deduplication check.
[0150] For example, if the similarity metric between the motion vectors of the first Merge candidate and the second Merge candidate is less than a threshold, then the motion vector of the first Merge candidate may be considered similar to the motion vector of the second Merge candidate. Similarly, if the similarity metric between the first codec information of the first Merge candidate and the second Merge candidate is greater than a threshold, then the first codec information of the first Merge candidate may be considered dissimilar to the codec information of the second Merge candidate. By way of example and not limitation, the similarity metric may be a difference, SAD, etc.
[0151] According to another embodiment of the present disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. In the method, a plurality of Merge candidates for a current video block of the video are obtained. A deduplication check is applied to the plurality of Merge candidates by examining the first codec information and motion information of the plurality of Merge candidates. The first codec information is different from the motion information. In addition, a bitstream is generated based on the application.
[0152] According to still some other embodiments of the present disclosure, a method for storing a bitstream of a video is provided. In this method, a plurality of Merge candidates for a current video block of the video are obtained. A duplicate removal check is applied to the plurality of Merge candidates by checking first codec information and motion information of the plurality of Merge candidates. The first codec information is different from the motion information. Further, a bitstream is generated based on the application and stored in a non-transitory computer-readable recording medium.
[0153] The embodiments of the present disclosure may be described according to the following items, and the features may be combined in any reasonable manner.
[0154] Item 1. A method for video processing, comprising: obtaining a set of motion vectors for a current video block of a video for conversion between the current video block of the video and a bitstream of the video, where the current video block is coded and decoded using a sub-block-based coding tool; applying decoder-side motion vector refinement (DMVR) processing to the set of motion vectors; and performing the conversion based on the application.
[0155] Item 2. The method according to Item 1, where the sub-block-based coding tool includes a sub-block-based temporal motion vector prediction (SbTMVP) mode or an affine mode.
[0156] Item 3. The method according to any one of Items 1 to 2, where the DMVR processing may include one of the following: prediction unit (PU)-level DMVR processing, which outputs a motion offset based on a PU; coding unit (CU)-level DMVR processing, which outputs a motion offset based on a CU; sub-PU-level DMVR processing, which outputs a motion offset based on a sub-PU; sub-CU-level DMVR processing, which outputs a motion offset based on a sub-CU; multi-pass DMVR processing, including PU-level DMVR processing and sub-PU-level DMVR processing; or multi-pass DMVR processing, including CU-level DMVR processing and sub-CU-level DMVR processing.
[0157] Item 4. The method according to any one of Items 1 to 3, where the set of motion vectors is bi-directionally coded and decoded.
[0158] Item 5. The method according to any one of Items 1 to 4, where the set of motion vectors satisfies the DMVR condition.
[0159] Item 6. The method according to any one of Items 1 to 5, wherein a first motion vector in the set of motion vectors points to a forward reference picture for a current picture including the current video block, a second motion vector in the set of motion vectors points to a backward reference picture for the current picture, and a picture order count (POC) distance between the forward reference picture and the current picture is the same as a POC distance between the current picture and the backward reference picture.
[0160] Item 7. The method according to any one of Items 1 to 6, wherein the set of motion vectors is used to obtain a first CU-level reference block in a forward reference picture for a current picture including the current video block and a second CU-level reference block in a backward reference picture for the current picture.
[0161] Item 8. The method according to any one of Items 1 to 7, wherein the set of motion vectors includes sub-block-based motion vectors for determining prediction of the current video block.
[0162] Item 9. The method according to any one of Items 1 to 8, wherein sub-block-based motion compensation is performed to generate two predictions for the current video block in two prediction directions, and a bilateral matching cost is determined as a distortion between the two predictions.
[0163] Item 10. The method according to any one of Items 1 to 9, wherein a motion offset is added to each motion vector of each sub-block in the sub-blocks of the current video block.
[0164] Item 11. The method according to Item 10, wherein after a PU-level reference block or a CU-level reference block is obtained, sub-block motion vectors are determined, and sub-block-based motion compensation is performed by adding the motion offset to each motion vector of each sub-block in the sub-blocks of the current video block.
[0165] Item 12. The method according to any one of Items 10 to 11, wherein a first motion offset is added to motion vectors of all sub-blocks of the current video block in a first prediction direction.
[0166] Item 13. The method according to Item 12, wherein a second motion offset is added to motion vectors of all sub-blocks of the current video block in a second prediction direction different from the first prediction direction, and the second motion offset is opposite to the first motion offset.
[0167] Item 14. The method according to Item 10, wherein the motion offset depends on steps of the DMVR process.
[0168] Item 15. The method according to any one of Items 1 to 9, wherein the motion vectors for obtaining two PU or CU level reference blocks for the current video block are refined by adding a motion offset to the motion vectors.
[0169] Item 16. The method according to Item 15, wherein before the two PU or CU level reference blocks are obtained, a first motion offset is added to the motion vectors in a first prediction direction for all sub - blocks of the current video block, a second motion offset is added to the motion vectors in a second prediction direction different from the first prediction direction for all sub - blocks of the current video block, and the second motion offset is opposite to the first motion offset.
[0170] Item 17. The method according to any one of Items 15 to 16, wherein the two PU or CU level reference blocks are obtained based on the refined motion vectors.
[0171] Item 18. The method according to Item 15, wherein the motion offset depends on the steps of the DMVR process.
[0172] Item 19. A method for video processing, comprising: obtaining a motion - compensated prediction of a current video block of a video for conversion between the current video block of the video and a bit - stream of the video, the current video block being encoded or decoded using an intra - frame template matching mode or an intra - block copy (IBC) mode; applying a sample refinement process to the motion - compensated prediction; and performing the conversion based on the application.
[0173] Item 20. The method according to Item 19, wherein a cost metric based on mean removal is used as a criterion for determining a motion vector or a block vector for the current video block.
[0174] Item 21. The method according to Item 20, wherein the current video block is encoded or decoded using the intra - frame template matching mode and a local illumination compensation (LIC) mode, or the current video block is encoded or decoded using the intra - frame template matching mode and an overlapping - block motion compensation (OBMC) mode, or the current video block is encoded or decoded using the IBC mode and the LIC mode, or the current video block is encoded or decoded using the IBC mode and the OBMC mode.
[0175] Item 22. The method according to any one of Items 19 to 22, wherein the current video block is a screen content video block or a camera - captured content video block.
[0176] Item 23. The method according to any one of Items 20 to 22, wherein a cost function based on mean removal is used to measure one of the following: the difference between two templates to be compared, the error between two templates to be compared, the cost between two templates to be compared, the sum of absolute differences (SAD) between two templates to be compared, the mean-removed SAD (MRSAD) between two templates to be compared, or the sum of absolute transform differences (SATD) between two templates to be compared.
[0177] Item 24. The method according to any one of Items 20 to 23, wherein the first template is determined based on samples adjacent to the current video block, and the second template is determined based on samples adjacent to the block pointed to by the candidate motion vector or candidate block vector for the current video block.
[0178] Item 25. The method according to Item 24, wherein the average value between the first template and the second template is determined.
[0179] Item 26. The method according to Item 25, wherein the accumulated difference is obtained by accumulating all the sample differences between the first template and the second template, and the average value is obtained based on the result of dividing the accumulated difference by the number of samples in the first template, the number of samples in the first template being equal to the number of samples in the second template.
[0180] Item 27. The method according to any one of Items 24 to 25, wherein each sample difference among the sample differences between the first template and the second template is subtracted by the average value to obtain the mean-removed difference, and a first cost metric is determined based on the mean-removed difference to obtain the cost metric based on mean removal.
[0181] Item 28. The method according to Item 27, wherein the first cost metric includes one of difference, error, SAD, MRSAD or SATD.
[0182] Item 29. The method according to any one of Items 19 to 28, wherein the sample refinement process includes LIC or OBMC.
[0183] Item 30. The method according to any one of Items 19 to 29, wherein the sample refinement process is applied to all samples within the current video block.
[0184] Item 31. The method according to any one of Items 19 to 29, wherein the sample refinement process is applied to some samples within the current video block.
[0185] Item 32. The method according to Item 31, wherein the partial samples include at least one of the following: the samples at the left boundary of the current video block, or the samples at the top boundary of the current video block.
[0186] Item 33. The method according to any one of Items 19 to 32, wherein the same parameters for the sample refinement process are used for all samples to be refined.
[0187] Item 34. The method according to any one of Items 19 to 32, wherein different parameters for the sample refinement process are used for different samples to be refined.
[0188] Item 35. The method according to any one of Items 19 to 32, wherein the parameters for the sample refinement process depend on the positions of the samples to be refined.
[0189] Item 36. The method according to Item 35, wherein the parameters depend on at least one of the following: the relative position between the sample and the left template of the current video block, the relative position between the sample and the upper template of the current video block, or the relative position between the sample and the boundary of the current video block.
[0190] Item 37. The method according to any one of Items 19 to 32, wherein the parameters for the sample refinement process depend on the cost metric between the samples adjacent to the current video block and the samples adjacent to another video block different from the current video block.
[0191] Item 38. The method according to Item 37, wherein the cost metric includes one of difference, error, SAD, MRSAD, or SATD.
[0192] Item 39. The method according to any one of Items 37 to 38, wherein the other video block is pointed to by the motion vector or block vector for the current video block.
[0193] Item 40. The method according to any one of Items 37 to 39, wherein the cost metric is determined based on a template-based scheme.
[0194] Item 41. The method according to any one of Items 19 to 40, wherein the prediction of the current video block is generated by mixing the prediction of the current video block generated by the top linear model (LM-T) mode and the prediction of the current video block generated by the left linear model (LM-L) mode.
[0195] Item 42. The method according to any one of Items 19 to 40, wherein the prediction of the current video block for the LM-TL mode is generated by mixing the prediction of the current video block generated based on the upper neighboring samples of the current video block and the prediction of the current video block generated based on the left neighboring samples of the current video block.
[0196] Item 43. The method according to any one of Items 19 to 40, wherein the prediction of the current video block for the LIC mode is generated by mixing the LIC prediction of the current video block generated based on the upper neighboring samples of the current video block and the LIC prediction of the current video block generated based on the left neighboring samples of the current video block.
[0197] Item 44. The method according to any one of Items 19 to 40, wherein the target prediction of the current video block is generated by mixing the first prediction of the current video block generated based on the neighboring samples in the first direction and the second prediction of the current video block generated based on the neighboring samples in the second direction.
[0198] Item 45. The method according to Item 44, wherein the direction that has a major influence on the target prediction is determined.
[0199] Item 46. The method according to Item 44, wherein it is determined whether the target prediction mainly comes from the first prediction or the second prediction.
[0200] Item 47. The method according to Item 46, wherein if the cost metric of the neighboring samples in the first direction is less than the cost metric of the neighboring samples in the second direction, the target prediction mainly comes from the first prediction.
[0201] Item 48. The method according to Item 46, wherein if the cost metric of the neighboring samples in the first direction is greater than the cost metric of the neighboring samples in the second direction, the target prediction mainly comes from the first prediction.
[0202] Item 49. The method according to any one of Items 47 to 48, wherein the cost metric includes one of difference, error, SAD, MRSAD, or SATD.
[0203] Item 50. The method according to Item 44, wherein the mixing weights of the first prediction and the second prediction are uniform.
[0204] Item 51. The method according to Item 44, wherein a first weight is assigned to all samples of the first prediction, and a second weight is assigned to all samples of the second prediction.
[0205] Item 52. The method according to Item 51, wherein the first weight is equal to the second weight.
[0206] Item 53. The method according to Item 51, wherein the first weight is different from the second weight.
[0207] Item 54. The method according to Item 51, wherein the first weight and the second weight are determined based on information about which prediction contributes more to the target prediction.
[0208] Item 55. The method according to Item 54, wherein if the first prediction contributes more to the target prediction, the first weight is greater than the second weight, or if the second prediction contributes more to the target prediction, the second weight is greater than the first weight.
[0209] Item 56. The method according to Item 44, wherein the mixing weights of the first prediction and the second prediction are based on samples.
[0210] Item 57. The method according to Item 56, wherein different mixing weights are assigned to different samples.
[0211] Item 58. The method according to any one of Items 56 to 57, wherein the mixing weight of a sample depends on the position of the sample.
[0212] Item 59. The method according to Item 58, wherein if the position of a first sample is closer to the direction determined to be the main contributor to the target prediction than the position of a second sample, the mixing weight of the first sample is greater than the mixing weight of the second sample.
[0213] Item 60. A method for video processing, comprising: obtaining a plurality of Merge candidates for a current video block of a video for conversion between the current video block of the video and a bitstream of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first codec information and motion information of the plurality of Merge candidates, the first codec information being different from the motion information; and performing the conversion based on the application.
[0214] Item 61. The method according to Item 60, wherein the first codec information includes at least one of the following: a bidirectional prediction (BCW) index having codec unit-level weights, or a LIC flag.
[0215] Item 62. The method according to any one of Items 60 to 61, wherein the plurality of Merge candidates include a first Merge candidate and a second Merge candidate, and if the motion vector of the first Merge candidate is the same as the motion vector of the second Merge candidate and the first codec information of the first Merge candidate is the same as the codec information of the second Merge candidate, then during the duplicate check, the first Merge candidate is determined to be different from the second Merge candidate.
[0216] Item 63. The method according to any one of Items 60 to 61, wherein the plurality of Merge candidates include a first Merge candidate and a second Merge candidate, and if the motion vector of the first Merge candidate is similar to the motion vector of the second Merge candidate and the first codec information of the first Merge candidate is not similar to the codec information of the second Merge candidate, then during the duplicate check, the first Merge candidate is determined to be not similar to the second Merge candidate.
[0217] Item 64. The method according to any one of Items 60 to 63, wherein the plurality of Merge candidates are included in one of the following: a regular Merge list, a Merge list based on a Merge mode with motion vector difference (MMVD), a Merge list based on template matching (TM), a Merge list based on bilateral matching (BM), a Merge list based on DMVR, an affine DMVR Merge list, a combined inter-frame and intra-frame prediction (CIIP) Merge list, a CIIP TM Merge list, a geometric partitioning mode (GPM) Merge list, a GPM TM Merge list, an SbTMVP Merge list, or an SbTMVP TM Merge list.
[0218] Item 65. The method according to any one of Items 1 to 64, wherein the transformation includes encoding the current video block into the bitstream.
[0219] Item 66. The method according to any one of Items 1 to 64, wherein the transformation includes decoding the current video block from the bitstream.
[0220] Item 67. An apparatus for video processing, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 66.
[0221] Item 68. A non-transitory computer-readable storage medium stores instructions that cause a processor to execute the method according to any one of Items 1 to 66.
[0222] Item 69. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a video processing device, where the method includes: obtaining a set of motion vectors for a current video block of the video, the current video block being encoded and decoded using sub-block-based encoding and decoding tools; applying decoder-side motion vector refinement (DMVR) processing to the set of motion vectors; and generating the bitstream based on the application.
[0223] Item 70. A method for storing a bitstream of a video includes: obtaining a set of motion vectors for a current video block of the video, the current video block being encoded and decoded using sub-block-based encoding and decoding tools; applying decoder-side motion vector refinement (DMVR) processing to the set of motion vectors; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
[0224] Item 71. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a video processing device, where the method includes: obtaining a motion-compensated prediction of a current video block of the video, the current video block being encoded and decoded using an intra-frame template matching mode or an intra-block copy (IBC) mode; applying sample refinement processing to the motion-compensated prediction; and generating the bitstream based on the application.
[0225] Item 72. A method for storing a bitstream of a video includes: obtaining a motion-compensated prediction of a current video block of the video, the current video block being encoded and decoded using an intra-frame template matching mode or an intra-block copy (IBC) mode; applying sample refinement processing to the motion-compensated prediction; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
[0226] Item 73. A non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by a video processing device, where the method includes: obtaining a plurality of Merge candidates for a current video block of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first encoding and decoding information and motion information of the plurality of Merge candidates, the first encoding and decoding information being different from the motion information; and generating the bitstream based on the application.
[0227] Article 74. A method for storing a bitstream of a video, comprising: obtaining a plurality of Merge candidates for a current video block of the video; applying a deduplication check to the plurality of Merge candidates by checking first codec information and motion information of the plurality of Merge candidates, the first codec information being different from the motion information; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0228] Figure 45 FIG. shows a block diagram of a computing device 4500 in which various embodiments of the present disclosure may be implemented. The computing device 4500 may be implemented as the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300), or may be included in the source device 110 (or the video encoder 114 or 200) or the destination device 120 (or the video decoder 124 or 300).
[0229] It should be understood that Figure 45 the computing device 4500 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0230] As Figure 45 shown, the computing device 4500 includes a general computing device 4500. The computing device 4500 may include at least one or more processors or processing units 4510, a memory 4520, a storage unit 4530, one or more communication units 4540, one or more input devices 4550, and one or more output devices 4540.
[0231] In some embodiments, the computing device 4500 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistant (PDA), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 4500 may support any type of interface to the user (such as a "wearable" circuitry, etc.).
[0232] The processing unit 4510 can be a physical processor or a virtual processor, and can implement various processes based on the programs stored in the memory 4520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 4500. The processing unit 4510 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0233] The computing device 4500 generally includes various computer storage media. Such media can be any media accessible by the computing device 4500, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4520 can be volatile memory (e.g., registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 4530 can be any removable or non-removable media, and can include machine-readable media, such as a memory, a flash drive, a magnetic disk, or other media that can be used to store information and / or data and can be accessed in the computing device 4500.
[0234] The computing device 4500 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 45 it can provide a disk drive for reading from and / or writing to a removable non-volatile magnetic disk, and an optical disk drive for reading from and / or writing to a removable non-volatile optical disk. In this case, each drive can be connected to a bus (not shown) via one or more data media interfaces.
[0235] The communication unit 4540 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 4500 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Therefore, the computing device 4500 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0236] The input device 4550 can be one or more of various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and so on. The output device 4560 can be one or more of various output devices, such as a display, a speaker, a printer, and so on. With the aid of the communication unit 4540, the computing device 4500 can also communicate with one or more external devices (not shown), such as a storage device and a display device, the computing device 4500 can also communicate with one or more devices that enable a user to interact with the computing device 4500, or if needed, the computing device 4500 can also communicate with any device (such as a network card, a modem, etc.) that enables the computing device 4500 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0237] In some embodiments, some or all components of the computing device 4500 can also be arranged in a cloud computing architecture rather than being integrated in a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses a suitable protocol to provide services via a wide area network (such as the Internet). For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0238] In an embodiment of the present disclosure, the computing device 4500 can be used to implement video encoding / decoding. The memory 4520 can include one or more video codec modules 4525 having one or more program instructions. These modules are accessible and executable by the processing unit 4510 to perform the functions of the various embodiments described herein.
[0239] In an example embodiment of performing video encoding, an input device 4550 may receive video data as an input 4570 to be encoded. The video data may be processed, for example, by a video codec module 4525 to generate an encoded bitstream. The encoded bitstream may be provided as an output 4580 via an output device 4560.
[0240] In an example embodiment of performing video decoding, an input device 4550 may receive an encoded bitstream as an input 4570. The encoded bitstream may be processed, for example, by a video codec module 4525 to generate decoded video data. The decoded video data may be provided as an output 4580 via an output device 4560.
[0241] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Accordingly, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: obtaining, for a conversion between a current video block of a video and a bitstream of the video, a set of motion vectors for the current video block, where the current video block is encoded and decoded using sub-block based coding tools; applying a decoder-side motion vector refinement (DMVR) process to the set of motion vectors; and performing the conversion based on the application.
2. The method according to claim 1, wherein the sub-block based coding tools include a sub-block based temporal motion vector prediction (SbTMVP) mode or an affine mode.
3. The method according to any one of claims 1 to 2, wherein the DMVR process may include one of the following: a prediction unit (PU) level DMVR process, the output of which is based on a PU motion offset, a coding unit (CU) level DMVR process, the output of which is based on a CU motion offset, a sub-PU level DMVR process, the output of which is based on a sub-PU motion offset, a sub-CU level DMVR process, the output of which is based on a sub-CU motion offset, a multi-pass DMVR process including a PU level DMVR process and a sub-PU level DMVR process, or a multi-pass DMVR process including a CU level DMVR process and a sub-CU level DMVR process.
4. The method according to any one of claims 1 to 3, wherein the set of motion vectors is bi-directionally encoded and decoded.
5. The method according to any one of claims 1 to 4, wherein the set of motion vectors satisfies the DMVR condition.
6. The method according to any one of claims 1 to 5, wherein a first motion vector in the set of motion vectors points to a forward reference picture for a current picture including the current video block, a second motion vector in the set of motion vectors points to a backward reference picture for the current picture, and a picture order count (POC) distance between the forward reference picture and the current picture is the same as a POC distance between the current picture and the backward reference picture.
7. The method according to any one of claims 1 to 6, wherein the set of motion vectors is used to obtain a first CU level reference block in a forward reference picture for a current picture including the current video block and a second CU level reference block in a backward reference picture for the current picture.
8. The method according to any one of claims 1 to 7, wherein the set of motion vectors includes sub-block based motion vectors for determining a prediction of the current video block.
9. The method according to any one of claims 1 to 8, wherein sub-block based motion compensation is performed to generate two predictions for the current video block in two prediction directions, and a bilateral matching cost is determined as a distortion between the two predictions.
10. The method according to any one of claims 1 to 9, wherein a motion offset is added to each motion vector of each sub-block of the sub-blocks of the current video block.
11. The method according to claim 10, wherein after the PU-level reference block or the CU-level reference block is obtained, the sub-block motion vector is determined, and motion compensation for the sub-blocks is performed by adding the motion offset to each motion vector of each sub-block in the sub-blocks of the current video block.
12. The method according to any one of claims 10 to 11, wherein a first motion offset is added to the motion vectors in a first prediction direction for all sub-blocks of the current video block.
13. The method according to claim 12, wherein a second motion offset is added to the motion vectors in a second prediction direction different from the first prediction direction for all sub-blocks of the current video block, and the second motion offset is opposite to the first motion offset.
14. The method according to claim 10, wherein the motion offset depends on the steps of the DMVR process.
15. The method according to any one of claims 1 to 9, wherein the motion vectors for obtaining two PU-level or CU-level reference blocks for the current video block are refined by adding a motion offset to the motion vectors.
16. The method according to claim 15, wherein before the two PU-level or CU-level reference blocks are obtained, a first motion offset is added to the motion vectors in a first prediction direction for all sub-blocks of the current video block, a second motion offset is added to the motion vectors in a second prediction direction different from the first prediction direction for all sub-blocks of the current video block, and the second motion offset is opposite to the first motion offset.
17. The method according to any one of claims 15 to 16, wherein the two PU-level or CU-level reference blocks are obtained based on the refined motion vectors.
18. The method according to claim 15, wherein the motion offset depends on the steps of the DMVR process.
19. A method for video processing, comprising: obtaining a motion-compensated prediction of a current video block of a video for conversion between the current video block of the video and a bitstream of the video, the current video block being encoded and decoded using an intra-template matching mode or an intra-block copy (IBC) mode; applying a sample refinement process to the motion-compensated prediction; and performing the conversion based on the application.
20. The method according to claim 19, wherein a cost metric based on mean removal is used as a criterion for determining a motion vector or a block vector for the current video block.
21. The method according to claim 20, wherein the current video block is encoded and decoded using the intra-template matching mode and a local illumination compensation (LIC) mode, or the current video block is encoded and decoded using the intra-template matching mode and an overlapped block motion compensation (OBMC) mode, or the current video block is encoded and decoded using the IBC mode and the LIC mode, or the current video block is encoded and decoded using the IBC mode and the OBMC mode.
22. The method according to any one of claims 19 to 22, wherein the current video block is a screen content video block or a camera-captured content video block.
23. The method according to any one of claims 20 to 22, wherein a cost function based on mean removal is used to measure one of the following: the difference between two templates to be compared, the error between two templates to be compared, the cost between two templates to be compared, the sum of absolute differences (SAD) between two templates to be compared, the mean-removed SAD (MRSAD) between two templates to be compared, or the sum of absolute transform differences (SATD) between two templates to be compared.
24. The method according to any one of claims 20 to 23, wherein the first template is determined based on samples adjacent to the current video block, and the second template is determined based on samples adjacent to the block pointed to by a candidate motion vector or a candidate block vector for the current video block.
25. The method according to claim 24, wherein an average value between the first template and the second template is determined.
26. The method according to claim 25, wherein the accumulated difference is obtained by accumulating all sample differences between the first template and the second template, and the average value is obtained based on the result of dividing the accumulated difference by the number of samples in the first template, the number of samples in the first template being equal to the number of samples in the second template.
27. The method according to any one of claims 24 to 25, wherein each sample difference in the sample differences between the first template and the second template is subtracted by the average value to obtain a mean-removed difference, and a first cost metric is determined based on the mean-removed difference to obtain the cost metric based on mean removal.
28. The method according to claim 27, wherein the first cost metric includes one of difference, error, SAD, MRSAD, or SATD.
29. The method according to any one of claims 19 to 28, wherein the sample refinement process includes LIC or OBMC.
30. The method according to any one of claims 19 to 29, wherein the sample refinement process is applied to all samples within the current video block.
31. The method according to any one of claims 19 to 29, wherein the sample refinement process is applied to partial samples within the current video block.
32. The method according to claim 31, wherein the partial samples include at least one of the following: samples at the left boundary of the current video block, or samples at the top boundary of the current video block.
33. The method according to any one of claims 19 to 32, wherein the same parameters for the sample process are used for all samples to be refined.
34. The method according to any one of claims 19 to 32, wherein different parameters for the sample refinement process are used for different samples to be refined.
35. The method according to any one of claims 19 to 32, wherein the parameters for the sample refinement process depend on the position of the sample to be refined.
36. The method according to claim 35, wherein the parameters depend on at least one of the following: the relative position between the sample and the left template of the current video block, the relative position between the sample and the upper template of the current video block, or the relative position between the sample and the boundary of the current video block.
37. The method according to any one of claims 19 to 32, wherein the parameters for the sample refinement process depend on the cost metric between the samples adjacent to the current video block and the samples adjacent to another video block different from the current video block.
38. The method according to claim 37, wherein the cost metric includes one of difference, error, SAD, MRSAD, or SATD.
39. The method according to any one of claims 37 to 38, wherein the other video block is pointed to by the motion vector or block vector for the current video block.
40. The method according to any one of claims 37 to 39, wherein the cost metric is determined based on a template-based scheme.
41. The method according to any one of claims 19 to 40, wherein the prediction of the current video block is generated by mixing the prediction of the current video block generated based on the top linear model (LM-T) mode and the prediction of the current video block generated based on the left linear model (LM-L) mode.
42. The method according to any one of claims 19 to 40, wherein the prediction of the current video block for the LM-TL mode is generated by mixing the prediction of the current video block generated based on the upper neighboring samples of the current video block and the prediction of the current video block generated based on the left neighboring samples of the current video block.
43. The method according to any one of claims 19 to 40, wherein the prediction of the current video block for the LIC mode is generated by mixing the LIC prediction of the current video block generated based on the upper neighboring samples of the current video block and the LIC prediction of the current video block generated based on the left neighboring samples of the current video block.
44. The method according to any one of claims 19 to 40, wherein the target prediction of the current video block is generated by mixing the first prediction of the current video block generated based on the neighboring samples in the first direction and the second prediction of the current video block generated based on the neighboring samples in the second direction.
45. The method according to claim 44, wherein the direction having a major influence on the target prediction is determined.
46. The method according to claim 44, wherein it is determined whether the target prediction mainly comes from the first prediction or the second prediction.
47. The method according to claim 46, wherein if the cost metric of the neighboring sample in the first direction is less than the cost metric of the neighboring sample in the second direction, the target prediction mainly comes from the first prediction.
48. The method according to claim 46, wherein if the cost metric of the neighboring sample in the first direction is greater than the cost metric of the neighboring sample in the second direction, the target prediction mainly comes from the first prediction.
49. The method according to any one of claims 47 to 48, wherein the cost metric includes one of difference, error, SAD, MRSAD or SATD.
50. The method according to claim 44, wherein the mixing weights of the first prediction and the second prediction are uniform.
51. The method according to claim 44, wherein a first weight is assigned to all samples of the first prediction, and a second weight is assigned to all samples of the second prediction.
52. The method according to claim 51, wherein the first weight is equal to the second weight.
53. The method according to claim 51, wherein the first weight is different from the second weight.
54. The method according to claim 51, wherein the first weight and the second weight are determined based on information about which prediction contributes more to the target prediction.
55. The method according to claim 54, wherein if the first prediction contributes more to the target prediction, the first weight is greater than the second weight, or if the second prediction contributes more to the target prediction, the second weight is greater than the first weight.
56. The method according to claim 44, wherein the mixing weights of the first prediction and the second prediction are sample-based.
57. The method according to claim 56, wherein different mixing weights are assigned to different samples.
58. The method according to any one of claims 56 to 57, wherein the mixing weight of a sample depends on the position of the sample.
59. The method according to claim 58, wherein if the position of a first sample is closer to the direction determined to be the main contributor to the target prediction than the position of a second sample, the mixing weight of the first sample is greater than the mixing weight of the second sample.
60. A method for video processing, comprising: obtaining a plurality of Merge candidates for a current video block of a video for conversion between the current video block of the video and a bitstream of the video; applying a duplicate removal check to the plurality of Merge candidates by examining first codec information and motion information of the plurality of Merge candidates, the first codec information being different from the motion information; and performing the conversion based on the application.
61. The method according to claim 60, wherein the first codec information includes at least one of the following: a bi-directional prediction (BCW) index having codec unit-level weights, or a LIC flag.
62. The method according to any one of claims 60 to 61, wherein the plurality of Merge candidates include a first Merge candidate and a second Merge candidate, and if the motion vector of the first Merge candidate is the same as the motion vector of the second Merge candidate and the first codec information of the first Merge candidate is the same as the codec information of the second Merge candidate, then the first Merge candidate is determined to be different from the second Merge candidate during the deduplication check.
63. The method according to any one of claims 60 to 61, wherein the plurality of Merge candidates include a first Merge candidate and a second Merge candidate, and if the motion vector of the first Merge candidate is similar to the motion vector of the second Merge candidate and the first codec information of the first Merge candidate is not similar to the codec information of the second Merge candidate, then the first Merge candidate is determined to be not similar to the second Merge candidate during the deduplication check.
64. The method according to any one of claims 60 to 63, wherein the plurality of Merge candidates are included in one of the following: a regular Merge list, a Merge list based on a Merge Mode with Motion Vector Difference (MMVD), a Merge list based on Template Matching (TM), a Merge list based on Bilateral Matching (BM), a Merge list based on DMVR, an affine DMVR Merge list, a Combined Inter and Intra Prediction (CIIP) Merge list, a CIIP TM Merge list, a Geometric Partition Mode (GPM) Merge list, a GPM TM Merge list, a SbTMVP Merge list, or a SbTMVP TM Merge list.
65. The method according to any one of claims 1 to 64, wherein the transformation includes encoding the current video block into the bitstream.
66. The method according to any one of claims 1 to 64, wherein the transformation includes decoding the current video block from the bitstream.
67. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 66.
68. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of claims 1 to 66.
69. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by an apparatus for video processing, wherein the method comprises: obtaining a set of motion vectors for a current video block of the video, the current video block being coded and decoded using sub-block-based codec tools; applying a decoder-side motion vector refinement (DMVR) process to the set of motion vectors; and Generate the bitstream based on the application.
70. A method for storing a bitstream of a video, comprising: obtaining a set of motion vectors for a current video block of the video, the current video block being encoded and decoded using sub-block-based encoding and decoding tools; applying a decoder-side motion vector refinement (DMVR) process to the set of motion vectors; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
71. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a device for video processing, wherein the method comprises: obtaining a motion-compensated prediction of a current video block of the video, the current video block being encoded and decoded using an intra-template matching mode or an intra-block copy (IBC) mode; applying a sample refinement process to the motion-compensated prediction; and generating the bitstream based on the application.
72. A method for storing a bitstream of a video, comprising: obtaining a motion-compensated prediction of a current video block of the video, the current video block being encoded and decoded using an intra-template matching mode or an intra-block copy (IBC) mode; applying a sample refinement process to the motion-compensated prediction; and generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.
73. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method executed by a device for video processing, wherein the method comprises: obtaining a plurality of Merge candidates for a current video block of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first encoding and decoding information and motion information of the plurality of Merge candidates, the first encoding and decoding information being different from the motion information; and generating the bitstream based on the application.
74. A method for storing a bitstream of a video, comprising: obtaining a plurality of Merge candidates for a current video block of the video; applying a duplicate removal check to the plurality of Merge candidates by checking first encoding and decoding information and motion information of the plurality of Merge candidates, the first encoding and decoding information being different from the motion information; generating the bitstream based on the application; and storing the bitstream in a non-transitory computer-readable recording medium.