Method and device for video processing and medium
By sorting motion candidates in the video unit in the video processing and determining the target value parameters, the problem of insufficient video encoding and codec performance in the prior art is solved, and more efficient video data processing is achieved.
Patent Information
- Application Number
- CN202380069573.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-27
- Filing Date
- 2023-09-26
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video encoding and decoding technologies have shortcomings in improving codec performance, especially in the fine-grained processing and motion prediction of videos.
A method for video processing is proposed, by sorting multiple motion candidates in the current video unit of the video, determining target value parameters related to cost differences, and thus optimizing the encoding and decoding performance of the video. This method performs determination of the target value at a level below the strip level to achieve a more refined parameter value setting.
By more detailed determination of parameter values in video processing, the video encoding and decoding performance is improved and the processing efficiency of video data is enhanced.
Smart Images

Figure CN120077661A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure generally relate to video processing technologies, and more particularly, to diversity reordering. Background Art
[0002] Nowadays, digital video capabilities are being applied to all aspects of people's lives. For video encoding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, there is an overall expectation to further improve the encoding / decoding performance of video encoding / decoding technologies. Summary of the Invention
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method includes: determining a target value of a parameter from a set of candidate values of the parameter for sorting multiple motion candidates of a current video unit with respect to the conversion between the current video unit of a video and the bitstream of the video, where the parameter is associated with a cost difference related to the multiple motion candidates, and the current video unit is part of a slice of the video; and performing the conversion based on the target value.
[0005] According to the method of the first aspect of the present disclosure, a target value of a parameter (e.g., lambda in diversity reordering) is determined for a video unit that is part of a slice. In other words, the determination of the target value is performed at a level lower than the slice level. Compared with conventional solutions where the value of lambda is determined at the picture level or slice level, the proposed method can advantageously enable the value of the parameter to be determined in a more refined manner, and thus the encoding / decoding performance can be improved.
[0006] In a second aspect, a device for video processing is proposed. The device includes a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0007] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing. The method includes: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of a video, where the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a slice of the video; and generating a bitstream based on the target value.
[0009] In a fifth aspect, a method for storing a bitstream of a video is proposed. The method includes: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of a video, where the parameter is associated with a cost difference related to the plurality of motion candidates; generating a bitstream based on the target value; and storing the bitstream in a non-transitory computer-readable recording medium.
[0010] The present invention content is provided to introduce a selection of concepts further described below in the detailed description in a simplified form. The present invention content is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features, and advantages of the example embodiments of the present disclosure will become more apparent from the following detailed description with reference to the accompanying drawings. In the example embodiments of the present disclosure, the same reference numerals generally refer to the same components.
[0012] Figure 1 A block diagram showing an example video codec system according to some embodiments of the present disclosure is shown;
[0013] Figure 2 A block diagram showing a first example video encoder according to some embodiments of the present disclosure is shown;
[0014] Figure 3 A block diagram showing an example video decoder according to some embodiments of the present disclosure is shown;
[0015] Figure 4 An example of an encoder block diagram is shown;
[0016] Figure 5 67 intra prediction modes are shown;
[0017] Figure 6 Reference sample points for wide-angle intra prediction are shown;
[0018] Figure 7 The problem of discontinuity in the case of a direction exceeding 45° is shown;
[0019] Figure 8 Shows the MMVD search points;
[0020] Figure 9 Shows the illustration for the symmetric MVD pattern;
[0021] Figure 10 Shows the extended CU region used in BDOF;
[0022] Figure 11 Shows the top and left neighboring blocks used in CIIP weight derivation;
[0023] Figure 12 Shows the affine motion model based on control points;
[0024] Figure 13 Shows the affine MVF for each sub-block;
[0025] Figure 14 Shows the position of the inherited affine motion prediction value;
[0026] Figure 15 Shows the control point motion vector inheritance;
[0027] Figure 16 Shows the positioning of the candidate positions for constructing the affine Merge mode;
[0028] Figure 17 Shows the illustration of the motion vector usage of the proposed combination method;
[0029] Figure 18 Shows the sub-block MV VSB and pixel Δv(i,j) (black dashed arrow);
[0030] Figure 19A Shows the spatial neighboring blocks used by ATVMP;
[0031] Figure 19B Shows the derivation of the sub-CU motion field by applying the motion displacement from the spatial neighbors and scaling the motion information from the corresponding co-located sub-CU;
[0032] Figure 20 Shows the local illumination compensation;
[0033] Figure 21 Shows that no subsampling is performed for the short side;
[0034] Figure 22 Shows the decoding-side motion vector refinement;
[0035] Figure 23 Shows the diamond region in the search area;
[0036] Figure 24 The location of the spatial Merge candidate is shown;
[0037] Figure 25 shows candidate pairs considered for redundancy check of spatial merge candidates;
[0038] Figure 26 A diagram showing motion vector scaling for temporal Merge candidates is shown;
[0039] Figure 27 The time domain Merge candidate C is shown 0 and C 1 Candidate location of
[0040] Figure 28 The VVC spatial domain neighboring blocks of the current block are shown;
[0041] Figure 29 shows a diagram of a virtual block in the i-th search round;
[0042] Figure 30 An example of GPM partitions grouped at the same angle is shown;
[0043] Figure 31 Unidirectional prediction MV selection for geometric partitioning mode is shown;
[0044] Figure 32 shows the blending weights w using the geometric partitioning mode 0 Example generation of;
[0045] Figure 33 The spatial neighboring blocks used to derive spatial Merge candidates are shown;
[0046] Figure 34 It shows that template matching is performed on the search area around the initial MV;
[0047] Figure 35 A diagram showing sub-blocks of an OBMC application is shown;
[0048] Figure 36 The SBT position, type and transformation type are shown;
[0049] Figure 37 The neighboring sample points used to calculate the SAD are shown;
[0050] Figure 38 shows neighboring samples used to calculate SAD for sub-CU level motion information;
[0051] Figure 39 The classification process is shown;
[0052] Figure 40 Shows the reordering process in the encoder;
[0053] Figure 41 Shows the reordering process in the decoder;
[0054] Figure 42 Shows the IBC reference region depending on the current CU position;
[0055] Figure 43 Shows an example of symmetry in a screen content picture;
[0056] Figure 44A Shows an illustration of BV adjustment for horizontal flipping;
[0057] Figure 44B Shows an illustration of BV adjustment for vertical flipping;
[0058] Figure 45 Shows the intra-frame template matching search region used;
[0059] Figure 46 Shows a flowchart of a method for video processing according to an embodiment of the present disclosure;
[0060] Figure 47 Shows a block diagram of a computing device in which various embodiments of the present disclosure can be implemented.
[0061] Throughout all the figures, the same or similar reference numerals generally refer to the same or similar elements. Detailed Description
[0062] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the description of these embodiments is for illustrative purposes only and to assist those skilled in the art in understanding and implementing the present disclosure, and does not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways other than those described below.
[0063] In the following description and claims, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present disclosure pertains.
[0064] As used herein, the terms "one embodiment", "embodiment", "exemplary embodiment", etc. indicate that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment must include the specific feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Further, when a specific feature, structure, or characteristic is described in connection with an exemplary embodiment, it is submitted that such feature, structure, or characteristic, whether or not explicitly described, is within the knowledge of those skilled in the art in relation to other embodiments.
[0065] It should be understood that although terms such as "first" and "second" may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the listed terms.
[0066] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the exemplary embodiments. As used herein, the singular forms "a", "an", and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprises", "comprising", "has", "having", "includes", and / or "including" when used herein specify the presence of the stated features, elements, and / or components, etc., but do not preclude the presence or addition of one or more other features, elements, components, and / or combinations thereof. Exemplary Environment
[0067] Figure 1 is a block diagram showing an exemplary video codec system 100 that may utilize the techniques of the present disclosure. As shown, the video codec system 100 may include a source device 110 and a destination device 120. The source device 110 may also be referred to as a video encoding device, and the destination device 120 may also be referred to as a video decoding device. In operation, the source device 110 may be configured to generate encoded video data, and the destination device 120 may be configured to decode the encoded video data generated by the source device 110. The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0068] The video source 112 may include sources such as video capture devices. Examples of video capture devices include, but are not limited to, an interface for receiving video data from a video content provider, a computer graphics system for generating video data, and / or combinations thereof.
[0069] Video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming an encoded representation of the video data. The bitstream may include encoded pictures and associated data. The encoded pictures are the encoded representations of the pictures. The associated data may include sequence parameter sets, picture parameter sets, and other syntax structures. The I / O interface 116 may include a modulator / demodulator and / or a transmitter. The encoded video data may be directly transmitted to the destination device 120 via the I / O interface 116 over the network 130A. The encoded video data may also be stored on the storage medium / server 130B for access by the destination device 120.
[0070] The destination device 120 may include an I / O interface 126, a video decoder 124, and a display device 122. The I / O interface 126 may include a receiver and / or a modulator. The I / O interface 126 may obtain the encoded video data from the source device 110 or the storage medium / server 130B. The video decoder 124 may decode the encoded video data. The display device 122 may display the decoded video data to the user. The display device 122 may be integrated with the destination device 120 or may be external to the destination device 120, which is configured to interface with an external display device.
[0071] The video encoder 114 and the video decoder 124 may operate according to video compression standards such as the High Efficiency Video Coding (HEVC) standard, the Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0072] Figure 2 is a block diagram showing an example of a video encoder 200 according to some embodiments of the present disclosure. The video encoder 200 may be Figure 1 an example of the video encoder 114 in the system 100 shown.
[0073] The video encoder 200 may be configured to implement any or all of the techniques of the present disclosure. In Figure 2 the example, the video encoder 200 includes multiple functional components. The techniques described in the present disclosure may be shared among the various components of the video encoder 200. In some examples, a processor may be configured to execute any or all of the techniques described in the present disclosure.
[0074] In some embodiments, the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transformation unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transformation unit 211, a reconstruction unit 212, a buffer 213, and an entropy encoding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra prediction unit 206.
[0075] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an Intra Block Copy (IBC) unit. The IBC unit may perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0076] Furthermore, although some components (such as the motion estimation unit 204 and the motion compensation unit 205) may be integrated, for explanatory purposes, these components are shown separately in the Figure 2 examples.
[0077] The segmentation unit 201 may segment a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.
[0078] The mode selection unit 203 may select, for example, one of a plurality of codec modes (intra codec or inter codec) based on an error result, and provide the resulting intra-coded block or inter-coded block to the residual generation unit 207 to generate residual block data, and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 may select a Combined Intra and Inter Prediction (CIIP) mode in which the prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 may also select a resolution for the motion vector for the block (e.g., sub-pixel accuracy or integer pixel accuracy).
[0079] To perform inter prediction on a current video block, the motion estimation unit 204 may generate motion information for the current video block by comparing one or more reference frames from the buffer 213 with the current video block. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and the decoded samples of a picture from the buffer 213 other than the picture associated with the current video block.
[0080] The motion estimation unit 204 and the motion compensation unit 205 may perform different operations on a current video block, e.g., depending on whether the current video block is in an I-slice, a P-slice, or a B-slice. As used herein, an "I-slice" may refer to a portion of a picture composed of macroblocks, all of which are based on macroblocks within the same picture. Additionally, as used herein, in some aspects, a "P-slice" and a "B-slice" may refer to a portion of a picture composed of macroblocks that are independent of macroblocks within the same picture.
[0081] In some examples, the motion estimation unit 204 may perform uni-directional prediction on a current video block, and the motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. The motion estimation unit 204 may then generate a reference index and a motion vector, where the reference index indicates the reference picture in list 0 or list 1 that contains the reference video block, and the motion vector indicates the spatial shift between the current video block and the reference video block. The motion estimation unit 204 may output the reference index, the prediction direction indicator, and the motion vector as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.
[0082] Alternatively, in other examples, the motion estimation unit 204 may perform bi-directional prediction on a current video block. The motion estimation unit 204 may search the reference pictures in list 0 to find one reference video block for the current video block, and may also search the reference pictures in list 1 to find another reference video block for the current video block. The motion estimation unit 204 may then generate a plurality of reference indexes and a plurality of motion vectors, where the plurality of reference indexes indicate the plurality of reference pictures in list 0 and list 1 that contain the plurality of reference video blocks, and the plurality of motion vectors indicate the plurality of spatial shifts between the plurality of reference video blocks and the current video block. The motion estimation unit 204 may output the plurality of reference indexes and the plurality of motion vectors of the current video block as the motion information of the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the plurality of reference video blocks indicated by the motion information of the current video block.
[0083] In some examples, the motion estimation unit 204 may output a complete set of motion information for use in the decoding process of a decoder. Alternatively, in some embodiments, the motion estimation unit 204 may signal the motion information of the current video block by referring to the motion information of another video block. For example, the motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0084] In one example, the motion estimation unit 204 may indicate a value in the syntax associated with the current video block, and this value indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0085] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in the syntax associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0086] As discussed above, the video encoder 200 may signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0087] The intra prediction unit 206 may perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 may generate prediction data for the current video block based on the decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and a plurality of syntax elements.
[0088] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the (multiple) predicted video blocks of the current video block from the current video block. The residual data of the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.
[0089] In other examples, such as in the skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.
[0090] The transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.
[0091] After the transform processing unit 208 generates the transform coefficient video block associated with the current video block, the quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0092] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform to the transformed coefficient video block respectively to reconstruct the residual video block from the transformed coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current video block for storage in the buffer 213.
[0093] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation can be performed to reduce the video block effect artifacts in the video block.
[0094] The entropy coding unit 214 can receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 can perform one or more entropy coding operations to generate the entropy-coded data and output a bitstream including the entropy-coded data.
[0095] Figure 3 is a block diagram illustrating an example of a video decoder 300 according to some embodiments of the present disclosure. The video decoder 300 can be Figure 1 an example of the video decoder 124 in the system 100 shown.
[0096] The video decoder 300 can be configured to perform any or all of the techniques of the present disclosure. In Figure 3 the example, the video decoder 300 includes a plurality of functional components. The techniques described in the present disclosure can be shared among the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in the present disclosure.
[0097] In Figure 3 the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally opposite to the encoding process described with respect to the video encoder 200.
[0098] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream can include entropy-encoded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information from the entropy-decoded video data, the motion information including motion vectors, motion vector precision, reference picture list indices, and other motion information. The motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge mode. AMVP is used, including derivation based on several most likely candidates from neighboring PB data and reference pictures. Motion information generally includes horizontal and vertical motion vector shift values, one or two reference picture indices, and, in the case of the prediction region in a B slice, also an indication of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to motion information derived from spatially neighboring blocks or temporally neighboring blocks.
[0099] The motion compensation unit 302 can generate motion-compensated blocks, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax element.
[0100] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of video blocks to calculate the interpolated values for sub-integer pixels of a reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 according to the received syntax information, and the motion compensation unit 302 can use the interpolation filter to generate a prediction block.
[0101] The motion compensation unit 302 can use at least part of the syntax information to determine the size of the blocks for encoding the (multiple) frames and / or (multiple) slices of the encoded video sequence, the partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, the mode indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-frame encoded block, and other information for decoding the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding / decoding, signal prediction, and residual signal reconstruction. A slice can be the entire picture or can also be a region of the picture.
[0102] The intra prediction unit 303 can use, for example, the intra prediction mode received in the bitstream to form a prediction block from spatially neighboring blocks. The inverse quantization unit 304 inverse quantizes (i.e., dequantizes) the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 305 applies an inverse transform.
[0103] The reconstruction unit 306 may obtain the decoded block, for example, by adding a residual block to a corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter may also be applied to filter the decoded block to remove block effect artifacts. The decoded video block is then stored in the cache 307, which provides reference blocks for subsequent motion compensation / intra prediction, and the cache 307 also produces the decoded video for presentation on a display device.
[0104] Some example embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for ease of understanding and does not limit the embodiments disclosed in the section to that section. Additionally, although some embodiments are described with reference to multi-functional video coding or other specific video codecs, the disclosed techniques are also applicable to other video coding techniques. Further, although some embodiments describe the video encoding steps in detail, it should be understood that the corresponding decoding steps for decoding will be implemented by a decoder. Additionally, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compression format to another or at different compression bitrates. 1. Brief Overview The present disclosure relates to video coding and decoding techniques. Specifically, the present disclosure relates to adaptive reordering for Merge candidates and diversity reordering of other coding tools in image / video coding and decoding. The present disclosure can be applied to existing video coding and decoding standards such as HEVC or multi-functional video coding (VVC). The present disclosure is also applicable to future video coding and decoding standards or video codecs. 2. Introduction Video coding standards have mainly evolved from the well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, and ISO / IEC developed MPEG-1 and MPEG-4 Visual. These two organizations jointly developed the H.264 / MPEG-2 video, H.264 / MMPEG-4 Advanced Video Coding (AVC), and H.264 / HEVC standards. Since H.262, video coding standards have been based on a hybrid video coding structure that utilizes temporal prediction plus transform coding. To explore future video coding technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into a reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Exploration Team (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was created to work on the VVC standard, with the goal of reducing the bitrate by 50% compared to HEVC. 2.1. Coding and Decoding Processes of Typical Video Codecs Figure 4 An example of the encoder block diagram of VVC is shown, which includes three loop filter blocks: the Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Different from the DF that uses predefined filters, SAO and ALF utilize the original samples of the current picture to reduce the mean square error between the original samples and the reconstructed samples by adding offsets and applying a Finite Impulse Response (FIR) filter, respectively, where the encoded side information signals the offsets and filter coefficients. ALF is located in the last processing stage of each picture and can be regarded as a tool to attempt to capture and repair the artifacts generated in the previous stage. 2.2. Intra Mode Coding and Decoding with 67 Intra Prediction Modes To capture any edge direction presented in natural videos, the number of directional intra modes is extended from 33 used in HEVC to 65, as Figure 5 shown, and the planar mode and DC mode remain unchanged. These denser directional intra prediction modes are applicable to all block sizes and to both luminance intra prediction and chrominance intra prediction. In HEVC, each intra-coded block has a square shape and the length of each of its sides is a power of 2. Therefore, no division operation is required to generate the intra prediction value using the DC mode. In VVC, blocks can have a rectangular shape, which generally requires a division operation for each block. To avoid the division operation for DC prediction, only the longer side is used to calculate the average value of non-square blocks. 2.2.1. Wide-angle Intra Prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index further depends on the block shape. The traditional angular intra prediction directions are defined as from 45 degrees to -135 degrees in the clockwise direction. In VVC, several traditional angular intra prediction modes are adaptively replaced with wide-angle intra prediction modes for non-square blocks. The original mode index is used to signal the replaced mode, and the original mode index is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode encoding and decoding methods remain unchanged. To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as Figure 6 shown. The number of replaced modes in the wide-angle direction mode depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2-1. Table 2-1 - Intra Prediction Modes Replaced by Wide-angle Modes Aspect ratio Replaced intra prediction mode W / H == 16 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 W / H == 8 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13 W / H == 4 Modes 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 W / H == 2 Modes 2, 3, 4, 5, 6, 7, 8, 9 W / H == 1 None W / H == 1 / 2 Modes 59, 60, 61, 62, 63, 64, 65, 66 W / H == 1 / 4 Modes 57, 58, 59, 60, 61, 62, 63, 64, 65, 66 W / H == 1 / 8 Modes 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66 W / H == 1 / 16 Modes 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66 As Figure 7 shown, in the case of wide-angle intra prediction, two vertically adjacent prediction samples can use two non-adjacent reference samples. Therefore, a low-pass reference sample filter and side smoothing are applied to wide-angle prediction to reduce the negative impact brought by the increased gap Δp α If the wide-angle mode represents a non-fractional offset. There are 8 modes in the wide-angle mode that meet this condition, which are [-14, -12, -10, -6, 72, 76, 78, 80]. When predicting a block through these modes, the samples in the reference buffer are directly copied without applying any interpolation. Through this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of non-fractional modes in traditional prediction modes and wide-angle modes. In VVC, 4:2:2 and 4:4:4 chroma formats as well as 4:2:0 chroma format are supported. The chroma derivation mode (DM) derivation table for the 4:2:2 chroma format is initially transplanted from HEVC, and the number of entries is extended from 35 to 67 to align with the extension of intra prediction modes. Since the HEVC specification does not support prediction angles below -135 degrees and above 45 degrees, the luminance intra prediction modes in the range from 2 to 5 are mapped to 2. Therefore, the chroma DM derivation table for the 4:2:2 chroma format is updated by replacing some values of the entries in the mapping table to more precisely target the prediction angles for chroma blocks. 2.3. Inter Prediction For each inter prediction CU, the motion parameters include the motion vector, the reference picture index and reference picture list use index, and additional information required for the new decoding features of VVC that will be used for inter prediction sample generation. The motion parameters may be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU is associated with a PU and has no significant residual coefficients, no coded motion vector difference (delta) or reference picture index. A Merge mode is specified, whereby the motion parameters of the current CU are obtained from neighboring CUs, including spatial candidates and temporal candidates, and additional items introduced in VVC. The Merge mode can be applied to any inter prediction CU, not just to skip mode. An alternative to the Merge mode is the explicit signaling of the motion parameters, where the motion vector, the corresponding reference picture index and reference picture list use flag for each reference picture list, and other required information are signaled explicitly for each CU. 2.4. Intra Block Copy (IBC) Intra Block Copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency of screen content material. Since the IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to the reference block that has been reconstructed within the current picture. The luminance block vectors of the CUs coded by IBC have integer precision. The chrominance block vectors are also rounded to integer precision. When used in combination with AMVR, the IBC mode can switch between 1-pixel and 4-pixel motion vector precision. The CUs coded by IBC are regarded as a third prediction mode in addition to the intra or inter prediction modes. The IBC mode is applicable to CUs with a width and height both less than or equal to 64 luminance samples. On the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD checks on blocks with a width or height not greater than 16 luminance samples. For non-Merge modes, block vector search is first performed using hash-based search. If the hash search does not return a valid candidate, a block-matching based local search will be performed. In the hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key calculation for each position in the current picture is based on 4×4 sub-blocks. For a current block with a larger size, the hash key is determined to match the hash key of the reference block when the hash keys of all 4×4 sub-blocks match the hash keys in the corresponding reference positions. If the hash keys of multiple reference blocks are found to match the hash key of the current block, the block vector cost for each matching reference is calculated, and the one with the minimum cost is selected. In block matching search, the search range is set to cover both the previous CTU and the current CTU. At the CU level, the IBC mode is signaled using flags, which can be signaled as the IBC AMVP mode or the IBC skip / Merge mode as follows: - IBC skip / Merge mode: The Merge candidate index is used to indicate which block vector from the list of neighboring candidate IBC decoded blocks is used to predict the current block. The Merge list includes spatial candidates, HMVP candidates, and paired candidates. - IBC AMVP mode: The block vector difference is coded in the same way as the motion vector difference. The block vector prediction method uses two candidates as prediction values, one from the left neighbor and one from the upper neighbor (if IBC decoded). When either neighbor is not available, the default block vector is used as the prediction value. A flag is signaled to indicate the block vector prediction value index. 2.5. Merge Mode with MVD (MMVD) In addition to the Merge mode, in the case where implicitly derived motion information is directly used for the prediction sample generation of the current CU, the Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag to specify whether the MMVD mode is used for the CU. In MMVD, after a Merge candidate is selected, it is further refined by the signaled MVD information. The further information includes the Merge candidate flag, an index for specifying the motion magnitude, and an index for indicating the motion direction. In the MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. The MMVD candidate flag is signaled to specify which one of the first Merge candidate and the second Merge candidate is used. The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. As Figure 8 shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2-2. Table 2-2 - Relationship between Distance Index and Predefined Offset Distance index 0 1 2 3 4 5 6 7 Offset (in luminance samples) 1 / 4 1 / 2 1 2 4 8 16 32 The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions, as shown in Table 2-3. It should be noted that the meaning of the MVD symbol can vary according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where two lists point to the same side of the current picture (i.e., both of the two referenced POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbols in Table 2-3 specify the signs of the MV offsets added to the starting MV. When the starting MV is a bidirectional prediction MV with two MVs pointing to different sides of the current picture (i.e., one referenced POC is greater than the POC of the current picture and the other referenced POC is less than the POC of the current picture) and the POC difference in List 0 is greater than the POC difference in List 1, then the symbols in Table 2-3 specify the signs of the MV offsets added to the List 0 MV component of the starting MV, and the signs for the List 1 MV have opposite values. Otherwise, if the POC difference in List 1 is greater than the POC difference in List 0, then the symbols in Table 2-3 specify the signs of the MV offsets added to the List 1 MV component of the starting MV, and the signs for the List 0 MV have opposite values. The MVD is scaled according to the POC differences in each direction. If the POC differences in the two lists are the same, no scaling is required. Otherwise, if the POC difference in List 0 is greater than the POC difference in List 1, then as Figure 26 described, the MVD of List 1 is scaled by defining the POC difference of L0 as td and the POC difference of L1 as tb. If the POC difference of L1 is greater than the POC difference of L0, the MVD of List 0 is scaled in the same way. If the starting MV is unidirectionally predicted, the MVD is added to the available MV. Table 2-3 - Signs of MV Offsets Specified by the Direction Index Direction index 00 01 10 11 x-axis + - N / A N / A y-axis N / A N / A + - 2.6. Symmetric MVD Coding and Decoding In VVC, in addition to the regular unidirectional prediction mode MVD signaling and bidirectional prediction mode MVD signaling, the symmetric MVD mode is applied for bidirectional prediction MVD signaling. In the symmetric MVD mode, the motion information including the reference picture indices of both List 0 and List 1 and the MVD of List 1 is not signaled but is derived. The decoding process of the symmetric MVD mode is as follows: 1. At the slice level, the variables BiDirPredFlag, RefIdxSymL0, and RefIdxSymL1 are derived as follows: – If mvd_l1_zero_flag is 1, then BiDirPredFlag is set to be equal to 0. – Otherwise, if the nearest reference picture in List-0 and the nearest reference picture in List-1 form a forward and backward reference picture pair or a backward and forward reference picture pair, then BiDirPredFlag is set to 1, and both the List-0 reference picture and the List-1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. 2. At the CU level, if the CU is bi-directionally predicted and decoded and BiDirPredFlag is equal to 1, then the symmetry mode flag indicating whether the symmetry mode is used is explicitly signaled. When the symmetry mode flag is true, only mvp_l0_flag, mvp_l1_flag, and MVD0 are explicitly signaled. The reference indices of List 0 and List 1 are respectively set to be equal to the reference picture pair. MVD1 is set to be equal to (-MVD0). The final motion vector is as shown in the following formula. Figure 9 A diagram showing the symmetric MVD mode is presented. In the encoder, symmetric MVD motion estimation starts from the initial MV evaluation. A set of initial MV candidates includes the MVs obtained from the uni-directional prediction search, the MVs obtained from the bi-directional prediction search, and the MVs from the AMVP list. The one with the lowest distortion rate cost is selected as the initial MV for the symmetric MVD motion search. 2.7. Bidirectional Optical Flow (BDOF) The bidirectional optical flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, is included in JEM. Compared with the JEM version, the BDOF in VVC is a simpler version and requires much less computation, especially in terms of the number of multiplications and the size of the multipliers. BDOF is used to refine the bi-directional prediction signal of a CU at the 4x4 sub-block level. BDOF is applied to a CU if the CU meets all of the following conditions: – The CU is decoded using the "true" bi-directional prediction mode, i.e., one of the two reference pictures is before the current picture in the display order, and the other of the two reference pictures is after the current picture in the display order. – The distances from the two reference pictures to the current picture (i.e., the POC differences) are the same. – Both of the two reference pictures are short-term reference pictures. – The CU is not decoded using the affine mode or the SbTMVP Merge mode. – The CU has more than 64 luma samples. – Both the CU height and the CU width are greater than or equal to 8 luma samples. – The BCW weight index indicates equal weights. – WP is not enabled for the current CU. – The CIIP mode is not used for the current CU. BDOF is only applied to the luminance component. As the name implies, the BDOF mode is based on the optical flow concept, which assumes that the motion of an object is smooth. For each 4x4 sub-block, the motion refinement (v x , v y ) is calculated by minimizing the difference between the L0 predicted samples and the L1 predicted samples. Then the motion refinement is used to adjust the bidirectional predicted sample values in the 4x4 sub-block. The following steps are applied during the BDOF process. First, by directly calculating the differences between two neighboring samples, the horizontal and vertical gradients of the two prediction signals, and k = 0, 1, are calculated, i.e., where I (k) (i, j) is the sample value at the coordinates (i, j) of the prediction signal in the list k, k = 0, 1, and shift1 is calculated based on the luminance bit depth bitDepth as shift1 = max(6, bitDepth - 6). Then, the autocorrelations and cross-correlations of the gradients S 1 , S 2 , S 3 , S 5 and S 6 are calculated as where where, Ω is a 6x6 window around the 4x4 sub-block, and n a and n b are respectively set to be equal to min(1, bitDepth - 11) and min(4, bitDepth - 8). Then using the cross-correlation terms and autocorrelation terms, the motion refinement (v x , v y ) is derived using the following method: where th′ BIO = 2 max(5,BD-7) , is the floor function, and Based on motion refinement and gradients, the following adjustments are calculated for each sample point in a 4x4 sub-block: Finally, the BD - OF samples of the CU are calculated by adjusting the bi - directional prediction samples in the manner shown below. These values are selected such that the multipliers in the BD - OF process do not exceed 15 bits, and the maximum bit - width of the intermediate parameters in the BD - OF process remains within 32 bits. To derive the gradient values, some prediction samples I (k) (i,j) in list k (k = 0,1) outside the current CU boundary need to be generated. As Figure 10 shown, BD - OF in VVC uses an extended row / column around the boundary of the CU. To control the computational complexity of generating prediction samples outside the boundary, the prediction samples (white positions) in the extended region are generated by directly taking the reference samples at nearby integer positions (using the floor() operation on the coordinates) without using interpolation, and the regular 8 - tap motion - compensated interpolation filter is used to generate the prediction samples (gray positions) within the CU. These extended sample values are only used for gradient calculation. For the remaining steps in the BD - OF process, if any sample values and gradient values outside the CU boundary are needed, these sample values and gradient values are filled (i.e., repeated) from their nearest neighbors. When the width and / or height of the CU is greater than 16 luma samples, the CU is divided into sub - blocks with width and / or height equal to 16 luma samples, and the sub - block boundaries are considered as the CU boundaries in the BD - OF process. The maximum unit size of the BD - OF process is limited to 16x16. For each sub - block, the BD - OF process can be skipped. When the SAD between the initial L0 prediction sample and the L1 prediction sample is less than the threshold, the BD - OF process is not applied to the sub - block. The threshold is set to be equal to (8*W*(H>>1), where W represents the sub - block width and H represents the sub - block height. To avoid the additional complexity of SAD calculation, the SAD calculated in the DVMR process between the initial L0 prediction sample and the L1 prediction sample is reused here. If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, then bi - directional optical flow is disabled. Similarly, if WP is enabled for the current block, i.e., for either of the two reference pictures, luma_weight_lx_flag is 1, then BD - OF is also disabled. When the CU is encoded / decoded in symmetric MVD mode or CIIP mode, BD - OF is also disabled. 2.8. Combined Inter - frame and Intra - frame Prediction (CIIP) In VVC, when a CU is encoded or decoded in the Merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), and if both the CU width and the CU height are less than 128 luma samples, an additional flag is signaled to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU. As the name implies, CIIP prediction combines the inter prediction signal with the intra prediction signal. The inter prediction signal P in the CIIP mode inter is derived using the same inter prediction process applied to the regular Merge mode; and the intra prediction signal P intra is derived in the regular intra prediction process with the planar mode. Then, the intra prediction signal and the inter prediction signal are combined using weighted averaging, where the weight values depend on the coding modes of the top neighboring block and the left neighboring block (depicted in Figure 11 ) and are calculated as follows: – If the top neighbor is available and is intra-coded, set isIntraTop to 1, otherwise set isIntraTop to 0; – If the left neighbor is available and is intra-coded, set isIntraLeft to 1, otherwise set isIntraLeft to 0; – If (isIntraLeft + isIntraTop) equals 2, set wt to 3; – Otherwise, if (isIntraLeft + isIntraTop) equals 1, set wt to 2; – Otherwise, set wt to 1. The CIIP prediction is formed as follows. P CIIP = ((4 - wt) * P inter + wt * P intra + 2) >> 2 (2 - 8) 2.9. Affine Motion Compensation Prediction In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are various motions such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensation prediction is applied. As Figure 12 shown, the affine motion field of a block is described by the motion information of two control points (4 parameters shown in Sub-picture 1200) or three control point motion vectors (6 parameters shown in Sub-picture 1202). For the 4-parameter affine motion model, the motion vector at the sampling position (x, y) in the block is derived as For a 6-parameter affine motion model, the motion vector at the sampled position (x, y) in the block is derived as follows: where (mv 0x , mv 0y ) is the motion vector of the top-left control point, (mv 1x , mv 1y ) is the motion vector of the top-right control point, and (mv 2x , mv 2y ) is the motion vector of the bottom-left control point. To simplify motion compensation prediction, block-based affine transform prediction is applied. To derive the motion vector for each 4x4 luma sub-block, the motion vector of the center sample of each sub-block (as shown in Figure 13 ) is calculated according to the above equation and rounded to 1 / 16 fractional precision. Then a motion compensation interpolation filter is applied to generate the prediction for each sub-block with the derived motion vector. The sub-block size of the chrominance components is also set to 4x4. The MV of a 4x4 chrominance sub-block is calculated as the average of the MVs of four corresponding 4x4 luma sub-blocks. Similar to translational motion inter prediction, there are also two affine motion inter prediction modes: affine Merge mode and affine AMVP mode. 2.9.1. Affine Merge Prediction The AF_MERGE mode can be applied to CUs with both width and height greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMVP candidates, and an index is signaled to indicate the one to be used for the current CU. The following three types of CPVM candidates are used to form the affine Merge candidate list: – Inherited affine Merge candidates inferred from the CPMV of neighboring CUs; – Constructed affine Merge candidate CPMVPs derived using the translational MVs of neighboring CUs; – Zero MV. In VVC, there are at most two inherited affine candidates, which are derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are in Figure 14Shown in. For the left prediction value, the scanning order is A0 -> A1, and for the upper prediction value, the scanning order is B0 -> B1 -> B2. Only the first inherited candidate from each side is selected. Duplicate removal checking is not performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive the CPMV candidates in the affine Merge list of the current CU. As shown, if the neighboring bottom-left block A is coded in the affine mode, the motion vectors v 2 , v 3 and v 4 of the top-left, top-right, and bottom-left corners of the CU containing block A are obtained. When block A is coded using the 4-parameter affine model, two CPMVs of the current CU are calculated according to v 2 and v 3 . In the case where block A is coded using the 6-parameter affine model, three CPMVs of the current CU are calculated according to v 2 , v 3 and v 4 . The constructed affine candidates mean that the candidates are constructed by combining the neighboring translational motion information of each control point. The motion information of the control points is derived from the specified spatial neighbors and temporal neighbors shown in Figure 16 . CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV 1 , the B2 -> B3 -> A2 block is checked, and the MV of the first available block is used. For CPMV 2 , the B1 > B0 block is checked, and for CPMV 3 , the A1 > A0 block is checked. TMVP is used as CPMV 4 (if available). After the MVs of the four control points are obtained, affine Merge candidates are constructed based on the motion information. The following combinations of the control point MVs are used to construct in sequence: {CPMV 1 , CPMV 2 , CPMV 3}, {CPMV 1 , CPMV 2 , CPMV 4}, {CPMV 1 , CPMV 3 , CPMV 4}, {CPMV 2 , CPMV 3 , CPMV 4}, {CPMV 1 , CPMV 2}, {CPMV 1, CPMV 3}。 Combinations of 3 CPMVs construct 6-parameter affine Merge candidates, and combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of the control point MVs are discarded. After checking the inherited affine Merge candidates and the constructed affine Merge candidates, if the list is still not full, zero MVs are inserted at the end of the list. 2.9.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with width and height both greater than or equal to 16. The affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether it is 4-parameter affine or 6-parameter affine. In this mode, the difference between the CPMV of the current CU and its predicted value CPMVP is signaled in the bitstream. The size of the affine AVMP candidate list is 2, and it is generated by sequentially using the following four types of CPMV candidates: – Inherited affine AMVP candidates inferred from the CPMVs of neighboring CUs; – Constructed affine AMVP candidate CPMVP derived using the translational MVs of neighboring CUs; – Translational MVs from neighboring CUs; – Zero MVs. The checking order of the inherited affine AMVP candidates is the same as that of the inherited affine Merge candidates. The only difference is that for AVMP candidates, only affine CUs with the same reference picture in the current block are considered. When inserting the inherited affine motion prediction values into the candidate list, the deduplication process is not applied. The constructed AMVP candidates are derived from the specified spatial neighbors shown in Figure 16 . The same checking order as in the construction of affine Merge candidates is used. In addition, the reference picture indices of the neighboring blocks are also checked. The first block in the checking order is used, which is inter-coded and has the same reference picture as the current CU. Only when the current CU is coded using the 4-parameter affine mode and both mv 0 and mv 1 are available, then they are added as a candidate in the affine AMVP list. When the current CU is coded using the 6-parameter affine mode and all three CPMVs are available, then they are added as a candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable. If the affine AMVP list candidate is still less than 2 after the inherited affine AMVP candidates and the constructed AMVP candidates are examined, mv 0 , mv 1 and mv 2 will be added in order as translational MVs to predict all control point MVs of the current CU when available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list. 2.9.3. Affine Motion Information Storage In VVC, the CPMV of an affine CU is stored in a separate cache. The stored CPMV is only used for the inherited CPMV in affine Merge mode and the inherited CPMV in affine AMVP mode for the most recently decoded CU. The sub-block MVs derived from the CPMV are used for motion compensation, MV derivation for the Merge / AMVP list of translational MVs, and deblocking. To avoid picture line caches for additional CPMVs, the inheritance of affine motion data from the CU above the CTU is processed differently from the inheritance from normal neighboring CUs. If the candidate CU for affine motion data inheritance is in the CTU upper row, the left-bottom and right-bottom sub-block MVs in the line cache are used for affine MVP derivation instead of the CPMV. In this way, the CPMV is only stored in the local cache. If the candidate CU is 6-parameter affine coded, the affine model is degraded to a 4-parameter model. As Figure 17 shown, along the top boundary of the CTU, the left-bottom and right-bottom sub-block motion vectors of the CU are used for affine inheritance of the CU in the CTU bottom. 2.9.4. Prediction Refinement Using Optical Flow for Affine Mode Compared with pixel-based motion compensation, sub-block-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve a more refined motion compensation granularity, prediction refinement using optical flow (PROF) is used to refine the sub-block-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after sub-block-based affine motion compensation is performed, the luminance prediction samples are refined by adding the differences derived from the optical flow equations. PROF is described as the following four steps: Step 1) Sub-block-based affine motion compensation is performed to generate the sub-block prediction I(i,j). Step 2) Using a 3-tap filter [-1,0,1], the spatial gradients g x (i,j) and g y (i,j) are calculated at each sample position. The gradient calculation is exactly the same as the gradient calculation in BDOF. g x(i,j) = (I(i + 1,j) >> shift1) - (I(i - 1,j) >> shift1) (2-11) g y (i,j) = (I(i,j + 1) >> shift1) - (I(i,j - 1) >> shift1) (2-12) shift1 is used to control the precision of the gradient. For gradient calculation, the sub-block (i.e., 4x4) prediction extends one sample point on each side. To avoid additional memory bandwidth and additional interpolation calculations, those extended sample points on the extended boundaries are copied from the nearest integer pixel positions in the reference picture. Step 3) The luminance prediction refinement is calculated by the following optical flow equation. Δ(i,j)g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) (2-13) where Δv(i,j) is the difference between the sample MV (denoted by v(i,j)) calculated for the sample position (i,j) and the sub-block MV of the sub-block to which the sample (i,j) belongs, as Figure 18 shown. Δv(i,j) is quantized in units of 1 / 32 luminance sample precision. Since the affine model parameters and the sample position relative to the sub-block center do not change from sub-block to sub-block, Δv(i,j) can be calculated for the first sub-block and reused for other sub-blocks in the same CU. Let dx(i,j) and dy(i,j) be the horizontal and vertical offsets from the sample position (i,j) to the center (x SB ,y SB ) of the sub-block, Δv(x,y) can be derived by the following equation, To maintain accuracy, the input of the sub-block (x SB ,y SB ) is calculated as ((W SB –1) / 2,(H SB –1) / 2), where W SB and H SB are the sub-block width and sub-block height respectively. For the 4-parameter affine model, For the 6-parameter affine model, where (v 0x ,v 0y), (v 1x , v 1y ), (v 2x , v 2y ), (v are the top - left, top - right, and bottom - left control point motion vectors, and w and h are the width and height of the CU. Step 4) Finally, the luminance prediction refinement ΔI(i, j) is added to the sub - block prediction I(i, j). The final prediction I’ is generated by the following equation. I'(i, j) = I(i, j)+Δ(i, j) (2 - 18) PROF is not applicable to two cases of affine - coded CUs: 1) all control - point MVs are the same, which indicates that the CU only has translational motion; 2) the affine motion parameters are greater than the specified limit because the sub - block - based affine MC is degraded to CU - based MC to avoid large memory - access bandwidth requirements. Fast coding methods are applied to reduce the coding complexity of affine motion estimation using PROF. PROF is not applied in the affine motion estimation stage in the following two cases: a) if the CU is not a root block and the parent block of the CU does not select the affine mode as its best mode, then PROF is not applied because the probability that the current CU selects the affine mode as the best mode is low; b) if the magnitudes of all four affine parameters (C, D, E, F) are less than a predefined threshold and the current picture is not a low - latency picture, then PROF is not applied because the improvement introduced by PROF for this case is small. In this way, the affine motion estimation using PROF can be accelerated. 2.10. Sub - block - based Temporal Motion Vector Prediction (SbTMVP) VVC supports the sub - block - based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the reference picture to improve the motion vector prediction and the Merge mode of the CUs in the current picture. The same reference picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects: – TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub - CU level; – TMVP extracts the temporal motion vector from the co - located block in the reference picture (the co - located block is the bottom - right block or the center block relative to the current CU), while SbTMVP applies a motion displacement before extracting the temporal motion information from the reference picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks of the current CU. The SbTVMP process is shown in Figure 19A and 19B ). Figure 19A shows the spatial neighboring blocks used by ATVMP, andFigure 19B The derivation of the sub - CU motion field is shown by applying the motion displacement from the spatial neighborhood and scaling the motion information from the corresponding co - located CU. SbTMVP predicts the motion vectors of the sub - CUs within the current CU in two steps. In the first step, Figure 19A the spatial neighbor A1 in is examined. If A1 has a motion vector using the co - located picture as its reference picture, this motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0). Figure 19B In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub - CU level motion information (motion vector and reference index) from the co - located picture as Figure 19B shown. The example in assumes that the motion displacement is set to the motion of block A1. Then, the motion information of the corresponding block (the smallest motion grid covering the central sample) for each sub - CU in the co - located picture is used to derive the motion information of the sub - CU. After the motion information of the co - located sub - CU is identified, the motion information is converted to the motion vector and reference index of the current sub - CU in a similar way to the TMVP process in HEVC, where temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU. In VVC, a combined sub - block - based Merge list containing both SbTMVP candidates and affine Merge candidates is used for signaling the sub - block - based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP prediction value is added as the first entry in the list of sub - block - based Merge candidates, followed by the affine Merge candidates. The size of the sub - block - based Merge list is signaled in the SPS, and the maximum allowed size of the sub - block - based Merge list in VVC is 5. The sub - CU size used in SbTMVP is fixed at 8x8, and like the affine Merge mode, the SbTMVP mode is only applicable to CUs with width and height both greater than or equal to 8. The encoding logic for additional SbTMVP Merge candidates is the same as that for other Merge candidates, i.e., for each CU in a P - slice or B - slice, an additional RD check is performed to decide whether to use the SbTMVP candidate. 2.11. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of the CU and the predicted motion vector) is signaled in units of quarter luminance samples. In VVC, a CU-level Adaptive Motion Vector Resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be encoded and decoded with different precisions. Depending on the current CU's mode (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows: – Normal AMVP mode: quarter luminance sample, half luminance sample, integer luminance sample, or quadruple luminance sample. – Affine AMVP mode: quarter luminance sample, integer luminance sample, or 1 / 16 luminance sample. If the current CU has at least one non-zero MVD component, the CU-level MVD resolution indication is conditionally signaled. If all MVD components (i.e., both the horizontal MVD and the vertical MVD of reference list L0 and reference list L1) are zero, the quarter luminance sample MVD resolution is assumed. For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter luminance sample MVD precision is used for the CU. If the first flag is 0, no further signaling is required and quarter luminance sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half luminance sample or other MVD precision (integer or quadruple luminance sample) is used for normal AMVP CUs. In the case of half luminance sample, the half luminance sample position uses a 6-tap interpolation filter instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer luminance sample or quadruple luminance sample MVD precision is used for normal AMVP CUs. In the case of affine AMVP CUs, the second flag is used to indicate whether integer luminance sample MVD precision or 1 / 16 luminance sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter luminance sample, half luminance sample, integer luminance sample, or quadruple luminance sample), the motion vector prediction value of the CU is rounded to the same precision as the MVD before being added to the MVD. The motion vector prediction value is rounded to zero (i.e., a negative motion vector prediction value is rounded to positive infinity, and a positive motion vector prediction value is rounded to negative infinity). The encoder uses RD checking to determine the motion vector resolution of the current CU. To avoid always performing four CU-level RD checks for each MVD resolution, in VTM11, the RD check for MVD accuracy outside the quarter-luma samples is only conditionally invoked. For the normal AVMP mode, first, the RD cost for the MVD accuracy of the quarter-luma samples and the RD cost for the MV accuracy of the integer-luma samples are calculated. Then, the RD cost for the MVD accuracy of the integer-luma samples is compared with the RD cost for the MVD accuracy of the quarter-luma samples to decide whether it is necessary to further check the RD cost for the MVD accuracy of the four-luma samples. When the RD cost for the MVD accuracy of the quarter-luma samples is much smaller than the RD cost for the MVD accuracy of the integer-luma samples, the RD check for the MVD accuracy of the four-luma samples is skipped. Then, if the RD cost for the MVD accuracy of the integer-luma samples is significantly greater than the best RD cost for the previously tested MVD accuracy, the check for the MVD accuracy of the half-luma samples is skipped. For the affine AMVP mode, if the rate-distortion cost of the affine inter mode is not selected after checking the rate-distortion costs of the affine Merge / skip mode, the Merge / skip mode, the normal AMVP mode with the MVD accuracy of the quarter-luma samples, and the affine AMVP mode with the MVD accuracy of the quarter-luma samples, the MV accuracy of the 1 / 16-luma samples and the affine inter mode with the 1-pixel MV accuracy are not checked. Additionally, in the affine inter mode with the 1 / 16-luma samples and the MVD accuracy of the quarter-luma samples, the affine parameters obtained in the affine inter mode with the MVD accuracy of the quarter-luma samples are used as the starting search point. 2.12. Bi-directional prediction with CU-level weights (BCW) In HEVC, the bi-directional prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-directional prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. P bi-pred ((8 - w)*P 0 + w*P 1 + 4) >> 3 (2 - 19) In weighted-average bi-directional prediction, five weights are allowed, w ∈ {-2, 3, 4, 5, 10}. For each bi-directional prediction CU, the weight w is determined in one of two ways: 1) for non-Merge CUs, the weight index is signaled after the motion vector difference; 2) for Merge CUs, the weight index is deduced from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., the CU width multiplied by the CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights are used (w ∈ {3, 4, 5}). – At the encoder, a fast search algorithm is applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. When combined with AMVR, if the current picture is a low-delay picture, the unequal weights for 1-pixel and 4-pixel motion vector precisions are only conditionally checked. – When combined with affine, the affine ME is performed for unequal weights if and only if the affine mode is selected as the current best mode. – When the two reference pictures in bi-prediction are the same, the unequal weights are only conditionally checked. – When specific conditions are met, the unequal weights are not searched, depending on the POC distance between the current picture and its reference picture, the coding / decoding QP, and the temporal level. The BCW weight index is coded using a context-coded bit followed by a bypass-coded bit. The first context-coded bit indicates whether equal weights are used; and if unequal weights are used, the bypass coding signals additional bits to indicate which unequal weight is used. Weight prediction (WP) is a coding / decoding tool supported by the H.264 / AVC and HEVC standards for efficient coding / decoding of video content in fading scenarios. Support for WP is also added in the VVC standard. WP allows signaling of weight parameters (weights and offsets) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the (multiple) weights and (multiple) offsets for the corresponding (multiple) reference pictures are applied. WP and BCW are designed for different types of video content. To avoid interaction between WP and BCW (which would complicate the VVC decoder design), if a CU uses WP, the BCW weight index is not signaled and w is presumed to be 4 (i.e., equal weights are applied). For a MergeCU, the weight index is presumed from neighboring blocks based on the Merge candidate index. This can be applied to both the normal Merge mode and the inherited affine Merge mode. For the constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index of a CU using the constructed affine Merge mode is simply set to be equal to the BCW index of the first control point MV. In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded / decoded using the CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights. 2.13. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is an encoding and decoding tool used to address the problem of local illumination changes between the current picture and its temporal reference picture. LIC is based on a linear model, where a scaling factor and an offset are applied to reference samples to obtain the predicted samples of the current block. Specifically, LIC can be mathematically modeled by the following equation: P(x,y) = α·P r (x + v x , y + v y ) + β where P(x,y) is the predicted signal of the current block at coordinates (x,y); P r (x + v x , y + v y ) is the reference block pointed to by the motion vector (v x , v y ); α and β are the corresponding scaling factor and offset applied to the reference block. Figure 20 shows the LIC process. In Figure 20 , when LIC is applied to a block, the Least Mean Square Error (LMSE) method is adopted. By minimizing the difference between the neighboring samples of the current block (i.e., the template T in Figure 20 ) and the corresponding reference samples in the temporal reference picture (i.e., T0 or T1 in Figure 20 ), the values of the LIC parameters (i.e., α and β) are derived. Additionally, to reduce the computational complexity, both the template samples and the reference template samples are subsampled (adaptive subsampling) to derive the LIC parameters, i.e., only the shaded samples in Figure 20 are used to derive α and β. To improve the encoding and decoding performance, the short side is not subsampled, as shown in Figure 21 . 2.14. Decoder - side Motion Vector Refinement (DMVR) To improve the accuracy of the MVs in the Merge mode, decoder - side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bidirectional prediction operation, refined MVs are searched around the initial MVs in the reference picture list L0 and the reference picture list L1. The BM method calculates the distortion between two candidate blocks in the reference picture list L0 and list L1. As shown in Figure 22 , the Sum of Absolute Differences (SAD) between two blocks based on each MV candidate (e.g., MV0’ and MV1’) around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bidirectional prediction signal. In VVC, the application of DMVR is restricted and is only applied to the CUs encoded and decoded using the following modes and features: – CU - level Merge mode with bidirectional prediction MVs; – One reference picture is past and the other reference picture is future with respect to the current picture; – The distances from the two reference pictures to the current picture (i.e., POC differences) are the same; – Both reference pictures are short-term reference pictures; – The CU has more than 64 luma samples; – Both the CU height and the CU width are greater than or equal to 8 luma samples; – The BCW weight index indicates equal weights; – WP is not enabled for the current block; – The CIIP mode is not used for the current block. The refined MV derived through the DMVR process is used to generate inter-prediction samples and is also used for temporal motion vector prediction in future picture coding. The original MV is used for the deblocking process and is also used for spatial motion vector prediction in future CU coding. Additional features of DMVR are mentioned in the following sub-articles. 2.14.1. Search Scheme In DVMR, the search points are around the initial MV, and the MV offset follows the MV difference mirroring rule. In other words, any point checked by DMVR represented by a candidate MV pair (MV0, MV1) follows the following two equations: MV0′ = MV0 + MV_offset (2-20) MV1' = MV1 - MV_offset (2-21) where MV_offset represents the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range is two integer luma samples starting from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase. The integer sample offset search uses a 25-point full search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than the threshold, the integer sample phase of DMVR terminates. Otherwise, the SADs of the remaining 24 points are calculated and checked in raster scan order. The point with the minimum SAD is selected as the output of the integer sample offset search phase. To reduce the impact of DMVR refinement uncertainty, it is proposed to support the original MV during the DMVR process. The SAD between the reference blocks pointed to by the initial MV candidates reduces the SAD value by 1 / 4. After the integer sample point search, fractional sample point refinement is performed. To save computational complexity, the fractional sample point refinement is derived using the parametric error surface equation instead of performing an additional search using SAD comparison. The fractional sample point refinement is conditionally invoked based on the output of the integer sample point search phase. When the integer sample point search phase ends at the center with the minimum SAD in the first iteration or the second iteration search, the fractional sample point refinement is further applied. In the sub-pixel offset estimation based on the parametric error surface, the cost at the center position and the costs at four neighboring positions from the center are used to fit a two-dimensional parabolic error surface equation of the following form: E(x,y) = A(x - x min ) 2 + B(y - y min ) 2 + C(2 - 22) where (x min , y min ) corresponds to the fractional position with the minimum cost, and C corresponds to the minimum cost value. By solving the above equation using the cost values of five search points, (x min , y min ) is calculated as follows. x min = (E(-1,0) - E(1,0)) / (2(E(-1,0) + E(1,0) - 2E(0,0))) (2 - 23) y min = (E(0,-1) - E(0,1)) / (2((E(0,-1) + E(0,1) - 2E(0,0)))(2 - 24) The values of x min and y min are automatically limited between -8 and 8 because all cost values are positive and the minimum value is E(0,0). This corresponds to a half-pixel offset with 1 / 16 pixel MV accuracy in VVC. The calculated fraction (x min , y min ) is added to the integer distance refined MV to obtain a sub-pixel accurate refined delta MV. 2.14.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of the MV is 1 / 16 luma samples. An 8-tap interpolation filter is used to interpolate samples at fractional positions. In DMVR, the search points are centered around the initial fractional pixel MV with integer sample offsets, so samples at these fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect is that by using the bilinear filter, within a 2-sample search range, compared to the normal motion compensation process, DVMR does not access more reference samples. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter is applied to generate the final prediction. To avoid accessing more reference samples of the normal MC process, samples will be filled from those available samples that are not needed for the interpolation process based on the original MV but are needed for the interpolation process based on the refined MV. 2.14.3. Maximum DMVR Processing Unit When the width and / or height of the CU is greater than 16 luma samples, it will be further divided into sub-blocks with a width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16. 2.15. Multi-pass Decoder-side Motion Vector Refinement Multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the coded and decoded block. In the second pass, BM is applied to each 16x16 sub-block within the coded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bidirectional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction. 2.15.1. First Pass - Block-based Bilateral Matching MV Refinement In the first pass, the refined MV is derived by applying BM to the coded and decoded block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between two reference blocks in L0 and L1. BM performs a local search to derive the integer sample accuracy intDeltaMV and half-pixel sample accuracy halfDeltaMV. The local search applies a 3×3 square search pattern to loop within the search range [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of distortion between reference blocks. When the bilCost at the center point of the 3×3 search pattern has the minimum cost, the intDeltaMV or halfDeltaMV local search terminates. Otherwise, the current minimum-cost search point becomes the new center point of the 3×3 search pattern, and the search for the minimum cost continues until it reaches the end of the search range. The current fractional sample refinement is further applied to derive the final deltaMV. Then, the refined MV after the first pass is derived as: · MV0_pass1 = MV0 + deltaMV; · MV1_pass1 = MV1 - deltaMV. 2.15.2. Second Pass - Sub-Block Based Bilateral Matching MV Refinement In the second pass, the refined MV is derived by applying BM to 16×16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) obtained in the first pass in the reference picture lists L0 and L1. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1. For each sub-block, BM full search is performed to derive the integer sample accuracy intDeltaMV. The full search has a search range of [–sHor, sHor] in the horizontal direction and [–sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block size, and the maximum values of sHor and sVer are 8. The bilateral matching cost is calculated by applying a cost factor to the SATD cost between two reference sub-blocks, as: bilCost = satdCost * costFactor. The search area (2*sHor + 1)*(2*sVer + 1) is divided into 5 diamond search zones, as Figure 23As shown. Each search area is assigned a costFactor, which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in raster scan order from the upper left corner of the area to the lower right corner. When the minimum bilCost in the current search area is less than a threshold equal to sbW*sbH, the full pixel search is terminated; otherwise, the full pixel search continues to the next search area until all search points have been checked. The BM performs a local search to derive the half - sample accuracy halfDeltaMv. The search pattern and cost function are the same as those defined in Section 2.9.1. Existing VVC DMVR fractional - sample refinement is further applied to derive the final deltaMV(sbIdx2). Then, the refined MV for the second pass is derived as: · MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2); · MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2). 2.15.3. Third Pass - Sub - block - based Bidirectional Optical Flow MV Refinement In the third pass, the refined MV is derived by applying BDOF to 8x8 grid sub - blocks. For each 8×8 sub - block, starting from the refined MV of the parent - child block in the second pass, BDOF refinement is applied to derive the scaled Vx and Vy without clipping. The derived bioMv(Vx, Vy) is rounded to 1 / 16 - sample accuracy and clipped between - 32 and 32. The refined MV for the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as: · MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv; · MV1_pass3(sbIdx3) = MV0_pass2(sbIdx2) - bioMv. 2.16. Sample - based BDOF In sample - based BDOF, instead of deriving the motion refinement (Vx, Vy) on a block basis, BDOF is performed for each sample. The coding / decoding block is divided into 8 sub-blocks of 8×8. For each sub-block, the BDOF is determined by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to the sub-block, for each sample point in the sub-block, a sliding 5x5 window is used, and the existing BDOF process is applied to each sliding window to derive Vx and Vy. The derived motion refinement (Vx, Vy) is applied to adjust the bi-predicted sample value for the central sample point of the window. 2.17. Extended Merge Prediction In VVC, the Merge candidate list is constructed by sequentially including the following five types of candidates: (1) Spatial MVPs from spatial neighboring CUs; (2) Temporal MVPs from co-located CUs; (3) History-based MVPs from the FIFO table; (4) Pairwise-averaged MVPs; (5) Zero MV. The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU code in the Merge mode, the index of the best Merge candidate is encoded using truncated unary binary (TU). The first binary bit of the Merge index is coded using context, while bypass coding is used for the other binary bits. The derivation process for each category of Merge candidates is provided in this section. Similar to what is done in HEVC, VVC also supports parallel derivation of the Merge candidate lists for all CUs within a certain sized region. 2.17.1. Spatial Candidate Derivation The derivation of spatial Merge candidates in VVC is the same as that in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates at the shown positions, at most four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Only when one or more CUs at positions B0, A0, B1, A1 are unavailable (e.g., because it belongs to another strip or slice) or are intra-coded, is position B2 considered. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates, which ensures that candidates with the same motion information are excluded from the list, thus improving the coding / decoding efficiency. To reduce the computational complexity, not all possible candidate pairs are considered in the mentioned redundancy check. Figure 24 The positions of the spatial Merge candidates are shown, and Figure 25 The candidate pairs considered for the redundancy check of the spatial Merge candidates are shown. Instead, only Figure 25Pairs linked by arrows are used, and a candidate is added to the list only if the corresponding candidate for redundancy check does not have the same motion information. 2.17.2. Temporal Candidate Derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, scaled motion vectors are derived based on co-located CUs belonging to co-located reference pictures. The reference picture list to be used for deriving the co-located CUs is signaled explicitly in the slice header. As Figure 26 shown by the dashed line in, the scaled motion vector for the temporal Merge candidate is obtained, which is scaled from the motion vector of the co-located CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the co-located picture and the co-located picture. The reference picture index of the temporal Merge candidate is set to be equal to zero. The position of the temporal candidate is selected between candidates C0 and C1, as Figure 27 shown. If the CU at position C0 is unavailable, intra-coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal Merge candidate. 2.17.3. History-based Merge Candidate Derivation After spatial MVP and TMVP, history-based MVP (HMVP) Merge candidates are added to the Merge list. In this method, the motion information of previously decoded blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. When a new CTU row is encountered, the table is reset (emptied). Whenever there is a non-sub-block inter-coded CU, the associated motion information is added as a new HMVP candidate to the last entry of the table. The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a constrained first-in-first-out (FIFO) rule is used, where redundancy check is first applied to find if there is the same HMVP in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are moved forward. HMVP candidates can be used in the Merge candidate list construction process. The several most recent HMVP candidates in the table are checked in order and inserted into the candidate list after the TMVP candidates. Redundancy check is applied to the HMVP candidates for spatial or temporal Merge candidates. To reduce the number of redundancy check operations, the following simplification is introduced: Is the number of HMPV candidates for Merge list generation set to (N <= 4)? M : (8 - N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table? Once the total number of available Merge candidates reaches one less than the maximum allowed Merge candidates, the Merge candidate list construction process from HMVP is terminated. 2.17.4. Pairwise average Merge candidate derivation Pairwise average candidates are generated by averaging predefined candidate pairs in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices of the Merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are both available in a list, they are averaged even if the two motion vectors point to different reference pictures; if only one motion vector is available, that motion vector is directly used; if no motion vector is available, this list is kept invalid. When the Merge list is not full after adding pairwise average Merge candidates, zero MVPs are inserted at the end until the maximum Merge candidate number is reached. 2.17.5. Merge estimation region The Merge estimation region (MER) allows for independent derivation of the Merge candidate list for a CU within the same Merge estimation region (MER). For generating the Merge candidate list of the current CU, candidate blocks within the same MER as the current CU are not included. Additionally, the update process for the candidate list of the history-based motion vector prediction values is updated only when (xCb + cbWidth) >> Log2ParMrgLevel is greater than xCb >> Log2ParMrgLevel and (yCb + cbHeight) >> Log2parMrglevel is greater than (yCb >> Log2ParMrgLevel), where (xCb, yCb) is the top-left luma sample position of the current CU in the picture and (cbWidth, cbHeight) is the CU size. The MER size is selected at the encoder side and signaled in the sequence parameter set in the form of log2_parallel_merge_level_minus2. 2.18. New Merge candidates 2.18.1. Non-adjacent Merge candidate derivation In VVC Figure 28 The five spatial neighboring blocks and one temporal neighboring block shown are used to derive Merge candidates. It is proposed to derive additional Merge candidates from positions not adjacent to the current block using the same style as in VVC. To achieve this, for each search round i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated by the following formula: Offsetx = -i × gridX, Offsety = -i × gridY where Offsetx and Offsety represent the offsets of the upper left corner of the virtual block relative to the upper left corner of the current block, and gridX and gridY are the width and height of the search grid. Second, the width and height of the virtual block are calculated by the following formula: newWidth = i × 2 × gridX + currWidth newHeight = i × 2 × gridY + currHeight. where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block. gridX and gridY are currently set to currWidth and currHeight respectively. Figure 29 The relationship between the virtual block and the current block is shown. After generating the virtual block, block A i 、B i 、C i 、D i and E i can be regarded as the VVC spatial neighboring blocks of the virtual block, and their positions are obtained using the same style as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i 、B i 、C i 、D i and E i are the spatial neighboring blocks used in the VVC Merge mode. When constructing the Merge candidate list, deduplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non - adjacent spatial neighboring blocks are used. The non - adjacent spatial Merge candidates are inserted into the Merge list after the temporal Merge candidates in the order of B 1 ->A 1 ->C 1 ->D 1 ->E 1 2.18.2. STMVP It is proposed to use three spatial Merge candidates and one temporal Merge candidate to derive an average candidate as the STMVP candidate. STMVP is inserted before the top - left spatial Merge candidate. The STMVP candidate is deduplicated together with all previous Merge candidates in the Merge list. For spatial candidates, the first three candidates in the current Merge candidate list are used. For temporal candidates, the same position as that of the VTM / HEVC in the same position is used. For spatial candidates, the first, second, and third candidates inserted before STMVP in the current Merge candidate list are denoted as F, S, and T. The temporal candidate with the same position as that of the VTM / HEVC used in TMVP is denoted as Col. The motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows: 1) If the reference indices of all four Merge candidates are valid and all equal to zero in the prediction direction X (X = 0 or 1), then mvLX=(mvLX_F + mvLX_S + mvLX_T + mvLX_Col)>>2. 2) If the reference indices of three of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1), mvLX=(mvLX_F×3 + mvLX_S×3 + mvLX_Col×2)>>3, or mvLX=(mvLX_F×3 + mvLX_T×3 + mvLX_Col×2)>>3, or mvLX=(mvLX_S×3 + mvLX_T×3 + mvLX_Col×2)>>3. 3) If the reference indices of two of the four Merge candidates are valid and equal to zero in the prediction direction X (X = 0 or 1), mvLX=(mvLX_F + mvLX_Col)>>1, or mvLX=(mvLX_S + mvLX_Col)>>1, or mvLX=(mvLX_T + mvLX_Col)>>1. Note: If the temporal candidate is not available, the STMVP mode is turned off. 2.18.3. Merge List Sizes If both non - adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 8. 2.19. Geometric Partitioning Mode (GPM) In VVC, the geometric partitioning mode is supported for inter - prediction. The CU - level flag is used as a Merge mode to signal the geometric partitioning mode. Other Merge modes include the regular Merge mode, MMVD mode, CIIP mode, and sub - block Merge mode. For each possible CU size w×g = 2 m ×2 n , where m,n∈{3…6} excluding 8x64 and 64x8, the geometric partitioning mode supports a total of 64 partitions. When this mode is used, the CU is divided into two parts by a geometrically - positioned line ( Figure 30 ). The position of the dividing line is mathematically derived from the angle and offset parameters of a specific partition. Each part in the geometric partitioning of the CU is inter - predicted using its own motion; only unidirectional prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. The unidirectional prediction motion constraint is applied to ensure the same as traditional bidirectional prediction, and each CU only requires two motion - compensated predictions. The unidirectional prediction motion for each partition is derived using the process described in 2.19.1. If the geometric partitioning mode is used for the current CU, the geometric partitioning index (angle and offset) indicating the partitioning mode of the geometric partitioning and two Merge indices (one for each partition) are further signaled. The number of maximum GPM candidate sizes is explicitly signaled in the SPS, and the syntax binarization for the GPM Merge indices is specified. After predicting each part of the geometric partitioning, as in 2.19.2, a hybrid process with adaptive weights is used to adjust the sample values along the geometric partitioning edge. This is the prediction signal for the entire CU, and the transform process and quantization process will be applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partitioning mode is stored as described in 2.19.3. 2.19.1. Unidirectional Prediction Candidate List Construction In 2.17, the unidirectional prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process. Let n denote the index of the unidirectional prediction motion in the geometric unidirectional prediction candidate list. The LX motion vector (where X equals the parity of n) of the n-th extended Merge candidate is used as the n-th unidirectional prediction motion vector of the geometric partitioning pattern. These motion vectors are marked with "x" in Figure 31 . If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(1-X) motion vector of the same candidate is used as the unidirectional prediction motion vector of the geometric partitioning pattern. 2.19.2. Blending along the geometric partitioning edge After using its own motion prediction for each part of the geometric partitioning, blending is applied to the two prediction signals to derive the samples around the geometric partitioning edge. The blending weight for each position of the CU is derived based on the distance between the independent position and the partitioning edge. The distance from the position (x, y) to the partitioning edge is derived as: where i, j are the indices of the angle and offset of the geometric partitioning, which depend on the geometric partitioning index transmitted through the signal. p x,j and p y,j 's signs depend on the angle index i. The weight for each part of the geometric partitioning is derived as follows: wIdxL(x, y) = partIdx? 32 + d(x, y) : 32 - d(x, y) (2-29) partIdx depends on the angle index i. An example of the weight w 0 is shown in Figure 32 . 2.19.3. Motion field storage for geometric partitioning patterns Mv1 from the first part of the geometric partitioning, Mv2 from the second part of the geometric partitioning, and the combined Mv of Mv1 and Mv2 are stored in the motion field of the CU encoded / decoded for the geometric partitioning pattern. The type of motion vector stored for each independent position in the motion field is determined as: sType = abs(motionIdx) < 32? 2 : (motionIdx <= 0? (1 - partdx) : partdx) (2-32) where motionIdx equals d(4x + 2, 4y + 2), which is recalculated according to Equation (2-18). partIdx depends on the angle index. If sType is equal to 0 or 1, Mv0 or Mv1 is stored in the corresponding motion field. Otherwise, if sType is equal to 2, the combined Mv from Mv1 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), Mv1 and Mv2 are simply combined to form a bi-predictive motion vector. Otherwise, if Mv1 and Mv2 are from the same list, only the uni-predictive motion Mv2 is stored. 2.20. Multiple Hypothesis Prediction In Multiple Hypothesis Prediction (MHP), on top of the inter-frame AMVP mode, regular Merge mode, affine Merge, and MMVD mode, up to two additional prediction values are signaled. The resulting overall prediction signal is iteratively accumulated using each additional prediction signal. p n+1 =(1 - α n+1 )p n +α n+1 h n+1 The weight factor α is specified according to Table 2-4 below. Table 2-4 - Weight Factors for MHP add_hyp_weight_idx α 0 1 / 4 1 -1 / 8 For the inter-frame AMVP mode, MHP is applied only when unequal weights in BCW are selected in the bi-predictive mode. The additional hypothesis can be either the Merge mode or the AMVP mode. In the case of the Merge mode, the motion information is indicated by the Merge index, and the Merge candidate list is the same as in the geometric partitioning mode. In the case of the AMVP mode, the reference index, MVP index, and MVD are signaled. 2.21. Non-Adjacent Spatio-Temporal Candidates Non-adjacent spatio-temporal Merge candidates are inserted after the TMVP in the regular Merge candidate list. The pattern of the spatio-temporal Merge candidates is shown in Figure 33 . The distance between the non-adjacent spatio-temporal candidate and the current coding block is based on the width and height of the current coding block. 2.22. Template Matching (TM) Template Matching (TM) is a decoder-side MV derivation method for refining the motion information of the current CU by finding the closest match between a template in the current picture (i.e., the top and / or left neighboring blocks of the current CU) and a block in the reference picture (i.e., of the same size as the template). As Figure 34As shown, within the [-8, +8] pixel search range, a better MV is searched around the initial motion of the current CU. The previously proposed template matching is adopted with two modifications: the search step size is determined based on the AMVR mode, and in the Merge mode, TM can be cascaded by leveraging the bilateral matching process. In the AMVP mode, an MVP candidate is determined by selecting the one that achieves the minimum difference between the current block template and the reference block template based on the template matching error, and then TM only performs MV refinement on this specific MVP candidate. TM refines this MVP candidate by using iterative diamond search starting from the full pixel MVD accuracy within the [-8, +8] pixel search range (or 4 pixels for the 4-pixel AMVR mode). The AMVP candidate can be further refined by using cross search with full pixel MVD accuracy (or 4 pixels for the 4-pixel AMVR mode), followed by using half pixels and quarter pixels in sequence depending on the AMVR mode as specified in Table 2-5. This search process ensures that the MVP candidate still maintains the same MV accuracy as indicated by the AMVR mode after the TM process. Table 2-5 - Search patterns of AMVR and Merge mode using AMVR. In the Merge mode, a similar search method is applied to the Merge candidates indicated by the Merge index. As shown in Table 2-5, TM can be performed until 1 / 8 pixel MVD accuracy, or skip those accuracies that exceed half pixel MVD accuracy, depending on whether an alternative interpolation filter is used based on the merged motion information (i.e., used when AMVR is in the half pixel mode). Additionally, when the TM mode is enabled, the template matching can work in an independent process or an additional MV refinement process between the block-based bilateral matching (BM) method and the sub-block-based bilateral matching method, depending on whether BM can be enabled according to its enabling condition check. 2.23. Overlapped Block Motion Compensation (OBMC) Overlapped block motion compensation (OBMC) has been used in H.263 before. In JEM, different from H.263, OBMC can be turned on and off using the syntax at the CU level. When OBMC is used in JEM, OBMC is performed on all motion compensation (MC) block boundaries except the right and bottom boundaries of the CU. In addition, it applies to both the luma and chroma components. In JEM, the MC block corresponds to the coding / decoding block. When a CU is coded / decoded using sub-CU modes (including sub-CU Merge, affine, and FRUC modes), each sub-block of the CU is an MC block. To process the CU boundary in a unified way, OBMC is performed on all MC block boundaries at the sub-block level, where the sub-block size is set to be equal to 4×4, as Figure 35 shown. When OBMC is applied to the current sub-block, in addition to the current motion vector, the motion vectors of the four connected neighboring sub-blocks, if available and different from the current motion vector, are also used to derive the predicted block of the current sub-block. These multiple predicted blocks based on multiple motion vectors are combined to generate the final predicted signal of the current sub-block. Denote the predicted block based on the motion vector of the neighboring sub-block as P N , where N indicates the indices of the neighboring upper, lower, left, and right sub-blocks, and the predicted block based on the motion vector of the current sub-block is denoted as P C . When P N is based on the motion information of the neighboring sub-blocks that contains the same motion information as the current sub-block, OBMC is not performed from P N . Otherwise, each sample of P N is added to the same sample in P C , i.e., the four rows / columns of P N are added to P C . The weight factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for P N , and the weight factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for P C . Except for small MC blocks, (i.e., when the height or width of the coding / decoding block is equal to 4, or when a CU is coded / decoded using sub-CU modes), where only two rows / columns of P N are added to P C . In this case, the weight factors {1 / 4, 1 / 8} are used for P N , while the weight factors {3 / 4, 7 / 8} are used for P C . For P N generated based on the motion vectors of the vertical (horizontal) neighboring sub-blocks, the samples in the same row (column) of P N are added to P C with the same weight factor. In JEM, for a CU with a size less than or equal to 256 luma samples, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU. For a CU with a size greater than 256 luma samples or not coded using the AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its effect is taken into account during the motion estimation stage. OBMC uses the prediction signals formed from the motion information of the top neighboring block and the left neighboring block to compensate the top and left boundaries of the original signal of the current CU, and then the normal motion estimation process is applied. 2.24. Multiple Transform Selection (MTS) for Kernel Transform In addition to DCT-II already adopted in HEVC, the Multiple Transform Selection (MTS) scheme is also used for residual coding of inter- and intra-coded blocks. It uses multiple selected transforms from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 2-6 shows the basis functions of the selected DST / DCT. Table 2-6 - Transform Basis Functions of DCT-II / VIII and DSTVII for N-Point Input To maintain the orthogonality of the transform matrix, the quantization of the transform matrix is more accurate than that of the transform matrix in HEVC. To keep the intermediate values of the transform coefficients within the 16-bit range, all coefficients are 10 bits after horizontal and vertical transforms. To control the MTS scheme, separate enable flags are specified at the SPS level for intra and inter, respectively. When MTS is enabled at the SPS, a CU-level flag is signaled to indicate whether MTS is applied. Here, MTS is only applicable to luma. The MTS signaling is skipped when one of the following conditions is met. – The position of the last significant coefficient of the luma TB is less than 1 (i.e., only DC). – The last significant coefficient of the luma TB is within the MTS zeroing region. If the MTS CU flag is equal to zero, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two additional flags are signaled to indicate the transform types in the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 2-7. A unified transform selection for ISP and implicit MTS is used by eliminating the intra mode and block shape dependencies. If the current block is in ISP mode or if the current block is an intra block and both intra and inter explicit MTS are turned on, only DST7 is used for the horizontal and vertical transform kernels. In terms of the transform matrix precision, 8-bit primary transform kernels are used. Therefore, all the transform kernels used in HEVC remain the same, including 4-point DCT-2 and DST-7, 8-point, 16-point, and 32-point DCT-2. In addition, for other transform kernels (including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7, and DCT-8), 8-bit primary transform kernels are used. Table 2-7 - Transform and Signaling Mapping Table To reduce the complexity of large-size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with a size (width or height, or both width and height) equal to 32, the high-frequency transform coefficients are set to zero. Only the coefficients within the 16×16 low-frequency region are retained. As in HEVC, the residual of a block can be coded and decoded using the transform skip mode. To avoid redundancy in syntax coding, when the CU-level MTS_CU_flag is not equal to 0, the transform skip flag is not signaled. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. In addition, when MTS is enabled for an inter-coded block, implicit MTS can still be enabled. 2.25. Sub-Block Transform (SBT) In VTM, a sub-block transform is introduced for inter-predicted CUs. In this transform mode, for a CU, only a sub-part of the residual block is coded and decoded. When the cu_cbf of an inter-predicted CU is equal to 1, the cu_sbt_flag can be signaled to indicate whether the entire residual block or a sub-part of the residual block is coded and decoded. For the former case, the inter MTS information is further parsed to determine the transform type of the CU. In the latter case, a part of the residual block is coded and decoded by a presumed adaptive transform while the other part of the residual block is set to zero. When SBT is used for an inter-coded CU, the SBT type and SBT position information are signaled in the bitstream. There are two SBT types and two SBT positions, as Figure 36As shown. For SBT-V (or SBT-H), the TU width (or height) can be equal to half of the CU width (or height) or 1 / 4 of the CU width (or height), resulting in a 2:2 partition or a 1:3 / 3:1 partition. The 2:2 partition is similar to the binary tree (BT) partition, while the 1:3 / 3:1 partition is similar to the asymmetric binary tree (ABT) partition. In the ABT partition, only small regions contain non-zero residuals. If one dimension of the CU is 8 (in terms of luma samples), a 1:3 / 3:1 partition along that dimension is not allowed. A CU has at most 8 SBT modes. Position-dependent transform kernel selection is applied to the luma transform blocks in SBT-V and SBT-H (chroma TBs always use DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, the horizontal and vertical transforms for each SBT position are Figure 36 specified in. For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of the residual TU is greater than 32, the transforms for both dimensions are set to DCT-2. Thus, the sub-block transforms jointly specify the TU slice, cbf, and the horizontal and vertical kernel transform types of the residual block. SBT is not applied to the CUs coded with the combined inter-intra mode. 2.26. Adaptive Merge Candidate Reordering Based on Template Matching To improve the coding efficiency, after constructing the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. The Merge candidates are arranged in the list according to the ascending template matching cost. It is operated in subgroups. The template matching cost is measured by the SAD (sum of absolute differences) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes bi-predicted motion information, the corresponding reference sample is the average of the corresponding reference samples in reference list 0 and the corresponding reference samples in reference list 1, as Figure 37 shown in. If the Merge candidate contains motion information at the sub-CU level, the corresponding reference sample consists of the neighboring samples of the corresponding reference sub-block, as Figure 38 shown in. The sorting process is operated in subgroups, as Figure 39 shown in. The first three Merge candidates are sorted together. The next three Merge candidates are sorted together. The template size (width of the left template or height of the upper template) is 1. The subgroup size is 3. 2.27. Adaptive Merge Candidate List An adaptive Merge candidate list is proposed, also known as Adaptive Reordering of Merge Candidates (ARMC). It can be assumed that the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first subgroup, and the subsequent 3 Merge candidates are taken as the second subgroup (i.e., the last subgroup). For the encoder, after the Merge candidate list is constructed, some Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, as Figure 40 shown. More specifically, the template matching costs of the Merge candidates in all subgroups except the last subgroup are calculated; then, the Merge candidates in their own subgroups except the last subgroup are reordered; finally, the final Merge candidate list is obtained. For the decoder, after the Merge candidate list is constructed, as Figure 41 shown, some / none of the Merge candidates are adaptively reordered in ascending order of the Merge candidate cost. In Figure 41 , the subgroup where the selected (for signal transmission) Merge candidate is located is called the selected subgroup. FIG. 57 shows the reordering process in the decoder. More specifically, if the selected Merge candidate is in the last subgroup, the Merge candidate list construction process is terminated after the selected Merge candidate is derived, no reordering is performed, and the Merge candidate list is not changed; otherwise, the following process is performed: After all the Merge candidates in the selected subgroup are derived, the Merge candidate list construction process is terminated; the template matching costs of the Merge candidates in the selected subgroup are calculated; the Merge candidates in the selected subgroup are reordered; finally, a new Merge candidate list is obtained. For both the encoder and the decoder, the template matching cost is derived as a function of T and RT, where T is the set of sample points in the template and RT is the set of reference sample points for the template. When deriving the reference sample points of the template of the Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision. The reference sample points (RT) for the template used for bidirectional prediction are derived by weighted averaging of the reference sample points (RT 0 ) of the template in reference list 0 and the reference sample points (RT 1 ) of the template in reference list 1. RT = ((8 - w) * RT 0 + w * RT 1 + 4) >> 3 (2 - 33) The weights of the reference templates in reference list 0 (8 - w) and the weights of the reference templates in reference list 1 (w) are determined by the BCW index of the Merge candidate. The BCW indices equal to {0, 1, 2, 3, 4} correspond to w equal to {-2, 3, 4, 5, 10}, respectively. If the local illumination compensation (LIC) flag of the Merge candidate is true, the LIC method is used to derive the reference sample points of the template. The template matching cost is calculated based on the sum of absolute differences (SAD) between T and RT. The template size is 1. This means that the width of the left - hand template and / or the height of the upper template is 1. If the codec mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered. If the codec mode is GPM, the Merge candidates used to derive the unidirectional prediction candidate list are not reordered. 2.28. Geometric prediction mode with motion vector difference In the geometric prediction mode with motion vector difference (GMVD), each geometric partition in GPM can decide whether to use GMVD. If GMVD is selected for a geometric region, the MV of that region is calculated as the sum of the MV of the Merge candidate and the MVD. All other processing remains the same as in GPM. Using GMVD, the MVD is signaled in the form of a direction - and - distance pair. There are nine candidate distances (1 / 4 - pixel, 1 / 2 - pixel, 1 - pixel, 2 - pixel, 3 - pixel, 4 - pixel, 6 - pixel, 8 - pixel, 16 - pixel), and eight candidate directions (four horizontal / vertical directions and four diagonal directions). Additionally, when pic_fpel_mmvd_enabled_flag equals 1, the MVD in GMVD is also left - shifted by 2 bits as in MMVD. 2.29. Affine MMVD In affine MMVD, an affine Merge candidate (referred to as the base affine Merge candidate) is selected, and the MV of the control points is further refined by the signaled MVD information. The MVD information of the MVs of all control points is the same in one prediction direction. When the starting MV is a bi - directional prediction MV where the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture while the other reference POC is less than the POC of the current picture), the MV offsets added to the list 0 MV component of the starting MV and the MV offset of the list 1 MV have opposite values; otherwise, when the starting MV is a bi - directional prediction MV where both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture, or both are less than the POC of the current picture), the MV offsets added to the list 0 MV component of the starting MV and the MV offset of the list 1 MV are the same. 2.30. Adaptive Decoder - side Motion Vector Refinement (ADMVR) In ECM - 2.0, if the selected Merge candidate satisfies the DMVR condition, the multi - pass decoder - side motion vector refinement (DMVR) method is applied in the regular Merge mode. In the first pass, bilateral matching (BM) is applied to the coded - decoded block. In the second pass, BM is applied to each 16x16 sub - block within the coded - decoded block. In the third pass, the MVs in each 8x8 sub - block are refined by applying bidirectional optical flow (BDOF). The adaptive decoder - side motion vector refinement method consists of two new Merge modes that are introduced to refine the MV only in one direction (L0 or L1) of the bi - directional prediction of the Merge candidate that satisfies the DMVR condition. The multi - pass DMVR process is applied to the selected Merge candidate to refine the motion vector. However, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is set to zero. Similar to the regular Merge mode, the Merge candidates for the proposed Merge mode are derived from spatially neighboring coded - decoded blocks, TMVP, non - adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The same Merge candidate list (i.e., the ADMVR Merge list) is used by the two proposed Merge modes, and the Merge index is coded - decoded in the regular Merge mode. 2.31. IBC with Template Matching The proposal also uses template matching with IBC for both the IBC Merge mode and the IBC AMVP mode. The IBC-TM Merge list has been modified compared to that used by the regular IBC Merge mode such that candidates are selected according to a deduplication method with motion distances between candidates in the regular TM Merge mode. End-zero motion completion (which is meaningless for intra-coding) has been replaced by motion vectors to the left (-W, 0), top (0, -H), and top-left (-W, -H) CUs, and then the list is executed using the left CU without deduplication if needed. In the IBC-TM Merge mode, the selected candidates are refined using a template matching method before the RDO or decoding process. The IBC-TM Merge mode has been made to compete with the regular IBC Merge mode, and the TM Merge flag is signaled. In the IBC-TM AMVP mode, up to 3 candidates are selected from the IBC Merge list. Each of these 3 selected candidates is refined using a template matching method and sorted according to their resulting template matching costs. Then usually only the first two are considered during the motion estimation process. The template matching refinement for both the IBC-TM Merge and AMVP modes is very simple because the IBC motion vectors are constrained to be integers and within the reference region as shown in Figure 42 Therefore, in the IBC-TM Merge mode, all refinements are performed with integer precision, and in the IBC-TM AMVP mode, all refinements are performed with integer or 4-pixel precision. In both cases, the refined motion vectors in each refinement step must comply with the constraints of the reference region. 2.32. IBC Merge Mode with Block Vector Difference The IBC Merge mode with block vector difference is as follows. The distance set is {1 pixel, 2 pixels, 4 pixels, 8 pixels, 12 pixels, 16 pixels, 24 pixels, 32 pixels, 40 pixels, 48 pixels, 56 pixels, 64 pixels, 72 pixels, 80 pixels, 88 pixels, 96 pixels, 104 pixels, 112 pixels, 120 pixels, 128 pixels}, and the BVD directions are two horizontal directions and two vertical directions. The base candidate is selected from the top five candidates in the re-ordered IBC Merge list. And for each base candidate, all possible MBVD refinement positions (20x4) are re-ordered based on the SAD cost between the template (one row above and one column to the left of the current block) and the reference for each refinement position. Finally, the top 8 refinement positions with the lowest template SAD cost are kept as available positions and thus used for MBVD index coding and decoding. 2.33. Reconstruction Re-ordered IBC (RR-IBC) Screen content coding and decoding tools similar to Intra Block Copy (IBC) generate prediction blocks by directly copying previously coded and decoded reference regions in the same picture. Symmetry is often observed in video content, especially in text character regions and computer-generated graphics in screen content sequences, such as Figure 43 shown. Therefore, specific screen content coding and decoding tools that consider symmetry will effectively compress such video content. A Reconstruction-Reordering IBC (RR-IBC) mode for screen content video coding and decoding is proposed. When RR-IBC is applied, the samples in the block are flipped according to the flipping type of the current block. On the encoder side, the original block is flipped before motion search and residual calculation, while the prediction block is derived without flipping. On the decoder side, the reconstructed block is flipped to restore the original block. Two flipping methods, horizontal flipping and vertical flipping, are supported for blocks coded and decoded with RR-IBC. First, the block syntax flag for IBC AMVP coding and decoding is signaled, indicating whether the reconstruction is flipped. If it is flipped, another flag specifying the flipping type is further signaled. For IBC Merge, without syntax signaling, the flipping type is inherited from neighboring blocks. Considering horizontal or vertical symmetry, the current block and the reference block are usually horizontally or vertically aligned. Therefore, when horizontal flipping is applied, the vertical component of BV is not signaled and is presumed to be equal to 0. Similarly, when vertical flipping is applied, the horizontal component of BV is not signaled and is presumed to be equal to 0. To better utilize the symmetry feature, a flipping-aware BV adjustment method is applied to refine the block vector candidates. Figure 44A An illustration of BV adjustment for horizontal flipping is shown, and Figure 44B an illustration of BV adjustment for vertical flipping is shown. For example, as Figure 44A and 44B shown, (x nbr, y nbr ) and (x cur , y cur ) represent the coordinates of the central samples of the neighboring block and the current block, respectively. BV nbr and BV cur represent the BV of the neighboring block and the current block, respectively. Instead of directly inheriting BV from the neighboring block, when the neighboring block is coded with horizontal flipping, the horizontal component of BV nbr (denoted as BV nbr h ) is added to the motion displacement to calculate the horizontal component of BV cur , i.e., BV cur h = 2(x nbr - xcur ) + BV nbr h . Similarly, when encoding and decoding are performed using a vertically flipped neighboring block, the motion displacement is added to BV nbr for the vertical component of (denoted as BV nbr v ) to calculate BV cur for the vertical component, i.e., BV cur v = 2(y nbr - y cur ) + BV nbr v . 2.34. Intra-frame Template Matching Intra-frame Template Matching Prediction (Intra TMP) is a special intra-frame prediction mode that copies the best prediction block from the reconstructed part of the current frame, and its L-shaped template matches the current template. For a predefined search range, the encoder searches for the template most similar to the current template in the reconstructed part of the current frame and uses the corresponding block as the prediction block. The encoder then signals the use of this mode, and the same prediction operation is performed on the decoder side. By matching the L-shaped temporary neighbor of the current block with another block in the predefined search region consisting of Figure 45 the prediction signal is generated as follows: R1: the current CTU; R2: the top-left CTU; R3: the upper CTU; R4: the left CTU. SAD is used as the cost function. Within each region, the decoder searches for the template with the minimum SAD relative to the current one and uses its corresponding block as the prediction block. The dimensions of all regions (SearchRange_w, SearchRange_h) are proportional to the block dimensions (BlkW, BlkH) with a fixed number of SAD comparisons per pixel. That is: SearchRange_w = a * BlkW; SearchRange_h = a * BlkH. where 'a' is a constant that controls the gain / complexity trade-off. In fact, 'a' is equal to 5. The intra-frame template matching tool is enabled for CUs with dimensions less than or equal to 64 in width and height. This maximum CU size for intra-frame template matching is configurable. When DIMD is not used for the current CU, the intra-frame template matching prediction mode is signaled at the CU level through a dedicated flag. 2.35. Diversity Reordering for ARMC This method modifies the reordering process of ARMC. The goal is to create diversity within the Merge candidate list. The proposed method looks for candidates that are overly redundant in the rate-distortion (RD) sense. In the proposed algorithm, a candidate is considered redundant if the cost difference between the candidate and its predecessor is lower than the lambda value (e.g., |D1 - D2| < λ, where D1 and D2 are the costs obtained during the first ARMC sorting and λ is the Lagrange parameter used in the RD criterion at the encoder side). The proposed algorithm is defined as follows: - Determine the minimum cost difference between a candidate and its predecessor among all candidates in the list, - If the minimum cost difference is greater than or equal to λ, the list is considered diverse enough and reordering stops. - If the minimum cost difference is lower than λ, the candidate is considered redundant and it is moved to another position in the table. This other position is the first position where the candidate is diverse enough compared to its predecessor. - The algorithm stops after a finite number of iterations (if the minimum cost difference is not lower than λ). This algorithm is applied to the regular, TM, BM, and affine Merge modes of ECM - 5.0. Similar algorithms are applied to the Merge MMVD and sign MVD prediction methods, which also use ARMC for reordering. The value of λ is set to be equal to the λ of the rate-distortion criterion used to select the best Merge candidate for the low-latency configuration at the encoder side and the value of λ corresponding to another QP for the random access configuration. A set of λ values corresponding to each signal-transmitted QP offset in the SPS or in the slice header is provided for QP offsets not present in the SPS. 3. Problem In the current design of diversity reordering, a set of thresholds (e.g., the lambda value (λ)) used to evaluate whether a candidate is too redundant in the rate-distortion sense is signaled corresponding to each signal-transmitted QP offset in the SPS or in the slice header. Depending on the QP offset of the slice, a constant threshold is used for the entire slice. However, the QP of the coded / decoded blocks in the slice is different from the slice QP, making it difficult to determine the threshold for diversity reordering and potentially limiting the coding / decoding performance. 4. Detailed Solution The following detailed solutions should be considered as examples to explain the general concept. These solutions should not be interpreted in a narrow way. Additionally, these solutions can be combined in any way. In the present disclosure, the diversity reordering may not be limited to the current art described in Section 2.35. In the present disclosure, Intra Block Copy (IBC) may not be limited to the current IBC techniques, but may be interpreted as a technique in which the reference (or prediction) block is obtained using samples in the current strip / slice / sub-picture / picture / other video unit (e.g., CTU row), excluding the conventional intra prediction methods. The parameter corresponding to the QP or QP offset (e.g., the lambda value for diversity reordering) may mean that the parameter is signaled in the bitstream or is derived using the codec information. Signaling of lambda value for diversity reordering 1. It is proposed that a set of N lambda values used for diversity reordering may be signaled in the bitstream. a. In one example, the lambda value may correspond to a specific quantization parameter (QP). b. In one example, when i < j, the corresponding QPs in a set of QPs, such as QP 1 , QP 2 , …, QP N , satisfy QP i < QP j . i. In one example, how to map lambda and its corresponding QP may be signaled in the bitstream. c. In one example, the lambda values corresponding to all QP values allowed to be used at the decoder may be signaled. i. In one example, QP 1 = 0 and QP N = 63, N = 64. ii. Alternatively, the lambda values corresponding to all QP values may be predefined. d. In one example, the lambda values corresponding to all QPs between the first QP (T 1 ) and the second QP (T 2 ) may be signaled, where QP 1 = T 1 and QP 2 = T 2 . 1) In one example, T 1 may be signaled in the bitstream. 2) In one example, T 1 or T 2 may depend on the picture / strip type and / or the temporal layer. 3) All lambda values corresponding to QPs between the first QP(T 1 ) and the second QP(T 2 ) can be predefined. e. In one example, the difference between two adjacent QPs in a set of QPs can be the same. i. In one example, for any two adjacent QPs (QP i and QP j , where i + 1 = j), ΔQ = QP j - QP i is the same. 1) In one example, ΔQ = 1, or ΔQ = 2, or ΔQ = 3, or ΔQ = 4, or ΔQ = 5. f. In one example, the lambda value can be signaled in a predictive manner. g. In one example, the decoder can determine the lambda value based on the QP of the current picture / strip / block and the mapping between the QP and the lambda value. i. In one example, the lambda value can be determined as the lambda value corresponding to the QP value in a set of QPs that is closest to the QP of the current picture / strip / block. 2. It is proposed that the determination of the lambda value used for diversity reordering of video units can depend on the codec information of the video units. a. In one example, the codec information can refer to the QP or the QP offset. b. In one example, when the lambda value corresponding to the first QP or QP offset (Q 1 ) is in a set of QPs, the lambda value can be used for diversity reordering of the video units. i. In one example, Q 1 can be equal to the current QP or QP offset (Q C ) of the video unit. ii. In another example, Q 1 can be equal to Q C + X or Q C - X, where X is an integer greater than 1. 1) In one example, X can depend on the picture / strip type and / or the codec configuration (e.g., random access or low latency). c. In one example, when the lambda value corresponding to the current QP or QP offset (Q C ) of the video unit is not in a set of QPs, another QP or QP offset (Q 2) The corresponding lambda value can be used for the diversity reordering of video units. i. In one example, when Q 2 is closest to Q 1 the lambda value corresponding to Q2 can be used. 1) In one example, compared with the absolute difference between Q 1 and other QP or QP offsets corresponding to available lambda values, the absolute difference between Q 2 and Q 1 is the smallest. ii. In one example, when there are two closest QP or QP offsets (e.g., QP 2S <QP 2L ), the lambda value corresponding to the smaller QP or QP offset (e.g., QP 2S ) can be used. 1) Alternatively, when there are two closest QP or QP offsets (e.g., QP 2S <QP 2L ), the lambda value corresponding to the larger QP or QP offset (e.g., QP 2L ) can be used. d. In the above examples, a video unit may refer to a color component / subpicture / strip / slice / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transformation unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transformation block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel. 3. In one example, some or all of the lambda values used in the diversity reordering can be derived rather than signaled in the bitstream. 4. In one example, the lambda values used for the diversity reordering of inter prediction tools may be different from those used for the diversity reordering of IBC. General aspects 5. Whether and / or how to apply the methods disclosed above can be signaled at the sequence level / group of pictures level / picture level / strip level / slice group level, e.g., in the sequence header / picture header / SPS / VPS / DPS / DCI / PPS / PPS / strip header / slice group header. 6. Whether and / or how the methods disclosed above can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU rows / strips / slices / sub-pictures / other types of regions containing more than one sample or pixel. 7. Whether and / or how the methods disclosed above can be applied may depend on the transcoding information, such as block size, color format, mono / double tree splitting, color component, strip / picture type.
[0105] More details of embodiments of the present disclosure related to the variations of video blocks for video coding and decoding will be described below. Embodiments of the present disclosure should be considered as examples for explaining general concepts and should not be construed in a narrow manner. In addition, these embodiments can be applied individually or in combination in any way.
[0106] Figure 46 A flowchart of a method 4600 for video processing according to some embodiments of the present disclosure is shown. Method 4600 can be implemented during the conversion between a current video block of a video and the bitstream of the video. As Figure 46 shown, method 4600 starts at 4602, where a target value of a parameter is determined from a set of candidate values of parameters for sorting a plurality of motion candidates for a current video unit. The parameter is associated with a cost difference related to the plurality of motion candidates. By way of example and not limitation, the parameter may be represented by lambda (λ) and is used as a threshold for diversity reordering of a plurality of motion candidates as described in Section 2.35 above. In addition, the current video unit is part of a strip of the video. In other words, the current video unit is at a level lower than the strip level, e.g., at the slice level, etc. As used herein, the strip including the current video unit may also be referred to as the current strip, and the picture including the current video unit may also be referred to as the current picture.
[0107] In some embodiments, the current video unit may be a slice, a coding tree unit (CTU), a CTU row, one or more groups of CTUs, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), a block, a sub-block of a block, a sub-region within a block, etc. In addition, the current video unit may be rectangular or non-rectangular.
[0108] In some embodiments, the target value may be selected from a set of candidate values based on the codec information of the current video unit. Alternatively or additionally, the target value may be selected from a set of candidate values based on the codec information of the current slice. In some further embodiments, the target value may be selected from a set of candidate values based on the codec information of the current picture. By way of example and not limitation, the codec information may be a quantization parameter (QP) value or a QP offset. In some embodiments, the QP value may be determined based on a base QP value and a QP offset. The determination of the target value will be described in detail below.
[0109] At 4604, a transformation based on the target value is performed. In some embodiments, the transformation may include encoding the current video unit into a bitstream. Alternatively or additionally, the transformation may include decoding the current video unit from the bitstream. It should be understood that the above diagrams and / or examples are described for illustrative purposes only. The scope of the present disclosure is not limited in this regard.
[0110] In view of the above, a target value for a parameter (e.g., lambda in diversity reordering) is determined for a video unit that is part of a slice. In other words, the determination of the target value is performed at a level lower than the slice level. Compared with conventional solutions where the value of lambda is determined at the picture level or the slice level, the proposed method can advantageously enable the value of the parameter to be determined in a more refined manner, and thus the codec performance can be improved.
[0111] In some embodiments, a set of candidate values may be indicated in the bitstream. For example, a set of candidate values may be indicated in the bitstream in a predictive manner. Alternatively, a set of candidate values may be predefined. In this case, a set of candidate values may not be present in the bitstream.
[0112] In some embodiments, each candidate value in a set of candidate values may correspond to a quantization parameter (QP) value. One or more QP values corresponding to a set of candidate values constitute a set of QP values corresponding to the set of candidate values. By way of example and not limitation, the QP values in a set of QP values may be sorted in ascending order. For example, the QP values in a set of QP values (denoted as QP 1 、QP 2 、…、QP N ) satisfy QP i <QP j , where N is an integer and 1 ≤ i < j ≤ N.
[0113] In some embodiments, a set of candidate values may include candidate values corresponding to all QP values that are allowed to be used at the decoder. By way of example and not limitation, all QP values include 64 QP values ranging from 0 to 63. It should be understood that the specific values described herein are intended to be exemplary and not to limit the scope of the present disclosure.
[0114] Alternatively, a set of candidate values may include candidate values corresponding to a subset of QP values among all QP values that are allowed to be used at the decoder. In one example, the subset of QP values may include QP values between a first threshold and a second threshold. For example, the first threshold and / or the second threshold may be indicated in the bitstream. Additionally or alternatively, the first threshold or the second threshold may depend on the picture type of the current video unit, the slice type of the current video unit, the temporal layer of the current video unit, etc. In another example, each difference between two adjacent QP values in the subset of QP values may be the same. Further, each difference may be equal to a predetermined value, such as 1, 2, 3, 4, 5, etc. For example, for any two adjacent QPs (denoted as QP i and QP j , where i + 1 = j), the difference ΔQ = QP j - QP i is the same.
[0115] In such embodiments, only a portion of all candidate values corresponding to all QP values is signaled. Thereby, the bits required to encode and decode a set of candidate values can be reduced, and thus the encoding and decoding efficiency can be improved.
[0116] In some embodiments, at 4602, a first QP value may be determined based on the QP value of the current video unit, the QP value of the current slice, or the QP value of the current picture. Further, a target value may be selected from a set of candidate value sets based on the first QP value and the mapping between the QP value and a set of candidate values. As used herein, the mapping between the QP value and a set of candidate values indicates the correspondence between the QP value and a set of candidate values. In one example, the mapping between the QP value and a set of candidate values may be indicated in the bitstream. Alternatively, the mapping between the QP value and a set of candidate values may be predefined and not present in the bitstream.
[0117] In some embodiments, the first QP value may be equal to the QP value of the current video unit. Alternatively, the QP value of the current video unit is adjusted using a first adjustment value to obtain the first QP value. In one example, the first QP value may be equal to the sum of the first adjustment value and the QP value of the current video unit. In another example, the first QP value may be equal to the result of subtracting the first adjustment value from the QP value of the current video unit. By way of example and not limitation, the first adjustment value may be an integer greater than 1. Additionally or alternatively, the first adjustment value may depend on the picture type of the current video unit, the slice type of the current video unit, the encoding and decoding configuration of the current video unit, etc.
[0118] In some embodiments, if it is determined that the first QP value is included in a set of QP values corresponding to a set of candidate values, the target value may be determined as the candidate value corresponding to the first QP value in the set of candidate values. Additionally, if it is determined that the first QP value is not included in the set of QP values corresponding to the set of candidate values, a second QP value may be selected from the set of QP values based on the first QP value, and the target value is determined as the candidate value corresponding to the second QP value in the set of candidate values.
[0119] In some embodiments, the selected second QP value may be the QP value in the set of QP values that is closest to the first QP value. For example, the absolute difference between the first QP value and the second QP value may be the smallest absolute difference among the absolute differences between the first value and each QP value in the set of QP values. In some additional embodiments, there may be multiple QP values in the set of QP values that are closest to the first QP value. In such a case, the second QP value may be the smallest QP value among the multiple QP values. Alternatively, the second QP value may be the largest QP value among the multiple QP values.
[0120] In some embodiments, at 4602, the first QP offset may be determined based on the QP offset of the current video unit, the QP offset of the current slice, or the QP offset of the current picture. Additionally, the target value may be selected from the set of candidate values based on the first QP offset and the mapping between the QP offset and the set of candidate values. As used herein, the mapping between the QP offset and the set of candidate values indicates the correspondence between the QP offset and the set of candidate values. In one example, the mapping between the QP offset and the set of candidate values may be indicated in the bitstream. Alternatively, the mapping between the QP offset and the set of candidate values may be predefined and not present in the bitstream.
[0121] In some embodiments, the first QP offset may be equal to the QP offset of the current video unit. Alternatively, the QP offset of the current video unit may be adjusted using a second adjustment value to obtain the first QP offset. In one example, the first QP offset may be equal to the sum of the second adjustment value and the QP offset of the current video unit. In another example, the first QP offset may be equal to the result of subtracting the second adjustment value from the QP offset of the current video unit. By way of example and not limitation, the second adjustment value may be an integer greater than 1. Additionally or alternatively, the second adjustment value may depend on the picture type of the current video unit, the slice type of the current video unit, the codec configuration of the current video unit, etc.
[0122] In some embodiments, if it is determined that the first QP offset is included in a set of QP offsets corresponding to a set of candidate values, the target value may be determined as the candidate value corresponding to the first QP offset in the set of candidate values. Additionally, if it is determined that the first QP offset is not included in the set of QP offsets, a second QP offset may be selected from the set of QP offsets based on the first QP offset, and the target value is determined as the candidate value corresponding to the second QP offset in the set of candidate values.
[0123] In some embodiments, the selected second QP offset may be the QP offset in the set of QP offsets that is closest to the first QP offset. For example, the absolute difference between the first QP offset and the second QP offset may be the smallest absolute difference among the absolute differences between the first value and each QP offset in the set of QP offsets. In some additional embodiments, there may be multiple QP offsets in the set of QP offsets that are closest to the first QP offset. In this case, the second QP offset may be the smallest QP offset among the multiple QP offsets. Alternatively, the second QP offset may be the largest QP offset among the multiple QP offsets.
[0124] In some embodiments, at least a portion of the set of candidate values may not be present in the bitstream and may be determined at the decoder. Additionally, the set of candidate values for a parameter used in the diversity reordering of inter prediction tools may be different from the set of candidate values for a parameter used in the diversity reordering of intra block copy (IBC).
[0125] In some embodiments, whether to apply the method and / or how to apply the method may be indicated at the sequence level, group of pictures level, picture level, slice level, slice group level, etc. In some embodiments, whether to apply the method and / or how to apply the method may be indicated in the sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), slice header, slice group header, etc.
[0126] In some embodiments, whether to apply the method and / or how to apply the method may be indicated at a region including more than one sample or pixel. By way of example and not limitation, the region may include a prediction block (PB), transform block (TB), codec block (CB), prediction unit (PU), transform unit (TU), codec unit (CU), virtual pipeline data unit (VPDU), codec tree unit (CTU), CTU row, slice, picture, sub - picture, etc.
[0127] In some embodiments, whether to apply the method and / or how to apply the method may depend on the transcoded information. By way of example and not limitation, the transcoded information may include block size, color format, single / double tree segmentation, double tree segmentation, color component, stripe type, picture type, etc.
[0128] According to further embodiments of the present disclosure, there is provided a non-transitory computer-readable recording medium. The non-transitory computer-readable recording medium stores a bitstream generated by a method executed by a device for video processing. In the method, a target value of a parameter is determined from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of a video. The parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a stripe of the video. Further, the bitstream is generated based on the target value.
[0129] According to still further embodiments of the present disclosure, there is provided a method for storing a bitstream of a video. In the method, a target value of a parameter is determined from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of a video. The parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a stripe of the video. Further, the bitstream is generated based on the target value, and the bitstream is stored in a non-transitory computer-readable recording medium.
[0130] Embodiments of the present disclosure may be described according to the following articles, and the features of these articles may be combined in any reasonable manner.
[0131] Article 1. A method for video processing, comprising: determining, for a conversion between a current video unit of a video and a bitstream of the video, a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of the current video unit, wherein the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a stripe of the video; and performing the conversion based on the target value.
[0132] Article 2. The method according to Article 1, wherein the set of candidate values is predefined or indicated in the bitstream.
[0133] Article 3. The method according to any one of Articles 1 to 2, wherein each candidate value in the set of candidate values corresponds to a quantization parameter (QP) value.
[0134] Article 4. The method according to Article 3, wherein the QP values in a set of QP values corresponding to the set of candidate values are sorted in ascending order.
[0135] Article 5. The method according to any one of Articles 1 to 4, wherein the set of candidate values includes candidate values corresponding to all QP values allowed to be used at a decoder.
[0136] Item 6. The method according to any one of Items 1 to 4, wherein a set of candidate values includes candidate values corresponding to a subset of QP values among all QP values allowed to be used at the decoder.
[0137] Item 7. The method according to Item 6, wherein the subset of QP values includes QP values between a first threshold and a second threshold.
[0138] Item 8. The method according to Item 7, wherein the first threshold is indicated in the bitstream.
[0139] Item 9. The method according to any one of Items 7 to 8, wherein the first threshold or the second threshold depends on at least one of the following: the picture type of the current video unit, the slice type of the current video unit, or the temporal layer of the current video unit.
[0140] Item 10. The method according to Item 6, wherein each difference between two adjacent QP values in the subset of QP values is the same.
[0141] Item 11. The method according to Item 7, wherein each difference is equal to a predetermined value.
[0142] Item 12. The method according to any one of Items 5 to 11, wherein all QP values include 64 QP values ranging from 0 to 63.
[0143] Item 13. The method according to any one of Items 2 to 12, wherein a set of candidate values is indicated in the bitstream in a predictive manner.
[0144] Item 14. The method according to any one of Items 1 to 13, wherein determining the target value includes: selecting the target value from a set of candidate values based on one of the following: the codec information of the current video unit, the codec information of the slice, or the codec information of the picture including the current video unit.
[0145] Item 15. The method according to Item 14, wherein the codec information includes a QP value.
[0146] Item 16. The method according to Item 15, wherein selecting the target value includes: determining a first QP value based on one of the following: the QP value of the current video unit, the QP value of the slice, or the QP value of the picture; and selecting the target value from a set of candidate values based on the first QP value and the mapping between the QP value and the set of candidate values.
[0147] Item 17. The method according to Item 16, wherein the mapping between the QP value and the set of candidate values is predefined or indicated in the bitstream.
[0148] Item 18. The method according to any one of Items 16 to 17, wherein the first QP value is equal to the QP value of the current video unit.
[0149] Item 19. The method according to any one of Items 16 to 17, wherein determining the first QP value includes: adjusting the QP value of the current video unit by using a first adjustment value to obtain the first QP value.
[0150] Item 20. The method according to Item 19, wherein the first adjustment value is an integer and greater than 1.
[0151] Item 21. The method according to any one of Items 19 to 20, wherein the first adjustment value depends on at least one of the following: the picture type of the current video unit, the slice type of the current video unit, or the codec configuration used for the current video unit.
[0152] Item 22. The method according to any one of Items 16 to 21, wherein selecting a target value from a set of candidate values based on the first QP value and a mapping includes: if it is determined that the first QP value is included in a set of QP values corresponding to a set of candidate values, determining the target value as the candidate value corresponding to the first QP value in the set of candidate values.
[0153] Item 23. The method according to any one of Items 16 to 22, wherein selecting a target value from a set of candidate values based on the first QP value and a mapping includes: if it is determined that the first QP value is not included in a set of QP values corresponding to a set of candidate values, selecting a second QP value from the set of QP values based on the first QP value; and determining the target value as the candidate value corresponding to the second QP value in the set of candidate values.
[0154] Item 24. The method according to Item 23, wherein the second QP value is the QP value in the set of QP values that is closest to the first QP value.
[0155] Item 25. The method according to Item 24, wherein the absolute difference between the first QP value and the second QP value is the smallest absolute difference among the absolute differences between the first value and each QP value in the set of QP values.
[0156] Item 26. The method according to Item 23, wherein multiple QP values in the set of QP values are closest to the first QP value, and the second QP value is the smallest QP value among the multiple QP values.
[0157] Item 27. The method according to Item 23, wherein multiple QP values in the set of QP values are closest to the first QP value, and the second QP value is the largest QP value among the multiple QP values.
[0158] Item 28. The method according to Item 14, wherein the codec information includes a QP offset.
[0159] Item 29. The method according to Item 28, wherein selecting the target value includes: determining a first QP offset based on one of the following: the QP offset of the current video unit, the QP offset of the slice, or the QP offset of the picture; and selecting the target value from a set of candidate values based on the first QP offset and the mapping between the QP offset and the set of candidate values.
[0160] Item 30. The method according to Item 29, wherein the mapping between the QP offset and the set of candidate values is predefined or indicated in the bitstream.
[0161] Item 31. The method according to any one of Items 29 to 30, wherein the first QP offset is equal to the QP offset of the current video unit.
[0162] Item 32. The method according to any one of Items 29 to 30, wherein determining the first QP offset includes: adjusting the QP offset of the current video unit with a second adjustment value to obtain the first QP offset.
[0163] Item 33. The method according to Item 32, wherein the second adjustment value is an integer and greater than 1.
[0164] Item 34. The method according to any one of Items 32 to 33, wherein the second adjustment value depends on at least one of the following: the picture type of the current video unit, the slice type of the current video unit, or the codec configuration used for the current video unit.
[0165] Item 35. The method according to any one of Items 29 to 34, wherein selecting the target value from a set of candidate values based on the first QP offset and the mapping includes: if it is determined that the first QP offset is included in a set of QP offsets corresponding to the set of candidate values, determining the target value as the candidate value corresponding to the first QP offset in the set of candidate values.
[0166] Item 36. The method according to any one of Items 29 to 35, wherein selecting the target value from a set of candidate values based on the first QP offset and the mapping includes: if it is determined that the first QP offset is not included in a set of QP offsets corresponding to the set of candidate values, selecting a second QP offset from the set of QP offsets based on the first QP offset; and determining the target value as the candidate value corresponding to the second QP offset in the set of candidate values.
[0167] Item 37. The method according to Item 36, wherein the second QP offset is the QP offset in the set of QP offsets that is closest to the first QP offset.
[0168] Item 38. The method according to Item 37, wherein the absolute difference between the first QP offset and the second QP offset is the smallest absolute difference among the absolute differences between the first value and each QP offset in a set of QP offsets.
[0169] Item 39. The method according to Item 36, wherein multiple QP offsets in a set of QP offsets are closest to the first QP offset, and the second QP offset is the smallest QP offset among the multiple QP offsets.
[0170] Item 40. The method according to Item 36, wherein multiple QP offsets in a set of QP offsets are closest to the first QP offset, and the second QP offset is the largest QP offset among the multiple QP offsets.
[0171] Item 41. The method according to any one of Items 1 to 40, wherein the current video unit includes one of the following: a slice, a coding tree unit (CTU), a CTU row, one or more sets of CTUs, a coding unit (CU), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding block (CB), a prediction block (PB), a transform block (TB), a block, a sub-block of a block, or a sub-region within a block.
[0172] Item 42. The method according to any one of Items 1 to 41, wherein at least a part of a set of candidate values does not exist in the bitstream and is determined at the decoder.
[0173] Item 43. The method according to any one of Items 1 to 42, wherein a set of candidate values for a parameter used in the diversity reordering for inter prediction tools is different from a set of candidate values for a parameter used in the diversity reordering for intra block copy (IBC).
[0174] Item 44. The method according to any one of Items 1 to 43, wherein the parameter is represented by lambda (λ).
[0175] Item 45. The method according to any one of Items 1 to 44, wherein whether to apply the method and / or how to apply the method is indicated at one of the following: sequence level, picture group level, picture level, strip level, or slice group level.
[0176] Item 46. The method according to any one of Items 1 to 45, wherein whether to apply the method and / or how to apply the method is indicated in one of the following: sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptive parameter set (APS), strip header, or slice group header.
[0177] Item 47. The method according to any one of Items 1 to 44, wherein whether the method is applied and / or how the method is applied is indicated at a region including more than one sample or pixel.
[0178] Item 48. The method according to Item 47, wherein the region includes at least one of the following: a prediction block (PB), a transform block (TB), a coding / decoding block (CB), a prediction unit (PU), a transform unit (TU), a coding / decoding unit (CU), a virtual pipeline data unit (VPDU), a coding / decoding tree unit (CTU), a CTU row, a slice, a picture, or a sub-picture.
[0179] Item 49. The method according to any one of Items 1 to 49, wherein whether the method is applied and / or how the method is applied depends on the coded / decoded information.
[0180] Item 50. The method according to Item 49, wherein the coded / decoded information includes at least one of the following: block size, color format, single / double tree segmentation, double tree segmentation, color component, slice type, or picture type.
[0181] Item 51. The method according to any one of Items 1 to 50, wherein the conversion includes encoding a current video unit into a bitstream.
[0182] Item 52. The method according to any one of Items 1 to 50, wherein the conversion includes decoding a current video unit from a bitstream.
[0183] Item 53. An apparatus for video processing, including a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of Items 1 to 52.
[0184] Item 54. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method according to any one of Items 1 to 52.
[0185] Item 55. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream being generated by a method executed by an apparatus for video processing, wherein the method includes: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of the video, wherein the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a slice of the video; and generating the bitstream based on the target value.
[0186] Item 56. A method for storing a bitstream of video, comprising: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of the video, wherein the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a slice of the video; generating a bitstream based on the target value; and storing the bitstream in a non-transitory computer-readable recording medium. Example device
[0187] Figure 47 FIG. shows a block diagram of a computing device 4700 in which various embodiments of the present disclosure may be implemented. The computing device 4700 may be implemented as the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300), or may be included in the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).
[0188] It should be understood that Figure 47 the computing device 4700 shown in is for illustrative purposes only and does not imply any limitation to the functionality and scope of the embodiments of the present disclosure in any way.
[0189] As Figure 47 shown, the computing device 4700 includes a general-purpose computing device 4700. The computing device 4700 may include at least one or more processors or processing units 4710, a memory 4720, a storage unit 4730, one or more communication units 4740, one or more input devices 4750, and one or more output devices 4760.
[0190] In some embodiments, the computing device 4700 may be implemented as any user terminal or server terminal having computing capabilities. The server terminal may be a server provided by a service provider, a large computing device, etc. The user terminal may be, for example, any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, stations, units, devices, multimedia computers, multimedia tablet computers, Internet nodes, communicators, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 4700 may support any type of interface to the user (such as a "wearable" circuitry, etc.).
[0191] The processing unit 4710 can be a physical processor or a virtual processor, and can implement various processes based on the programs stored in the memory 4720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 4700. The processing unit 4710 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0192] The computing device 4700 generally includes various computer storage media. Such media can be any media accessible by the computing device 4700, including but not limited to volatile media and non-volatile media, or removable media and non-removable media. The memory 4720 can be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory), or any combination thereof. The storage unit 4730 can be any removable or non-removable media, and can include machine-readable media, such as a memory, a flash drive, a magnetic disk, or other media that can be used to store information and / or data and can be accessed in the computing device 4700.
[0193] The computing device 4700 can also include additional removable / non-removable storage media, volatile / non-volatile storage media. Although not shown in Figure 47 it, a disk drive for reading and / or writing to / from a removable non-volatile magnetic disk, and an optical disk drive for reading and / or writing to / from a removable non-volatile optical disk can be provided. In this case, each drive can be connected to a bus (not shown) via one or more data media interfaces.
[0194] The communication unit 4740 communicates with another computing device via a communication medium. Additionally, the functions of the components in the computing device 4700 can be implemented by a single computing cluster or multiple computer machines, which can communicate via a communication connection. Therefore, the computing device 4700 can operate in a networked environment using a logical connection with one or more other servers, networked personal computers (PCs), or other general network nodes.
[0195] The input device 4750 can be one or more of various input devices, such as a mouse, a keyboard, a trackball, a voice input device, and so on. The output device 4760 can be one or more of various output devices, such as a display, a speaker, a printer, and so on. With the aid of the communication unit 4740, the computing device 4700 can also communicate with one or more external devices (not shown), such as a storage device and a display device, the computing device 4700 can also communicate with one or more devices that enable a user to interact with the computing device 4700, or if needed, the computing device 4700 can also communicate with any device (such as a network card, a modem, etc.) that enables the computing device 4700 to communicate with one or more other computing devices. Such communication can be carried out via an input / output (I / O) interface (not shown).
[0196] In some embodiments, some or all components of the computing device 4700 can also be arranged in a cloud computing architecture instead of being integrated in a single device. In a cloud computing architecture, the components can be provided remotely and work together to implement the functions described in the present disclosure. In some embodiments, cloud computing provides computing, software, data access, and storage services, which do not require the end user to know the physical location or configuration of the system or hardware providing these services. In various embodiments, cloud computing uses appropriate protocols to provide services via a wide area network, such as the Internet. For example, a cloud computing provider provides an application via a wide area network, and the application can be accessed through a web browser or any other computing component. The software or components of the cloud computing architecture and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be consolidated or distributed at the locations of remote data centers. The cloud computing infrastructure can provide services through a shared data center, although to the user, they appear as a single access point. Thus, the cloud computing architecture can be used to provide the components and functions described herein from a service provider at a remote location. Alternatively, the components and functions described herein can be provided by a conventional server or installed directly or otherwise on a client device.
[0197] In an embodiment of the present disclosure, the computing device 4700 can be used to implement video encoding / decoding. The memory 4720 can include one or more video codec modules 4725 having one or more program instructions. These modules are accessible and executable by the processing unit 4710 to perform the functions of various embodiments described herein.
[0198] In an example embodiment of performing video encoding, an input device 4750 may receive video data as an input 4770 to be encoded. The video data may be processed, for example, by a video codec module 4725 to generate an encoded bitstream. The encoded bitstream may be provided as an output 4780 via an output device 4760.
[0199] In an example embodiment of performing video decoding, an input device 4750 may receive the encoded bitstream as an input 4770. The encoded bitstream may be processed, for example, by a video codec module 4725 to generate decoded video data. The decoded video data may be provided as an output 4780 via an output device 4760.
[0200] Although the present disclosure has been specifically shown and described with reference to preferred embodiments thereof, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the present application as defined by the appended claims. These variations are intended to be covered by the scope of the present application. Accordingly, the foregoing description of the embodiments of the present application is not intended to be limiting.
Claims
1. A method for video processing, comprising: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit with respect to a conversion between the current video unit of a video and a bitstream of the video, wherein the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a slice of the video; and performing the conversion based on the target value.
2. The method according to claim 1, wherein the set of candidate values is predefined or indicated in the bitstream.
3. The method according to any one of claims 1 to 2, wherein each candidate value in the set of candidate values corresponds to a quantization parameter (QP) value.
4. The method according to claim 3, wherein the QP values in a set of QP values corresponding to the set of candidate values are sorted in ascending order.
5. The method according to any one of claims 1 to 4, wherein the set of candidate values includes candidate values corresponding to all QP values allowed to be used at a decoder.
6. The method according to any one of claims 1 to 4, wherein the set of candidate values includes candidate values corresponding to a subset of QP values among all QP values allowed to be used at a decoder.
7. The method according to claim 6, wherein the subset of QP values includes QP values between a first threshold and a second threshold.
8. The method according to claim 7, wherein the first threshold is indicated in the bitstream.
9. The method according to any one of claims 7 to 8, wherein the first threshold or the second threshold depends on at least one of the following: a picture type of the current video unit, a slice type of the current video unit, or a temporal layer of the current video unit.
10. The method according to claim 6, wherein each difference between two adjacent QP values in the subset of QP values is the same.
11. The method according to claim 7, wherein each difference is equal to a predetermined value.
12. The method according to any one of claims 5 to 11, wherein the all QP values include 64 QP values ranging from 0 to 63.
13. The method according to any one of claims 2 to 12, wherein the set of candidate values is indicated in a predictive manner in the bitstream.
14. The method according to any one of claims 1 to 13, wherein determining the target value comprises: selecting the target value from the set of candidate values based on one of the following: encoding / decoding information of the current video unit, encoding / decoding information of the slice, or encoding / decoding information of a picture including the current video unit.
15. The method according to claim 14, wherein the encoding / decoding information includes a QP value.
16. The method according to claim 15, wherein selecting the target value comprises: determining a first QP value based on one of the following: the QP value of the current video unit, the QP value of the slice, or the QP value of the picture; and Select the target value from the set of candidate values based on the first QP value and the mapping between the QP value and the set of candidate values.
17. The method according to claim 16, wherein the mapping between the QP value and the set of candidate values is predefined or indicated in the bitstream.
18. The method according to any one of claims 16 to 17, wherein the first QP value is equal to the QP value of the current video unit.
19. The method according to any one of claims 16 to 17, wherein determining the first QP value comprises: Adjust the QP value of the current video unit using a first adjustment value to obtain the first QP value.
20. The method according to claim 19, wherein the first adjustment value is an integer and greater than 1.
21. The method according to any one of claims 19 to 20, wherein the first adjustment value depends on at least one of the following: the picture type of the current video unit, the slice type of the current video unit, or the codec configuration for the current video unit.
22. The method according to any one of claims 16 to 21, wherein selecting the target value from the set of candidate values based on the first QP value and the mapping comprises: If it is determined that the first QP value is included in a set of QP values corresponding to the set of candidate values, determine the target value as the candidate value in the set of candidate values corresponding to the first QP value.
23. The method according to any one of claims 16 to 22, wherein selecting the target value from the set of candidate values based on the first QP value and the mapping comprises: If it is determined that the first QP value is not included in a set of QP values corresponding to the set of candidate values, select a second QP value from the set of QP values based on the first QP value; and determine the target value as the candidate value in the set of candidate values corresponding to the second QP value.
24. The method according to claim 23, wherein the second QP value is the QP value in the set of QP values that is closest to the first QP value.
25. The method according to claim 24, wherein the absolute difference between the first QP value and the second QP value is the smallest absolute difference among the absolute differences between the first value and each QP value in the set of QP values.
26. The method according to claim 23, wherein multiple QP values in the set of QP values are closest to the first QP value, and the second QP value is the smallest QP value among the multiple QP values.
27. The method according to claim 23, wherein multiple QP values in the set of QP values are closest to the first QP value, and the second QP value is the largest QP value among the multiple QP values.
28. The method according to claim 14, wherein the codec information includes a QP offset.
29. The method according to claim 28, wherein selecting the target value comprises: Determine a first QP offset based on one of the following: the QP offset of the current video unit, the QP offset of the slice, or The QP offset of the picture; and selecting the target value from the set of candidate values based on the first QP offset and a mapping between the QP offset and the set of candidate values.
30. The method according to claim 29, wherein the mapping between the QP offset and the set of candidate values is predefined or indicated in the bitstream.
31. The method according to any one of claims 29 to 30, wherein the first QP offset is equal to the QP offset of the current video unit.
32. The method according to any one of claims 29 to 30, wherein determining the first QP offset comprises: adjusting the QP offset of the current video unit by a second adjustment value to obtain the first QP offset.
33. The method according to claim 32, wherein the second adjustment value is an integer and greater than 1.
34. The method according to any one of claims 32 to 33, wherein the second adjustment value depends on at least one of the following: the picture type of the current video unit, the slice type of the current video unit, or the codec configuration for the current video unit.
35. The method according to any one of claims 29 to 34, wherein selecting the target value from the set of candidate values based on the first QP offset and the mapping comprises: if it is determined that the first QP offset is included in a set of QP offsets corresponding to the set of candidate values, determining the target value as the candidate value in the set of candidate values corresponding to the first QP offset.
36. The method according to any one of claims 29 to 35, wherein selecting the target value from the set of candidate values based on the first QP offset and the mapping comprises: if it is determined that the first QP offset is not included in a set of QP offsets corresponding to the set of candidate values, selecting a second QP offset from the set of QP offsets based on the first QP offset; and determining the target value as the candidate value in the set of candidate values corresponding to the second QP offset.
37. The method according to claim 36, wherein the second QP offset is the QP offset in the set of QP offsets that is closest to the first QP offset.
38. The method according to claim 37, wherein the absolute difference between the first QP offset and the second QP offset is the smallest absolute difference among the absolute differences between the first value and each QP offset in the set of QP offsets.
39. The method according to claim 36, wherein multiple QP offsets in the set of QP offsets are closest to the first QP offset, and the second QP offset is the smallest QP offset among the multiple QP offsets.
40. The method according to claim 36, wherein multiple QP offsets in the set of QP offsets are closest to the first QP offset, and the second QP offset is the largest QP offset among the multiple QP offsets.
41. The method according to any one of claims 1 to 40, wherein the current video unit comprises one of the following: a slice, a coding tree unit (CTU), a CTU row, One or more groups of CTUs, Coding and decoding units (CUs), Prediction units (PUs), Transformation units (TUs), Coding and decoding tree blocks (CTBs), Coding and decoding blocks (CBs), Prediction blocks (PBs), Transformation blocks (TBs), Blocks, Sub-blocks of the said block, or Sub-regions within the said block.
42. The method according to any one of claims 1 to 41, wherein at least a part of the said group of candidate values does not exist in the bitstream and is determined at the decoder.
43. The method according to any one of claims 1 to 42, wherein a group of candidate values for the said parameter used in the diversity reordering for inter prediction tools is different from a group of candidate values for the said parameter used in the diversity reordering for intra block copy (IBC).
44. The method according to any one of claims 1 to 43, wherein the said parameter is represented by lambda (λ).
45. The method according to any one of claims 1 to 44, wherein whether to apply the said method and / or how to apply the said method is indicated at one of the following: Sequence level, Group of pictures level, Picture level, Slice level, or Slice group level.
46. The method according to any one of claims 1 to 45, wherein whether to apply the said method and / or how to apply the said method is indicated in one of the following: Sequence header, Picture header, Sequence parameter set (SPS), Video parameter set (VPS), Dependency parameter set (DPS), Decoding capability information (DCI), Picture parameter set (PPS), Adaptive parameter set (APS), Slice header, or Slice group header.
47. The method according to any one of claims 1 to 44, wherein whether to apply the said method and / or how to apply the said method is indicated at a region including more than one sample or pixel.
48. The method according to claim 47, wherein the said region includes at least one of the following: Prediction block (PB), Transformation block (TB), Coding and decoding block (CB), Prediction unit (PU), Transformation unit (TU), Coding and decoding unit (CU), Virtual pipeline data unit (VPDU), Coding and decoding tree unit (CTU), CTU row, Slice, Picture, or Sub-picture.
49. The method according to any one of claims 1 to 49, wherein whether to apply the said method and / or how to apply the said method depends on the coded and decoded information.
50. The method according to claim 49, wherein the said coded and decoded information includes at least one of the following: Block size, Color format, Single / double tree segmentation, Double tree segmentation, Color component, Slice type, or Picture type.
51. The method according to any one of claims 1 to 50, wherein the said conversion includes encoding the current video unit into the bitstream.
52. The method according to any one of claims 1 to 50, wherein the said conversion includes decoding the current video unit from the bitstream.
53. A device for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to execute the method according to any one of claims 1 to 52.
54. A non-transitory computer-readable storage medium storing instructions that cause a processor to execute the method according to any one of claims 1 to 52.
55. A non-transitory computer-readable recording medium storing a bitstream of a video, the bitstream being generated by a method executed by a device for video processing, wherein the method comprises: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of the video, wherein the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a slice of the video; and generating the bitstream based on the target value.
56. A method for storing a bitstream of a video, comprising: determining a target value of a parameter from a set of candidate values of the parameter for sorting a plurality of motion candidates of a current video unit of the video, wherein the parameter is associated with a cost difference related to the plurality of motion candidates, and the current video unit is part of a slice of the video; generating the bitstream based on the target value; and storing the bitstream in a non-transitory computer-readable recording medium.