Method and device for video processing and medium
By refining the second motion information using the first motion information during the refinement process of video units, the problem of improving encoding and decoding efficiency in the existing technology is solved, and more efficient video encoding and decoding is achieved.
Patent Information
- Application Number
- CN202480038077.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-06-07
- Filing Date
- 2024-06-05
- Publication Date
- 2026-01-02
AI Technical Summary
Existing video encoding and decoding technologies have room for improvement in encoding and decoding efficiency, especially in template matching and video compression, where further improvements are difficult.
By refining the second motion information using the first motion information during the video unit refinement process, and then performing conversion based on the refined motion information, the encoding and decoding performance is improved.
It improves the encoding and decoding performance of video codecs and enhances the efficiency of video processing.
Smart Images

Figure CN121264047A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure generally relate to video processing technology, and more particularly, to improvements on template matching. BACKGROUND
[0002] Nowadays, digital video capability is being applied to various aspects of people's life. For video coding / decoding, various types of video compression technologies have been proposed, such as MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 Advanced Video Coding (AVC), ITU-T H.265 High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard. However, it is generally desired to further improve the coding efficiency of video coding technology. SUMMARY
[0003] Embodiments of the present disclosure provide a solution for video processing.
[0004] In a first aspect, a method for video processing is proposed. The method comprises: obtaining first motion information and second motion information of a video unit of a video for conversion between the video unit and a bitstream of the video unit; refining the second motion information by using the first motion information during a refinement process of the video unit, wherein the refinement process is refinement or iterative refinement; and performing the conversion based on the refined first motion information and the second motion information. In this way, it can improve the coding performance.
[0005] In a second aspect, an apparatus for video processing is proposed. The apparatus comprises a processor and a non-transitory memory having instructions thereon. The instructions, when executed by the processor, cause the processor to perform the method according to the first aspect of the present disclosure.
[0006] In a third aspect, a non-transitory computer-readable storage medium is proposed. The non-transitory computer-readable storage medium stores instructions that cause a processor to perform the method according to the first aspect of the present disclosure.
[0007] In a fourth aspect, another non-transitory computer-readable recording medium is proposed. The non-transitory computer-readable recording medium stores a bitstream of a video generated by a method executed by an apparatus for video processing. The method comprises: obtaining first motion information and second motion information of a video unit of the video; refining the second motion information by using the first motion information during a refinement process of the video unit, wherein the refinement process is refinement or iterative refinement; and generating the bitstream based on the refined first motion information and the second motion information.
[0008] In a fifth aspect, a method for storing a bitstream of a video is suggested. The method includes obtaining first motion information and second motion information of a video unit of the video; refining the second motion information by using the first motion information during a refinement process of the video unit, wherein the refinement process is a refinement or an iterative refinement; generating the bitstream based on the refined first motion information and the second motion information; and storing the bitstream in a non-transitory computer-readable recording medium.
[0009] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which like reference characters refer to the like elements throughout. In the example embodiments of the present disclosure, like reference numerals refer to like elements throughout.
[0011] Figure 1 A block diagram showing an example video coding system is shown in accordance with some embodiments of the present disclosure; Figure 2 A block diagram showing a first example video encoder is shown in accordance with some embodiments of the present disclosure; Figure 3 A block diagram showing an example video decoder is shown in accordance with some embodiments of the present disclosure; Figure 4 An example of an encoder block diagram of VVC is shown; Figure 5 67 intra prediction modes are shown; Figure 6A And Figure 6B Reference samples for wide-angle intra prediction are shown; Figure 7 Discontinuity issues in cases where the direction exceeds 45° are shown; Figure 8A And Figure 8B MMVD search points are shown; Figure 9 An illustration for symmetric MVD mode is shown; Figure 10 Extended CU regions used in BDOF are shown; Figure 11 Top and left neighboring blocks used in CIIP weight derivation are shown; Figure 12A And Figure 12B Control point based affine motion model is shown; Figure 13 Affine MVFs for each sub-block are shown; Figure 14 Location of inherited affine motion predictor value is shown; Figure 15 Control point motion vector inheritance is shown; Figure 16 Location of candidate positions for constructed affine Merge mode is shown; Figure 17 Illustration of motion vector usage for proposed combination method is shown; Figure 18 Sub-block MV VSB and pixel delta v(i,j) (red arrows) are shown; Figure 19A and Figure 19B SbTMVP process in VVC is shown; Figure 20 Local illumination compensation is shown; Figure 21 No downsampling for short side is shown; Figure 22 Decoder-side motion vector refinement is shown; Figure 23 Diamond-shaped region in search region is shown; Figure 24 Location of spatial Merge candidate is shown; Figure 25 Pair of candidates considered for redundancy check for spatial Merge candidate is shown; Figure 26 Illustration of motion vector scaling for temporal Merge candidate is shown; Figure 27 Candidate positions, C0 and C1, for temporal Merge candidate are shown; Figure 28 VVC spatial neighboring blocks of current block are shown; Figure 29 Illustration of virtual blocks in i-th round of search is shown; Figure 30 Example of GPM partitioning grouped at the same angle is shown; Figure 31 Unidirectional prediction MV selection for geometric partition mode is shown; Figure 32 Example generation of hybrid weights using geometric partition mode is shown; Figure 33 Spatial neighboring blocks used to derive spatial Merge candidate are shown; Figure 34 Template matching execution on search region around initial MV is shown; Figure 35 Diagram showing sub-blocks for OBMC application is shown; Figure 36 SBT position, type and transform type is shown; Figure 37 Neighboring samples used to compute SAD is shown; Figure 38 Neighboring samples used to compute SAD for sub-CU level motion information is shown; Figure 39 Ordering process is shown; Figure 40 Reordering process in encoder is shown; Figure 41 Reordering process in decoder is shown; Figure 42 First and second HPTs are shown; Figure 43A And Figure 43B Spatial neighbors for derivation of affine Merge / AMVP candidates is shown; Figure 44 Diagram showing constructed affine Merge / AMVP candidates from non-adjacent neighbors to first type is shown; Figure 45 Different number of search points in same search mode is shown; Figure 46 Different number of search points in different search modes is shown; Figure 47 A flowchart of a method for video processing according to an embodiment of the present disclosure is shown; and Figure 48 A block diagram of a computing device in which various embodiments of the present disclosure can be implemented is shown.
[0012] Throughout the drawings, identical or similar reference numerals can designate identical or similar elements throughout the several views. DETAILED DESCRIPTION
[0013] The principles of the present disclosure will now be described with reference to some embodiments. It should be understood that the embodiments are described for illustrative purposes only and help the person skilled in the art to understand and implement the present disclosure, but do not imply any limitation on the scope of the present disclosure. The disclosure described herein can be implemented in various ways in addition to the ways described below.
[0014] In the following description and claims, unless otherwise defined, all scientific and technical terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs.
[0015] References in the present disclosure to “one embodiment,” “an embodiment,” “example embodiments,” etc., indicate that the embodiment described can include a particular feature, structure, or characteristic, but every embodiment can not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases are not necessarily referring to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with an example embodiment, it is submitted that
[0016] It should be understood that although the terms “first” and “second” etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element could be termed a second element, and, similarly, a second element could be termed a first element without departing from the scope of the example embodiments. As used herein, the term “and / or” includes any and all combinations of one or more of the associated terms.
[0017] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of example embodiments. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes” and / or “including,” when used herein, specify the presence of stated features, elements and / or components, but do not preclude the presence or addition of one or more other features, elements, components and / or combinations thereof.
[0018] Example Environment Figure 1 is a block diagram illustrating an example video coding system 100 that can utilize the techniques of this disclosure. As shown, video coding system 100 can include a source device 110 and a destination device 120. Source device 110 can also be referred to as a video encoding device, and destination device 120 can also be referred to as a video decoding device. In operation, source device 110 can be configured to generate encoded video data, and destination device 120 can be configured to decode the encoded video data generated by source device 110. Source device 110 can include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.
[0019] Video source 112 can include a source such as a video capture device. Examples of video capture devices include, but are not limited to, an interface to receive video data from a video content provider, a computer graphics system to generate video data, and / or a combination thereof.
[0020] Video data can comprise one or more pictures. Video encoder 114 encodes video data from video source 112 to generate a bitstream. The bitstream can include a sequence of bits that forms an encoded representation of the video data. The bitstream can include encoded pictures and associated data. An encoded picture is an encoded representation of a picture. The associated data can include sequence parameter sets, picture parameter sets, and other syntax structures. I / O interface 116 can include a modulator / demodulator and / or a transmitter. The encoded video data can be transmitted directly to destination device 120 by network 130A via I / O interface 116. The encoded video data can also be stored onto a storage medium / server 130B for access by destination device 120.
[0021] Destination device 120 can include an I / O interface 126, a video decoder 124, and a display device 122. I / O interface 126 can include a receiver and / or a modem. I / O interface 126 can acquire encoded video data from source device 110 or storage medium / server 130B. Video decoder 124 can decode the encoded video data. Display device 122 can display the decoded video data to a user. Display device 122 can be integrated with destination device 120, or can be external to destination device 120 which is configured to interface with an external display device.
[0022] Video encoder 114 and video decoder 124 can operate according to a video compression standard, such as the High Efficiency Video Coding (HEVC) standard, Versatile Video Coding (VVC) standard, and other existing and / or future standards.
[0023] Figure 2 is a block diagram illustrating an example of a video encoder 200 that can be Figure 1 an example of video encoder 114 in system 100 shown.
[0024] Video encoder 200 can be configured to implement any or all of the techniques of this disclosure. In Figure 2 example, video encoder 200 includes a plurality of functional components. The techniques described in this disclosure can be shared amongst the functional components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0025] In some embodiments, video encoder 200 can include partition unit 201, prediction unit 202, which can include mode select unit 203, motion estimation unit 204, motion compensation unit 205, and intra-prediction unit 206, residual generation unit 207, transform unit 208, quantization unit 209, inverse quantization unit 210, inverse transform unit 211, reconstruction unit 212, buffer 213, and entropy encoding unit 214.
[0026] In other examples, video encoder 200 can include more, less, or different functional components. In one example, prediction unit 202 can include an intra block copy (IBC) unit. The IBC unit can perform prediction in an IBC mode in which at least one reference picture is the picture in which the current video block is located.
[0027] Furthermore, although some components, such as motion estimation unit 204 and motion compensation unit 205, can be integrated, for purposes of explanation, these components are shown separately in Figure 2 examples.
[0028] Partition unit 201 can partition a picture into one or more video blocks. Video encoder 200 and video decoder 300 can support various video block sizes.
[0029] Mode select unit 203 can select one of a plurality of encoding modes (intra- or inter- coding) based on, for example, error results, and provide the resulting intra- or inter- coded block to residual generation unit 207 to generate residual block data and to reconstruction unit 212 to reconstruct the encoded block for use as a reference picture. In some examples, mode select unit 203 can select a combined inter-intra prediction (CIIP) mode in which prediction is based on both inter- and intra-prediction signals. In the case of inter-prediction, mode select unit 203 can also select a resolution for motion vectors (e.g., sub-pixel precision or integer pixel precision) for the block.
[0030] To perform inter-prediction for a current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 to the current video block. Motion compensation unit 205 can determine a predicted video block for the current video block based on the motion information and decoded samples of pictures from buffer 213 other than the picture in which the current video block is located.
[0031] Motion estimation unit 204 and motion compensation unit 205 can perform different operations on a current video block, e.g., depending on whether the current video block is in an I slice, a P slice, or a B slice. As used herein, an "I slice" can refer to a portion of a picture composed of macroblocks all of which are based on macroblocks within the same picture. Further, as used herein, a "P slice" and a "B slice" can refer, in some aspects, to portions of a picture composed of macroblocks that are independent of macroblocks in the same picture.
[0032] In some examples, motion estimation unit 204 can perform uni-prediction on a current video block, and motion estimation unit 204 can search a reference picture in List 0 or List 1 for a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference picture in List 0 or List 1 containing the reference video block, and a motion vector indicating a spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, a prediction direction indicator, and the motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information for the current video block.
[0033] Alternatively, in other examples, motion estimation unit 204 can perform bi-prediction on a current video block. Motion estimation unit 204 can search a reference picture in List 0 for one reference video block for the current video block, and can also search a reference picture in List 1 for another reference video block for the current video block. Motion estimation unit 204 can then generate multiple reference indices indicating multiple reference pictures in List 0 and List 1 containing the multiple reference video blocks, and multiple motion vectors indicating multiple spatial displacements between the multiple reference video blocks and the current video block. Motion estimation unit 204 can output the multiple reference indices and the multiple motion vectors for the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the multiple reference video blocks indicated by the motion information for the current video block.
[0034] In some examples, motion estimation unit 204 can output a full set of motion information for use in decoding processing by a decoder. Alternatively, in some embodiments, motion estimation unit 204 can reference motion information of another video block to signal motion information of the current video block. For example, motion estimation unit 204 can determine that the motion information of the current video block is sufficiently similar to the motion information of a neighboring video block.
[0035] In one example, the motion estimation unit 204 can indicate a value in a syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.
[0036] In another example, the motion estimation unit 204 can identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates a difference between a motion vector of the current video block and a motion vector of the indicated video block. The video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.
[0037] As discussed above, the video encoder 200 can signal motion vectors in a predictive manner. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and Merge mode signaling.
[0038] The intra prediction unit 206 can perform intra prediction on the current video block. When the intra prediction unit 206 performs intra prediction on the current video block, the intra prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block can include a predicted video block and various syntax elements.
[0039] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the prediction video block(s) for the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of samples in the current video block.
[0040] In other examples, such as in skip mode, there can be no residual data for the current video block for the current video block, and the residual generation unit 207 can not perform the subtraction operation.
[0041] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.
[0042] After the transform processing unit 208 generates the transform coefficient video blocks associated with the current video block, the quantization unit 209 can quantize the transform coefficient video blocks associated with the current video block based on one or more quantization parameter (QP) values associated with the current video block.
[0043] The inverse quantization unit 210 and the inverse transform unit 211 can apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 can add the reconstructed residual video block to corresponding samples from one or more prediction video blocks generated by the prediction unit 202 to produce a reconstructed video block associated with the current video block for storage in the buffer 213.
[0044] After the reconstruction unit 212 reconstructs the video block, an in-loop filtering operation can be performed to reduce video block effect artifacts in the video block.
[0045] The entropy encoding unit 214 can receive data from other functional components of the video encoder 200. When the entropy encoding unit 214 receives data, the entropy encoding unit 214 can perform one or more entropy encoding operations to generate entropy encoded data and output a bitstream that includes the entropy encoded data.
[0046] Figure 3 is a block diagram illustrating an example of a video decoder 300 in accordance with some embodiments of the disclosure, which can be an example of video decoder 124 in system 100 shown Figure 1 is an example of video decoder 124 in system 100 shown.
[0047] The video decoder 300 can be configured to perform any or all of the techniques of the disclosure. In Figure 3 example, the video decoder 300 includes a plurality of functional components. The techniques described in this disclosure can be shared amongst the various components of the video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.
[0048] In Figure 3 example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, and a reconstruction unit 306 and a buffer 307. In some examples, the video decoder 300 can perform a decoding process generally reciprocal to the encoding process described with respect to the video encoder 200.
[0049] Entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream can include entropy encoded video data (e.g., encoded blocks of video data). Entropy decoding unit 301 can decode the entropy encoded video data, and motion compensation unit 302 can determine motion information from the entropy decoded video data, including motion vectors, motion vector precision, reference picture list index, and other motion information. Motion compensation unit 302 can determine such information, for example, by performing AMVP and Merge modes. AMVP is used, including deriving a number of most probable candidates based on data from neighboring PBs and reference pictures. The motion information typically includes a horizontal motion vector displacement value and a vertical motion vector displacement value, one or two reference picture indices, and in the case of a prediction region in a B slice, an identification of which reference picture list is associated with each index. As used herein, in some aspects, "Merge mode" can refer to deriving motion information from a spatially or temporally neighboring block.
[0050] Motion compensation unit 302 can generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier for the interpolation filter used at sub-pixel precision can be included in the syntax elements.
[0051] Motion compensation unit 302 can use the interpolation filter used by video encoder 200 during encoding of the video block to calculate interpolated values for sub-integer pixels of the reference block. Motion compensation unit 302 can determine the interpolation filter used by video encoder 200 from the received syntax information, and motion compensation unit 302 can use the interpolation filter to generate the prediction block.
[0052] Motion compensation unit 302 can use at least some of the syntax information to determine the size of blocks used to encode frames and / or slices of the encoded video sequence, partitioning information describing how each macroblock of a picture of the encoded video sequence is partitioned, modes indicating how each partition is encoded, one or more reference frames (and reference frame lists) for each inter-coded block, and other information used to decode the encoded video sequence. As used herein, in some aspects, a "slice" can refer to a data structure that can be decoded independently of other slices of the same picture in terms of entropy coding, signal prediction, and residual signal reconstruction. A slice can be an entire picture, or can also be a region of a picture.
[0053] Intra prediction unit 303 can use intra prediction modes, e.g., received in the bitstream, to form a prediction block from spatial neighboring blocks. Dequantization unit 304 dequantizes quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.
[0054] The reconstruction unit 306 can obtain the decoded block, e.g., by adding the residual block to the corresponding prediction block generated by the motion compensation unit 302 or the intra prediction unit 303. If needed, a deblocking filter can also be applied to filter the decoded block in order to remove blocking artifacts. The decoded video blocks are then stored in the buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction, and which also produces decoded video for presentation on a display device.
[0055] Some example embodiments of the present disclosure will be described in detail below. It should be noted that the use of section headings in this document is for convenience only and not to be construed as limiting the embodiments disclosed in that section to that section only. Furthermore, although some embodiments are described in detail with reference to the versatile video coding or other specific video codec, the disclosed techniques are also applicable to other video coding technologies. Moreover, although some embodiments describe video encoding steps in detail, it should be understood that the corresponding decoding steps of the de- encoded will be implemented by a decoder. Furthermore, the term video processing includes video encoding or compression, video decoding or decompression, and video transcoding, in which video pixels are represented from one compressed format to another or at different compressed bit rates.
[0056] 1. Brief overview The present disclosure relates to video coding technologies. In particular, it relates to template matching for inter prediction, and how to apply it with other coding tools in image / video coding. It can be applied to existing video coding standards like HEVC or the versatile video coding (VVC). It can also be applicable to future video coding standards or video codecs.
[0057] 2. Introduction Video coding standards have evolved mainly through the development of the well-known ITU-T and ISO / IEC standards. The ITU-T produced H.261 and H.263 standards, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Coding (AVC) and H.265 / HEVC standards. From H.262, video coding standards are based on the hybrid video coding structure, where temporal prediction is combined with transform coding. To explore future video coding technologies beyond HEVC, the Joint Video Exploration Team (JVET) was founded by VCEG and MPEG jointly in 2015. Since then, many new methods have been adopted by JVET and put into the reference software named Joint Exploration Model (JEM). In April 2018, the Joint Video Team (JVT) between VCEG (Q6 / 16) and ISO / IEC JTC1 SC29 / WG11 (MPEG) was established, which is committed to the VVC standard, aiming to reduce 50% bit rate compared with HEVC.
[0058] 2.1. Coding process of a typical video codec Figure 4 An example of the encoder block diagram of VVC is shown, which contains three in-loop filters: Deblocking Filter (DF), Sample Adaptive Offset (SAO), and ALF. Unlike DF, which uses a pre-defined filter, SAO and ALF utilize the original samples of the current picture, respectively, by adding an offset and by applying a Finite Impulse Response (FIR) filter, and utilize the coded side information to signal the offset and filter coefficients to reduce the mean square error between the original samples and the reconstructed samples. ALF is located at the last processing stage of each picture and can be considered as a tool that tries to capture and fix artifacts caused by previous stages.
[0059] 2.2. Intra mode coding with 67 intra prediction modes To capture arbitrary edge directions presented in natural videos, as shown in Figure 5 The number of directional intra modes is extended from 33 used in HEVC to 65, and the planar mode and DC mode remain unchanged. These denser directional intra prediction modes are applied to all block sizes and to both luma intra prediction and chroma intra prediction.
[0060] In HEVC, each intra-coded block has a square shape and each side of it has a length that is a power of 2. Therefore, no partitioning operation is needed to generate an intra prediction value using the DC mode. In VVC, a block can have a rectangular shape, which in general case requires a partitioning operation to be used for each block. To avoid the partitioning operation for DC prediction, only the longer side is used to calculate the average value for non-square blocks.
[0061] 2.2.1. Wide-angle intra prediction Although 67 modes are defined in VVC, the exact prediction direction for a given intra prediction mode index also depends on the block shape. The regular angular intra prediction directions are defined as clockwise directions from 45 degrees to -135 degrees. In VVC, for non-square blocks, several regular angular intra prediction modes are adaptively replaced by wide-angle intra prediction modes. The replaced modes are signaled using the original mode index, which is remapped to the index of the wide-angle mode after parsing. The total number of intra prediction modes remains unchanged, i.e., 67, and the intra mode coding method remains unchanged.
[0062] To support these prediction directions, a top reference of length 2W+1 and a left reference of length 2H+1 are defined, as shown in Figure 6A and Figure 6B
[0063] The number of replaced modes in the wide-angle direction modes depends on the aspect ratio of the block. The replaced intra prediction modes are shown in Table 2-1.
[0064]
[0065] As shown in Figure 7 In the case of wide-angle intra prediction, two non-adjacent reference samples can be used for two vertically adjacent predicted samples. Therefore, a low-pass reference sample filter and edge smoothing are applied to wide-angle prediction to reduce the negative impact of the increased gap pα. If the wide-angle mode represents a non-fraction offset. There are 8 modes in the wide-angle mode that satisfy this condition, 8 modes are [-14, -12, -10, -6, 72, 76, 78, 80]. When a block is predicted by these modes, the samples in the reference buffer are directly copied without applying any interpolation. With this modification, the number of samples that need to be smoothed is reduced. In addition, it aligns the design of non-fraction modes in regular prediction modes with the wide-angle modes.
[0066] In VVC, 4:2:2 and 4:4:4 chroma formats are supported in addition to 4:2:0. The chroma derivation mode (DM) derivation table for 4:2:2 chroma format is initially ported from HEVC, extending the number of entries from 35 to 67 to align with the extension of intra prediction modes. Since HEVC specification does not support prediction angles lower than -135 degrees and higher than 45 degrees, the range of luma intra prediction modes from 2 to 5 is mapped to 2. Therefore, the chroma DM derivation table for 4:2:2 chroma format is updated by replacing some values of the mapping table to convert the prediction angles more accurately for chroma blocks.
[0067] 2.3. Inter prediction For each inter predicted CU, the motion parameters consist of motion vectors, reference picture indices and reference picture list usage indices, and additional information required for the inter predicted samples generation using new coding features of VVC. The motion parameters can be signaled in an explicit or implicit manner. When a CU is coded in skip mode, the CU is associated with one PU and does not have significant residual coefficients, no coded motion vector difference or reference picture index. Merge mode is specified, where the motion parameters for the current CU are obtained from neighboring CUs, including spatial candidates and temporal candidates, and additional scheduling introduced in VVC. Merge mode can be applied to any inter predicted CU, not only for skip mode. An alternative to merge mode is the explicit transmission of motion parameters, where the motion vectors for each reference picture list, the corresponding reference picture indices and reference picture list usage flags, and other required information are explicitly signaled for each CU.
[0068] 2.4. Intra block copy (IBC) Intra block copy (IBC) is a tool adopted in the HEVC extension on SCC. It is well known that it significantly improves the coding efficiency of screen content material. Since IBC mode is implemented as a block-level coding mode, block matching (BM) is performed at the encoder to find the best block vector (or motion vector) for each CU. Here, the block vector is used to indicate the displacement from the current block to a reference block that has already been reconstructed inside the current picture. The luma block vector of an IBC coded CU is in integer precision. The chroma block vector is also rounded to integer precision. When combined with AMVR, the IBC mode can switch between 1-pixel motion vector precision and 4-pixel motion vector precision. An IBC coded CU is considered as a third prediction mode different from intra or inter prediction modes. IBC mode is applicable to CUs with width and height less than or equal to 64 luma samples.
[0069] At the encoder side, hash-based motion estimation is performed for IBC. The encoder performs RD check for blocks whose width or height is not larger than 16 luma samples. For non-Merge mode, first a block vector search is performed using hash-based search. If the hash search does not return a valid candidate, a local search based on block matching will be performed.
[0070] In hash-based search, the hash key match (32-bit CRC) between the current block and the reference block is extended to all allowed block sizes. The hash key computation is based on 4x4 sub-blocks for each location in the current picture. For larger size current blocks, the hash key is determined to match the hash key of a reference block when all hash keys of all 4x4 sub-blocks match the hash keys in the corresponding reference locations. If multiple reference blocks' hash keys are found to match the hash key of the current block, the block vector cost of each matching reference is computed and the one with the smallest cost is selected.
[0071] In block matching search, the search range is set to cover both the previous CTU and the current CTU.
[0072] At CU level, IBC mode is signaled with a flag and it can be signaled as IBC AMVP mode or IBC Skip / Merge mode as follows: - IBC Skip / Merge mode: Merge candidate index is used to indicate which block vector from the list of IBC coded blocks from neighboring candidates is used to predict the current block. The Merge list consists of spatial candidates, HMVP candidates and paired candidates.
[0073] - IBC AMVP mode: Block vector difference is coded in the same way as motion vector difference. Block vector prediction method uses two candidates as the predictor, one from left neighbor and one from above neighbor (if IBC coded). When either neighbor is not available, the default block vector will be used as the predictor. A flag is signaled to indicate the block vector predictor index.
[0074] 2.5. Merge mode with MVD (MMVD) In addition to the Merge mode (where the implicitly derived motion information is directly used for the prediction sample generation of the current CU), Merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the regular Merge flag to specify whether MMVD mode is used for the CU.
[0075] In MMVD, after selecting the Merge candidate, the Merge candidate is further refined by MVD information signaled. The further information includes a Merge candidate flag, an index specifying the motion magnitude, and an index for the indication of the motion direction. In MMVD mode, one of the first two candidates in the Merge list is selected to be used as the MV basis. A MMVD candidate flag is signaled to specify which of the first Merge candidate and the second Merge candidate is used.
[0076] The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. As shown in Figure 8A and Figure 8B The offset is added to the horizontal component or the vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 2.
[0077]
[0078] The direction index represents the direction of the MVD relative to the starting point. The direction index can represent four directions as shown in Table 3. It is noted that the meaning of the sign of the MVD can change depending on the information of the starting MV. When the starting MV is a uni-prediction MV or bi-prediction MV and both lists point to the same side of the current picture (i.e., both reference POCs are greater than the POC of the current picture or both are less than the POC of the current picture), the sign in Table 3 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bi-prediction MV and the two MVs point to different sides of the current picture (i.e., one reference POC is greater than the POC of the current picture while the other reference POC is less than the POC of the current picture), and the difference of the POCs in list 0 is greater than the difference of the POCs in list 1, the sign in Table 3 specifies the sign of the MV offset added to the list 0 MV component of the starting MV and the sign of the list 1 MV has the opposite value. Otherwise, if the difference of the POCs in list 1 is greater than the difference of the POCs in list 0, the sign in Table 3 specifies the sign of the MV offset added to the list 1 MV component of the starting MV and the sign of the list 0 MV has the opposite value.
[0079] The MVD is scaled according to the difference of the POCs in each direction. If the difference of the POCs in the two lists is the same, no scaling is needed. Otherwise, if the difference of the POCs in list 0 is greater than the difference of the POCs in list 1, the MVD of list 1 is scaled as shown in Figure 26 If the difference of the POCs in list 1 is greater than the difference of the POCs in list 0, the MVD of list 0 is scaled in the same way. If the starting MV is uni-predicted, the MVD is added to the available MV.
[0080]
[0081] 2.6. Symmetric MVD coding In VVC, in addition to normal uni-prediction and bi-prediction mode MVD signaling, a symmetric MVD mode for bi-prediction MVD signaling is applied. In symmetric MVD mode, the motion information including the reference picture indices of both List 0 and List 1 and the MVD of List 1 is not signaled but derived.
[0082] The decoding process of symmetric MVD mode is as follows: 1. At slice level, the variables BiDirPredFlag, RefIdxSymL0 and RefIdxSymL1 are derived as follows: - If mvd_l1_zero_flag is equal to 1, BiDirPredFlag is set equal to 0.
[0083] - Otherwise, if the closest reference picture in List 0 and the closest reference picture in List 1 form a pair of forward and backward reference pictures or a pair of backward and forward reference pictures, BiDirPredFlag is set to 1 and both List 0 and List 1 reference pictures are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0.
[0084] 2. At CU level, if the CU is bi-predictively coded and BiDirPredFlag is equal to 1, a symmetric mode flag explicitly signaling whether the symmetric mode is used or not is signaled.
[0085] When the symmetric mode flag is true, only mvp_l0_flag, mvp_l1_flag and MVD0 are explicitly signaled. The reference indices of List 0 and List 1 are set equal to the pair of reference pictures, respectively. MVD1 is set equal to (-MVD0). The final motion vector is as shown in the following equation.
[0086] (2-1) Figure 9 A diagram is shown for symmetric MVD mode.
[0087] In the encoder, symmetric MVD motion estimation starts from the initial MV evaluation. A set of initial MV candidates, including the MVs obtained from uni-prediction search, the MVs obtained from bi-prediction search and the MVs from AMVP list. One MV with the lowest rate-distortion cost is selected as the initial MV for symmetric MVD motion search.
[0088] 2.7. Bidirectional optical flow (BDOF) The bidirectional optical flow (BDOF) tool is included in VVC. BDOF, previously known as BIO, was included in JEM. Compared to the JEM version, the BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the multiplier size.
[0089] BDOF is used to refine the bi-predicted signal of a CU at the 4x4 subblock level. BDOF is applied to a CU if the CU satisfies all the following conditions: - The CU is coded using the "true" bi-prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order.
[0090] - The distances (i.e., POC difference) from the two reference pictures to the current picture are the same.
[0091] - Both reference pictures are short-term reference pictures.
[0092] - The CU is not coded using the affine mode or the SbTMVP Merge mode.
[0093] - The CU has more than 64 luma samples.
[0094] - Both CU height and CU width are greater than or equal to 8 luma samples.
[0095] - The BCW weight index indicates equal weights.
[0096] - WP is not enabled for the current CU.
[0097] - CIIP mode is not used for the current CU.
[0098] BDOF is only applied to the luma component. As the name indicates, the BDOF mode is based on the optical flow concept which assumes that the motion of an object is smooth. For each 4x4 subblock, a motion refinement is computed by minimizing the difference between the L0 predicted samples and the L1 predicted samples. The motion refinement is then used to adjust the bi-predicted sample values in the 4x4 subblock. The following steps are applied in the BDOF process.
[0099] First, the horizontal and vertical gradients of the two prediction signals and , are computed by directly computing the difference between two neighboring samples, i.e.,
[0100] where is the list of coordinates of the prediction signal in , and shift1 is calculated based on the luma bit depth bitDepth as shift1 = max( 6, bitDepth-6).
[0101] Then, the auto-correlation and cross-correlation of the gradients , , , and are calculated as
[0102] where
[0103] where is a 6x6 window around the 4x4 sub-block, and and are set to equal min( 1, bitDepth 11 ) and min( 4, bitDepth 8 ), respectively.
[0104] Motion refinement is then derived from the cross-correlation term and the auto-correlation term using the following equation:
[0105] where , , . is a floor function, and .
[0106] Based on the motion refinement and the gradients, the following adjustment is calculated for each sample in the 4x4 sub-block:
[0107] Finally, the BDOF samples of the CU are calculated by adjusting the bi-prediction samples as follows:
[0108] These values are chosen such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process is kept within 32 bits.
[0109] To derive the gradient values, it is necessary to generate the list ( Some predicted samples in ) .like Figure 10 As shown, BDOF in VVC uses an extended row / column around the CU boundary. To control the computational complexity of generating prediction samples outside the boundary, prediction samples in the extended region (white area) are generated by directly taking reference samples at nearby integer positions without interpolation (using the floor() operation on the coordinates), and a normal 8-tap motion-compensated interpolation filter is used to generate prediction samples inside the CU (gray area). These extended sample values are only used in gradient calculations. For the remaining steps in the BDOF process, if any samples and gradient values outside the CU boundary are needed, the samples and gradient values are padded with their nearest neighbors (i.e., repeated).
[0110] When the width and / or height of a CU is greater than 16 luminance samples, it will be divided into sub-blocks with a width and / or height equal to 16 luminance samples, and the sub-block boundaries will be considered as CU boundaries in the BDOF process. The maximum cell size for the BDOF process is limited to 16×16. The BDOF process can be skipped for each sub-block. The BDOF process is not applied to the sub-block when the SAD between the initial L0 and L1 predicted samples is less than a threshold. The threshold is set to equal to (8 * W * ( H >> 1 ), where W indicates the sub-block width and H indicates the sub-block height. To avoid the additional complexity of SAD calculation, the SAD between the initial L0 and L1 predicted samples calculated in the DVMR process is reused here.
[0111] Bidirectional optical flow (BDOF) is disabled if BCW is enabled for the current block, meaning the BCW weight index indicates unequal weights. Similarly, BDOF is disabled if WP is enabled for the current block, meaning either luma_weight_lx_flag is 1 for either of the two reference images. BDOF is also disabled when the CU is encoded / decoded using symmetric MVD mode or CIIP mode.
[0112] 2.8. Intra-frame and Inter-frame Joint Prediction (CIIP) In VVC, when a CU is encoded and decoded in Merge mode, if the CU contains at least 64 luma samples (i.e., CU width multiplied by CU height equal to or greater than 64), and if both the CU width and CU height are less than 128 luma samples, an additional flag is transmitted via signaling to indicate whether Intra / Inter Joint Prediction (CIIP) mode is applied to the current CU. As its name suggests, CIIP prediction combines inter-frame prediction signals with intra-frame prediction signals. The inter-frame prediction signal in CIIP mode... The inter-frame prediction process is derived using the same procedure as the regular Merge mode; and the intra-frame prediction signal... The regular intra prediction process with the planar mode is followed. Then, the intra prediction signal and the inter prediction signal are combined using a weighted average, where the weight values depend on the coding modes of the top and left neighboring blocks (as shown in Figure 11 The coding mode of the top and left neighboring blocks is calculated as follows: - If the top neighbor is available and is intra coded, set isIntraTop to 1, otherwise set it to 0; - If the left neighbor is available and is intra coded, set islntraLeft to 1, otherwise set it to 0; - If (islntraLeft + isIntraTop) is equal to 2, set wt to 3; - Otherwise, if (islntraLeft + isIntraTop) is equal to 1, set wt to 2; - Otherwise, set wt to 1.
[0113] The CIIP prediction is formed as follows:
[0114] 2.9. Affine motion compensated prediction In HEVC, only translational motion model is applied for motion compensated prediction (MCP). In the real world, there are many kinds of motions, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VVC, block-based affine transform motion compensated prediction is applied. As shown in Figure 12A and Figure 12B The 4-parameter affine model is shown in Figure 12A and the 6-parameter affine model is shown in Figure 12B The affine motion field of a block is described by the motion information of two control points (4-parameter) or three control point motion vectors (6-parameter).
[0115] For the 4-parameter affine motion model, the motion vector at sample position (x, y) in the block is derived as: (2-9) For the 6-parameter affine motion model, the motion vector at sample position (x, y) in the block is derived as: x, y (2-10) where (x0, y0) is the motion vector of the top-left control point, (x1, y1) is the motion vector of the top-right control point, and (x2, y2) is the motion vector of the left-bottom control point. mv 0x , mv 0y mv 1x ,mv 1y ) is the motion vector of the top-right control point, and (x, y) is the motion vector of the bottom-left control point. mv 2x , mv 2y ) is the motion vector of the bottom-left control point.
[0116] To simplify the motion-compensated prediction, block-based affine transform prediction is applied. To derive the motion vector of each 4x4 luma sub-block, the motion vector of the center sample of each sub-block is calculated according to the above equation (as shown in Figure 13 ), and rounded to 1 / 16 fractional precision. Then a motion-compensated interpolation filter is applied to generate the prediction of each sub-block with the derived motion vector. The sub-block size for chroma components is also set to 4x4. The MV of a 4x4 chroma sub-block is calculated as the average of the MVs of the four corresponding 4x4 luma sub-blocks.
[0117] As with translational inter prediction, there are also two affine inter prediction modes: affine Merge mode and affine AMVP mode.
[0118] 2.9.1. Affine Merge prediction The AF_MERGE mode can be applied to a CU whose width and height are both greater than or equal to 8. In this mode, the CPMVs of the current CU are generated based on the motion information of spatial neighboring CUs. There can be at most five CPMVP candidates, and one of them is indicated by signaling an index to be used for the current CU. The following three types of CPMV candidates are used to form the affine Merge candidate list: - Inherited affine Merge candidate inferred from the CPMVs of neighboring CUs.
[0119] - Constructed affine Merge candidate CPMVP derived using the translational MVs of neighboring CUs.
[0120] - Zero MV.
[0121] In VVC, there are at most two inherited affine candidates, which are derived from the affine motion model of neighboring blocks, one from the left neighboring CU and one from the above neighboring CU. The candidate blocks are as shown in Figure 14The scanning order is A0->A1 for the left side predictor and B0->B1->B2 for the above predictor. Only the first inherited candidate from each side is selected. No de-duplication check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive CPMVP candidates in the affine Merge list of the current CU. If a neighboring bottom-left block A is coded in affine mode, as shown, the motion vectors of the top-left, top-right and bottom-left corners of the CU containing block A are obtained , and When block A is coded with a 4-parameter affine model, two CPMVs of the current CU are calculated according to , When block A is coded with a 6-parameter affine model, three CPMVs of the current CU are calculated according to , and .
[0122] Figure 15 Control point motion vector inheritance is shown.
[0123] Constructed affine candidates mean that the candidate is constructed by combining the neighboring translational motion information of each control point. The motion information for the control points is derived from the prescribed spatial and temporal neighbors as shown. Figure 16 CPMVk (k = 1, 2, 3, 4) denotes the k-th control point. For CPMV1, the B2->B3->A2 blocks are checked and the MV of the first available block is used. For CPMV2, the B1->B0 blocks are checked and for CPMV3, the A1->A0 blocks are checked. If available, TMVP is used as CPMV4.
[0124] After obtaining the MVs of the four control points, the affine Merge candidate is constructed based on this motion information. The following combinations of control point MVs are used to construct in order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, { CPMV1, CPMV2}, { CPMV1, CPMV3}.
[0125] The combinations of 3 CPMVs construct 6-parameter affine Merge candidates and the combinations of 2 CPMVs construct 4-parameter affine Merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the related combination of control point MVs is discarded.
[0126] After the inherited affine Merge candidates and the constructed affine Merge candidates are checked, if the list is still not full, a zero MV is inserted at the end of the list.
[0127] 2.9.2. Affine AMVP Prediction The affine AMVP mode can be applied to CUs with both width and height greater than or equal to 16. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and another flag is signaled to indicate whether it is a 4-parameter affine or a 6-parameter affine. In this mode, the difference between the current CU's CPVM and its predicted CPMVP is signaled in the bitstream. The affine AMVP candidate list is of size 2 and is generated by sequentially using the following four types of CPVM candidates: – Inherited affine AMVP candidates inferred from the CPMV of neighboring CUs.
[0128] – Constructed affine AMVP candidate CPMVP derived using the translation MV of neighboring CUs.
[0129] – Translation MV from the neighboring CU.
[0130] – Zero MV.
[0131] The checking order for inherited affine AMVP candidates is the same as that for inherited affine Merge candidates. The only difference is that, for AVMP candidates, only affine CUs with the same reference picture as those in the current block are considered. No deduplication process is applied when inserting inherited affine motion predictions into the candidate list.
[0132] The constructed AMVP candidate is from Figure 16 The specified spatial neighbors are derived using the same checking order as in the affine Merge candidate construction. Additionally, the reference picture index of neighboring blocks is checked. The block that is the first to be inter-coded in the checking order and has the same reference picture as the current CU is used. Only one is considered. When the current CU is encoded using a 4-parameter affine mode and both mv0 and mv1 are available, they are added as candidates in the affine AMVP list. When the current CU is encoded using a 6-parameter affine mode and all three CPMVs are available, they are added as candidates in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.
[0133] If the affine AMVP list candidates are still less than 2 after the inherited affine AMVP candidates and constructed AMVP candidates are checked, mv0, mv1 and mv2 will be added as translational MVs in order to predict all control point MVs of the current CU if available. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list.
[0134] 2.9.3. Affine motion information storage In VVC, the CPMVs of an affine CU are stored in a separate buffer. The stored CPMVs are only used to generate inherited CPMVs in affine Merge mode and affine AMVP mode for the most recently coded CU. The subblock MVs derived from the CPMVs are used for motion compensation, MV derivation for translational MVs in Merge / AMVP list and deblocking.
[0135] To avoid picture row buffer for additional CPMVs, the affine motion data inheritance from a CU of the above CTU is treated differently from that from regular neighboring CUs. If the candidate CU for affine motion data inheritance is in the above CTU row, the bottom-left and bottom-right subblock MVs in the row buffer, instead of the CPMVs, are used for affine MVP derivation. In this way, the CPMVs are only stored in the local buffer. If the candidate CU is 6-parameter affine coded, the affine model is downgraded to the 4-parameter model. As shown in Figure 17 along the top CTU boundary, the bottom-left and bottom-right subblock motion vectors of a CU are used for affine inheritance for the CU in the bottom CTU.
[0136] 2.9.4. Prediction refinement with optical flow for affine mode Compared with pixel-based motion compensation, subblock-based affine motion compensation can save memory access bandwidth and reduce computational complexity at the cost of prediction accuracy loss. To achieve more fine-grained motion compensation, prediction refinement with optical flow (PROF) is used to refine the subblock-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after the subblock-based affine motion compensation is performed, the luma prediction samples are refined by adding the difference derived by the optical flow equation. PROF is described as the following four steps: Step 1) Subblock-based affine motion compensation is performed to generate subblock prediction .
[0137] Step 2) Spatial gradient of the subblock prediction is calculated at each sample position using a 3-tap filter [-1, 0, 1] and The gradient calculation is exactly the same as in BDOF.
[0138] (2-11) (2-12) The precision used to control the gradient. Each side of a subblock (i.e. 4x4) prediction is extended by one sample for gradient calculation. To avoid additional memory bandwidth and additional interpolation calculation, those extended samples on the extended boundary are copied from the nearest integer pixel position in the reference picture.
[0139] Step 3) The luminance prediction refinement is calculated by the following optical flow equation.
[0140] (2-13) where is the sample MV (denoted by ) calculated for sample position is the difference between the sample MV and the subblock MV of the subblock the sample belongs to, as shown in Figure 18 . is quantized in unit of 1 / 32 luminance sample precision.
[0141] Since the affine model parameters and the sample position relative to the subblock center do not change from subblock to subblock, can be calculated for the first subblock and reused for other subblocks in the same CU. Let and be the horizontal and vertical offset from the sample position to the center of the subblock, can be derived by the following equations: (2-14) (2-15) To maintain accuracy, the center of the subblock is calculated as ( ( W SB - 1 ) / 2, ( H SB - 1 ) / 2), where W SB and H SB are the width and height of the subblock, respectively.
[0142] For 4-parameter affine model, (2-16) For 6-parameter affine model, (2-17) where , , is the motion vector of the top-left, top-right and bottom-left control points, and is the width and height of the CU.
[0143] Step 4) Finally, the luminance prediction refinement is added to the sub-block prediction . The final prediction I’ is generated as the following equation.
[0144] (2-18) For an affine coded CU, PROF is not applied in two cases: 1) all control point MVs are the same, which indicates the CU has only translational motion; 2) the affine motion parameters are larger than the specified limit, because the sub-block based affine MC is downgraded to CU based MC to avoid large memory access bandwidth requirement.
[0145] A fast encoding method is applied to reduce the encoding complexity of affine motion estimation with PROF. In the following two cases, PROF is not applied in the affine motion estimation stage: a) if the CU is not the root block and its parent block does not select affine mode as its best mode, then PROF is not applied because the probability that the current CU selects affine mode as the best mode is low; b) if the amplitudes of the four affine parameters (C, D, E, F) are all smaller than a pre-defined threshold, and the current picture is not a low-delay picture, then PROF is not applied because the improvement introduced by PROF is small in this case. In this way, the affine motion estimation with PROF can be accelerated.
[0146] 2.10. Sub-block based temporal motion vector prediction (SbTMVP) VVC supports a sub-block based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the co-located picture to improve the motion vector prediction and the Merge mode for a CU in the current picture. The same co-located picture used by TMVP is used for SbTVMP. SbTMVP is different from TMVP in the following two main aspects: - TMVP predicts the motion at the CU level, but SbTMVP predicts the motion at the sub-CU level; - While TMVP obtains the temporal motion vector from the co-located block in the co-located picture (the co-located block is relative to the bottom-right or center block of the current CU), SbTMVP applies a motion displacement before obtaining the temporal motion information from the co-located picture, where the motion displacement is obtained from the motion vector of one of the spatial neighboring blocks from the current CU.
[0147] The SbTVMP process is as followsFigure 19A and Figure 19B SbTMVP predicts the motion vector of a sub-CU within the current CU in two steps. In the first step, the spatial neighbor Al in Figure 19A is examined. If Al has a motion vector using a collocated picture as its reference picture, then this motion vector is selected as the motion displacement to be applied. If no such motion is identified, the motion displacement is set to (0, 0).
[0148] In the second step, the motion displacement identified in step 1 is applied (i.e., added to the coordinates of the current block) to obtain the sub-CU level motion information (motion vector and reference index) from the collocated picture as shown in Figure 19B . Figure 19B The example in assumes that the motion displacement is set to the motion of block Al. Then, for each sub-CU, the motion information of its corresponding block in the collocated picture (covering the smallest motion grid of the center sample) is used to derive the motion information for the sub-CU. After the motion information of the collocated sub-CU is identified, it is converted to the motion vector and reference index of the current sub-CU in a similar way as the TMVP process of HEVC, where the temporal motion scaling is applied to align the reference picture of the temporal motion vector with the reference picture of the current CU.
[0149] Figure 19B The derivation of the sub-CU motion field by applying the motion displacement from the spatial neighbor and scaling the motion information from the corresponding collocated sub-CU is shown.
[0150] In VVC, the combination of the combined subblock-based Merge list containing both SbTMVP candidates and affine Merge candidates is used for the signaling of the subblock-based Merge mode. The SbTMVP mode is enabled / disabled by a sequence parameter set (SPS) flag. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry of the list of subblock-based Merge candidates, followed by the affine Merge candidates. The size of the subblock-based Merge list is signaled in the SPS, and the maximum allowed size of the subblock-based Merge list is 5 in VVC.
[0151] The sub-CU size used in SbTMVP is fixed to 8x8, and as with the affine Merge mode, the SbTMVP mode is only applicable to CUs with both width and height greater than or equal to 8.
[0152] The coding logic of the additional SbTMVP Merge candidate is the same as that of other Merge candidates, i.e., for each CU in a P slice or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.
[0153] 2.11. Adaptive Motion Vector Resolution (AMVR) In HEVC, when use_integer_mv_flag in the slice header is equal to 0, the motion vector difference (MVD) (between the motion vector of a CU and the predicted motion vector) is signaled in quarter-luma-sample units. In VVC, a CU-level adaptive motion vector resolution (AMVR) scheme is introduced. AMVR allows the MVD of a CU to be coded with different precision. Depending on the mode for the current CU (normal AMVP mode or affine AMVP mode), the MVD of the current CU can be adaptively selected as follows: - Normal AMVP mode: quarter-luma-sample, half-luma-sample, integer-luma-sample, or quarter-luma-sample.
[0154] - Affine AMVP mode: quarter-luma-sample, integer-luma-sample, or 1 / 16-luma-sample.
[0155] The CU-level MVD resolution indication is conditionally signaled if the current CU has at least one non-zero MVD component. If all MVD components (i.e., both the horizontal MVD and the vertical MVD for reference list L0 and reference list L1) are zero, the quarter-luma-sample MVD resolution is inferred.
[0156] For a CU with at least one non-zero MVD component, a first flag is signaled to indicate whether quarter-luma-sample MVD precision is used for the CU. If the first flag is 0, no further signaling is needed and quarter-luma-sample MVD precision is used for the current CU. Otherwise, a second flag is signaled to indicate whether half-luma-sample or other MVD precision (integer or quarter-luma-sample) is used for the normal AMVP CU. In the case of half-luma-sample, a 6-tap interpolation filter is used for the half-luma-sample positions instead of the default 8-tap interpolation filter. Otherwise, a third flag is signaled to indicate whether integer-luma-sample MVD precision or quarter-luma-sample MVD precision is used for the normal AMVP CU. In the case of affine AMVP CU, the second flag is used to indicate whether integer-luma-sample MVD precision or 1 / 16-luma-sample MVD precision is used. To ensure that the reconstructed MV has the expected precision (quarter-luma-sample, half-luma-sample, integer-luma-sample, or quarter-luma-sample), the motion vector predictor for the CU will be rounded to the same precision as the MVD before being added with the MVD. The motion vector predictor is rounded to zero (i.e., negative motion vector predictor is rounded to positive infinity and positive motion vector predictor is rounded to negative infinity).
[0157] The encoder uses RD check to determine the motion vector resolution for the current CU. To avoid performing CU-level RD check four times for each MVD resolution, in VTM11, the RD check for MVD precision other than quarter luma sample is invoked conditionally. For normal AMVP mode, the RD cost of quarter luma sample MVD precision and integer luma sample MV precision are first calculated. Then, the RD cost of integer luma sample MVD precision is compared with the RD cost of quarter luma sample MVD precision to decide whether it is necessary to further check the RD cost of quarter luma sample MVD precision. When the RD cost of quarter luma sample MVD precision is much smaller than the RD cost of integer luma sample MVD precision, the RD check of quarter luma sample MVD precision is skipped. Then, the check of half luma sample MVD precision is skipped if the RD cost of integer luma sample MVD precision is significantly larger than the best RD cost of previously tested MVD precisions. For affine AMVP mode, if no affine inter mode is selected after checking the rate-distortion cost of affine Merge / skip mode, Merge / skip mode, quarter luma sample MVD precision normal AMVP mode and quarter luma sample MVD precision affine AMVP mode, the 1 / 16 luma sample MV precision and 1-pixel MV precision affine inter mode are not checked. Moreover, in 1 / 16 luma sample and quarter luma sample MV precision affine inter mode, the affine parameters obtained in quarter luma sample MV precision affine inter mode are used as the starting search points.
[0158] 2.12. Bi-prediction with CU-level weights (BCW) In HEVC, bi-predicted signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, bi-prediction mode is extended beyond simple averaging to allow weighted averaging of two prediction signals.
[0159]
[0160] Five weights are allowed in weighted average bi-prediction, For each bi-predicted CU, the weight w is determined in one of two ways: 1) for non-Merge CU, the weight index is signaled after the motion vector difference; 2) for Merge CU, the weight index is inferred from neighboring blocks based on the Merge candidate index. BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w e {3, 4, 5}) are used.
[0161] - At the encoder, fast search algorithms are applied to find the weight index without significantly increasing the encoder complexity. These algorithms are summarized as follows. The reader can refer to the VTM software for more details. When combined with AMVR, if the current picture is a low-delay picture, the unequal weights are only checked conditionally for 1-pel motion vector precision and 4-pel motion vector precision.
[0162] - When combined with affine, affine ME will be performed for unequal weights if and only if the affine mode is selected as the current best mode.
[0163] - When the two reference pictures in bi-prediction are the same, the unequal weights are only checked conditionally.
[0164] - The unequal weights are not searched when certain conditions are met, depending on the POC distance between the current picture and its reference pictures, the coded QP, and the temporal level.
[0165] The BCW weight index is coded using one context-coded bin followed by bypass-coded bins. The first context-coded bin indicates whether equal weights are used; if unequal weights are used, the additional bins are signaled using bypass coding to indicate which unequal weight is used.
[0166] Weighted prediction (WP) is a coding tool supported by H.264 / AVC and HEVC standards for efficiently coding video content with gradual transitions. Support for WP is also added to the VVC standard. WP allows the signaling of a weighting parameter (weight and offset) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weight(s) and offset(s) of the corresponding reference picture(s) are applied. WP and BCW are designed for different types of video content. To avoid the interaction between WP and BCW that would complicate the VVC decoder design, if a CU uses WP, the BCW weight index is not signaled and w is inferred to be 4 (i.e., equal weights are applied). For Merge CUs, the weight index is inferred from neighboring blocks based on the Merge candidate index. This can be applied to both normal Merge mode and inherited affine Merge mode. For constructed affine Merge mode, the affine motion information is constructed based on the motion information of up to 3 blocks. The BCW index for a CU using constructed affine Merge mode is simply set equal to the BCW index of the first control point MV.
[0167] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded with CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights.
[0168] 2.13. Local Illumination Compensation (LIC) Local Illumination Compensation (LIC) is a coding tool to address the local illumination change between the current picture and its temporal reference picture. LIC is based on a linear model, where scaling factors and offsets are applied to the reference samples to obtain the prediction samples of the current block. Specifically, LIC can be mathematically modeled by the following equation:
[0169] where is the prediction signal of the current block at coordinate ; is the reference block pointed by the motion vector ; and are the corresponding scaling factors and offsets applied to the reference block. Figure 20 The LIC process is shown. In Figure 20 , when LIC is applied for a block, the Least Mean Square Error (LMSE) method is employed to derive the values of the LIC parameters (i.e., and ) by minimizing the difference between the neighboring samples of the current block (i.e., the template T in Figure 20 ) and their corresponding reference samples (i.e., T0 or T1 in Figure 20 ) in the temporal reference picture. Additionally, to reduce the computational complexity, both the template samples and the reference template samples are down-sampled (adaptive down-sampling) to derive the LIC parameters, i.e., only the shaded samples in Figure 20 are used to derive and .
[0170] To improve the coding performance, as shown in Figure 21 , no down-sampling is performed for the short side.
[0171] 2.14. Decoder-side Motion Vector Refinement (DMVR) To improve the accuracy of the MVs of the Merge mode, decoder-side motion vector refinement based on bilateral matching (BM) is applied in VVC. In the bi-prediction operation, a refined MV is searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method computes the distortion between two candidate blocks in the reference picture list L0 and list L1. As shown in Figure 22As shown, the SAD between the two blocks based on each MV candidate (e.g., MV0' and MV1') around the initial MV is calculated. The MV candidate with the lowest SAD becomes the refined MV and is used to generate the bi-predicted signal.
[0172] In VVC, the application of DMVR is restricted and is only applied to CUs that are coded with the following modes and features: - CU-level Merge mode with bi-predicted MV.
[0173] - One reference picture is past and the other reference picture is future relative to the current picture.
[0174] - The distance (i.e., POC difference) from the two reference pictures to the current picture is the same.
[0175] - Both reference pictures are short-term reference pictures.
[0176] - The CU has more than 64 luma samples.
[0177] - Both CU height and CU width are greater than or equal to 8 luma samples.
[0178] - The BCW weight index indicates equal weights.
[0179] - WP is not enabled for the current block.
[0180] - CIIP mode is not used for the current block.
[0181] The refined MV derived by the DMVR process is used to generate the inter-predicted samples and is also used in temporal motion vector prediction for future picture coding. While the original MV is used in the deblocking process and is also used in spatial motion vector prediction for future CU coding.
[0182] Additional features of DMVR are mentioned in the following sub-entry.
[0183] 2.14.1. Search scheme In DVMR, the search points are around the initial MV and the MV offsets obey the MV difference mirroring rule. In other words, any point (denoted by the candidate MV pair (MV0, MV1)) examined by DMVR obeys the following two equations: (2-20) (2-21) where This represents the refinement offset between the initial MV and the refined MV in one of the reference images. The refinement search range is two integer luminance samples from the initial MV. The search includes an integer sample offset search phase and a fractional sample refinement phase.
[0184] A 25-point full search is applied to the integer sample offset search. First, the SAD of the initial MV pair is calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR is terminated. Otherwise, the SAD of the remaining 24 points is calculated and checked in raster scan order. The point with the smallest SAD is selected as the output of the integer sample offset search stage. To reduce the impact of DMVR refinement uncertainties, a bias towards the original MV is proposed during the DMVR process. The SAD between the reference blocks referenced by the initial MV candidates is reduced by 1 / 4 of the SAD value.
[0185] The integer sample search is followed by fractional sample refinement. To save computational complexity, fractional sample refinement is derived using surface equations of parameter error, rather than through an additional search with SAD comparisons. Fractional sample refinement is conditionally invoked based on the output of the integer sample search phase. Fractional sample refinement is further applied when the integer sample search phase terminates in the first or second iteration with the minimum SAD at the center.
[0186] In subpixel offset estimation based on parametric error surfaces, the cost at the center location and the costs at the four nearest neighbor locations are used to fit a two-dimensional parabolic error surface equation of the following form. (2-22) in( This corresponds to the score position with the minimum cost, and C corresponds to the minimum cost. The above equation is solved by using the costs of the five search points. Calculated as: (2-23) (2-24) and The value is automatically constrained between -8 and 8 because all cost values are positive, and the minimum value is... This corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated score ( It is added to the integer distance thinning MV to obtain the subpixel accurate thinning increment MV.
[0187] 2.14.2. Bilinear Interpolation and Sample Filling In VVC, the resolution of MV is 1 / 16 luma sample. The samples at fractional positions are interpolated using an 8-tap interpolation filter. In DMVR, the search points are around the initial fractional pixel MV with integer sample offset, so the samples at those fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter is used to generate the fractional samples for the search process in DMVR. Another important effect of using the bilinear filter is that with a 2-sample search range, DVMR does not access more reference samples compared to the normal motion compensation process. After the refined MV is obtained through the DMVR search process, the normal 8-tap interpolation filter is applied to generate the final prediction. To not access more reference samples than the normal MC process, the samples that are not needed for the interpolation process based on the original MV but needed for the interpolation process based on the refined MV will be padded from those available samples.
[0188] 2.14.3. Maximum DMVR processing unit When the width and / or height of a CU is larger than 16 luma samples, it will be further divided into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process is limited to 16x16.
[0189] 2.15. Multi-pass decoder-side motion vector refinement In this contribution, a multi-pass decoder-side motion vector refinement is applied instead of DMVR. In the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16x16 sub-block within the coded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bi-directional optical flow (BDOF). The refined MVs are stored for both spatial and temporal motion vector prediction.
[0190] 2.15.1. First pass - block-based bilateral matching MV refinement In the first pass, the refined MV is derived by applying BM to the coded block. Similar to decoder-side motion vector refinement (DMVR), the refined MV is searched around two initial MVs (MV0 and MV1) in the reference picture lists L0 and L1. The refined MVs (MV0_pass1 and MV1_pass1) are derived around the initial MVs based on the minimum bilateral matching cost between the two reference blocks in L0 and L1.
[0191] The BM performs a local search to derive the integer sample precision intDeltaMV and the half-pel sample precision halfDeltaMv. The local search applies a 3x3 square search pattern to loop through a search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum values of sHor and sVer are 8.
[0192] The bilateral matching cost is computed as: bilCost = mvDistanceCost + sadCost. When the block size cbW * cbH is greater than 64, the MRSAD cost function is applied to remove the DC effect of the distortion between the reference blocks. The intDeltaMV or halfDeltaMV local search is terminated when the bilCost at the center point of the 3x3 search pattern has the minimum cost. Otherwise, the current minimum cost search point becomes the new center point of the 3x3 search pattern and the search for the minimum cost continues until it reaches the end of the search range.
[0193] The existing fractional sample refinement is further applied to derive the final deltaMV. The refined MV after the first pass is then derived as: MV0_pass1 = MV0 + deltaMV, MV1_pass1 = MV1 - deltaMV.
[0194] 2.15.2. Second pass - sub-block based bilateral matching MV refinement In the second pass, the refined MV is derived by applying BM to 16x16 grid sub-blocks. For each sub-block, the refined MV is searched around the two MVs (MV0_pass1 and MV1_pass1) for the reference picture lists L0 and L1 obtained from the first pass. The refined MVs (MV0_pass2(sbIdx2) and MV1_pass2(sbIdx2)) are derived based on the minimum bilateral matching cost between the two reference sub-blocks in L0 and L1.
[0195] For each sub-block, the BM performs a full search to derive the integer sample precision intDeltaMV. The full search has a search range [-sHor, sHor] in the horizontal direction and [-sVer, sVer] in the vertical direction, where the values of sHor and sVer are determined by the block dimensions and the maximum values of sHor and sVer are 8.
[0196] The bilateral matching cost is computed by applying a cost factor to the SATD cost between the two reference sub-blocks as: bilCost = satdCost * costFactor. The search area (2 * sHor + 1) * (2 * sVer + 1) is divided into Figure 23 up to 5 diamond search areas as shown. Each search area is assigned a costFactor which is determined by the distance (intDeltaMV) between each search point and the starting MV, and each diamond area is processed in order starting from the center of the search area. In each area, the search points are processed in a raster scan order from the top-left corner to the bottom-right corner of the area. When the minimum bilCost within the current search area is smaller than a threshold (which is equal to sbW * sbH), the integer-pixel full search is terminated, otherwise, the integer-pixel full search continues to the next search area until all search points are checked.
[0197] The BM performs a local search to derive the half-pel precision halfDeltaMv. The search pattern and cost function are the same as defined in 2.9.1.
[0198] The existing VVC DMVR fractional sample refinement is further applied to derive the final deltaMV (sbIdx2). The refined MV at the second pass is then derived as: MV0_pass2(sbIdx2) = MV0_pass1 + deltaMV(sbIdx2), MV1_pass2(sbIdx2) = MV1_pass1 - deltaMV(sbIdx2).
[0199] 2.15.3. Third pass - sub-block based bi-directional optical flow MV refinement In the third pass, the refined MV is derived by applying BDOF to the 8x8 grid sub-blocks. For each 8x8 sub-block, the BDOF refinement is applied to derive the scaled Vx and Vy without clipping from the refined MV of the parent block at the second pass. The derived bioMv (Vx, Vy) is rounded to 1 / 16 sample precision and clipped between -32 and 32.
[0200] The refined MV at the third pass (MV0_pass3(sbIdx3) and MV1_pass3(sbIdx3)) is derived as: MV0_pass3(sbIdx3) = MV0_pass2(sbIdx2) + bioMv, MV1_pass3(sbldx3) = MV0_pass2(sbldx2) - bioMv.
[0201] 2.16. Sample-based BDOF In sample-based BDOF, instead of deriving motion refinements (Vx, Vy) on a block basis, motion refinements are performed for each sample.
[0202] A coded block is divided into 8x8 sub-blocks. For each sub-block, whether to apply BDOF is determined by checking the SAD between two reference sub-blocks against a threshold. If it is decided to apply BDOF to a sub-block, for each sample in the sub-block, a sliding 5x5 window is used, and the existing BDOF process is applied for each sliding window to derive Vx and Vy. The derived motion refinements (Vx, Vy) are applied to adjust the bi-predicted sample value for the center sample of the window.
[0203] 2.17. Extended Merge prediction In VVC, the Merge candidate list is constructed by including the following five types of candidates in order: (1) Spatial MVP from spatial neighboring CUs.
[0204] (2) Temporal MVP from collocated CUs.
[0205] (3) History-based MVP from a FIFO table.
[0206] (4) Pairwise average MVP.
[0207] (5) Zero MV.
[0208] The size of the Merge list is signaled in the sequence parameter set header, and the maximum allowed size of the Merge list is 6. For each CU coded in Merge mode, the index of the best Merge candidate is coded using truncated unary binarization (TU). The first bin of the Merge index is coded with a context, and bypass coding is used for the other bins.
[0209] The derivation process of each category of Merge candidates is provided in this section. As done in HEVC, VVC also supports parallel derivation of the Merge candidate list for all CUs within a certain size of region.
[0210] 2.17.1. Spatial candidate derivation The derivation of spatial Merge candidates in VVC is the same as in HEVC, except that the positions of the first two Merge candidates are swapped. Among the candidates located at Figure 24 positions shown in FIG. 6B, at most four Merge candidates are selected. The derivation order is B0, A0, B1, A1, and B2. Position B2 is considered only when one or more than one CU at positions B0, A0, B1, A1 is not available (e.g., because it belongs to another slice or tile) or is intra coded. After the addition of the candidate at position A1, a redundancy check is performed for the addition of the remaining candidates, which ensures that candidates with the same motion information are not included in the list, thereby improving coding efficiency. To reduce the computational complexity, not all possible pairs of candidates are considered in the mentioned redundancy check. Instead, only pairs linked by an arrow in FIG. 6B are considered, and a candidate is added to the list only if the corresponding candidate for the redundancy check does not have the same motion information. Figure 25
[0211] 2.17.2. Temporal candidate derivation In this step, only one candidate is added to the list. Specifically, in the derivation of the temporal Merge candidate, the scaled motion vector is derived based on a collocated CU belonging to a collocated reference picture. The reference picture list to be used for the derivation of the collocated CU is explicitly signaled in the slice header. As shown by the dashed line in FIG. 6C, the scaled motion vector for the temporal Merge candidate is obtained by scaling the motion vector of the collocated CU using the POC distances tb and td, where tb is defined as the POC difference between the reference picture of the current picture and the current picture, and td is defined as the POC difference between the reference picture of the collocated picture and the collocated picture. The reference picture index of the temporal Merge candidate is set equal to 0. Figure 26
[0212] As shown in FIG. 6D, the position for the temporal candidate is selected between candidates C0 and C1. If the CU at position C0 is not available, intra coded, or outside the current row of the CTU, position C1 is used. Otherwise, position C0 is used for the derivation of the temporal Merge candidate. Figure 27
[0213] 2.17.3. History-based Merge candidate derivation A history-based MVP (HMVP) Merge candidate is added to the Merge list, after the spatial MVP and the TMVP. In this method, the motion information of previously coded blocks is stored in a table and used as MVPs for the current CU. During the encoding / decoding process, a table with multiple HMVP candidates is maintained. When a new CTU row is encountered, the table is reset (emptied). As long as there is a non-subblock inter coded CU, the associated motion information is added to the last entry of the table as a new HMVP candidate.
[0214] The HMVP table size S is set to 6, which indicates that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained first-in-first-out (FIFO) rule is utilized, where a redundancy check is first applied to find if there is an identical HMVP in the table. If found, the identical HMVP is removed from the table and all the HMVP candidates after it are moved forward and the identical HMVP is inserted to the last entry of the table.
[0215] The HMVP candidates can be used in the Merge candidate list construction process. The latest few HMVP candidates in the table are checked in order and inserted into the candidate list, after the TMVP candidates. A redundancy check is applied for the HMVP candidates against the spatial or temporal Merge candidates.
[0216] To reduce the number of redundancy check operations, the following simplification is introduced: The number of HMPV candidates used for Merge list generation is set to (N <= 4)? M : (8 N), where N indicates the number of existing candidates in the Merge list and M indicates the number of available HMVP candidates in the table.
[0217] The Merge candidate list construction process from HMVP is terminated once the total number of available Merge candidates reaches the maximum allowed Merge candidates minus 1.
[0218] 2.17.4. Pairwise average Merge candidate derivation Pairwise average candidates are generated by averaging predefined pairs of candidates in the existing Merge candidate list, and the predefined pairs are defined as {(0, 1), (0, 2), (1, 2), (0, 3), (1, 3), (2, 3)}, where the numbers represent the Merge indices in the Merge candidate list. The averaged motion vector is calculated separately for each reference list. If both motion vectors are available in one list, the two motion vectors are averaged even if they point to different reference pictures; if only one motion vector is available, the motion vector is used directly; if no motion vector is available, the list is kept invalid.
[0219] When the Merge list is not full after adding the pairwise average Merge candidates, a zero-MVP is inserted at the end until the maximum number of Merge candidates is reached.
[0220] 2.17.5. Merge estimation region The Merge estimation region (MER) allows independent derivation of the Merge candidate list for CUs in the same Merge estimation region (MER). Candidate blocks that are within the same MER as the current CU are not included for the generation of the Merge candidate list of the current CU. In addition, the update process of the history-based motion vector predictor candidate list is only updated when ( xCb + cbWidth ) » Log2ParMrgLevel is greater than xCb » Log2ParMrgLevel and ( yCb + cbHeight ) » Log2ParMrgLevel is greater than ( yCb » Log2ParMrgLevel ), where ( xCb, yCb ) is the top-left luma sample position of the current CU in the picture and ( cbWidth, cbHeight ) is the CU size. The MER size is selected at the encoder side and is signaled in the sequence parameter set as log2_parallel_merge_level_minus2.
[0221] 2.18. New Merge candidates 2.18.1. Non-adjacent Merge candidate derivation In VVC, Figure 28 The five spatial neighboring blocks and one temporal neighbor shown are used to derive Merge candidates.
[0222] It is proposed to derive additional Merge candidates from positions that are non-adjacent to the current block using the same modes as in VVC. To achieve this, for each search pass i, a virtual block is generated based on the current block as follows: First, the relative position of the virtual block to the current block is calculated by: Offsetx = -i x gridX, Offsety = -i x gridY where Offsetx and Offsety represent the offset of the top-left corner of the virtual block to the top-left corner of the current block, and gridX and gridY are the width and height of the search grid.
[0223] Second, the width and height of the virtual block are calculated by: newWidth = i x 2 x gridX + currWidth, newHeight = i x 2 x gridY + currHeight.
[0224] where currWidth and currHeight are the width and height of the current block. newWidth and newHeight are the width and height of the new virtual block.
[0225] gridX and gridY are currently set to currWidth and currHeight, respectively.
[0226] Figure 29 The relationship between the virtual block and the current block is shown. After the virtual block is generated, blocks Ai, Bi, Ci, Di, and Ei can be regarded as the VVC spatial neighboring blocks of the virtual block, and their positions are obtained with the same mode as in VVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, blocks Ai, Bi, Ci, Di, and Ei are the spatial neighboring blocks used in the VVC Merge mode.
[0227] When constructing the Merge candidate list, de-duplication is performed to ensure that each element in the Merge candidate list is unique. The maximum search round is set to 1, which means that five non-adjacent spatial neighboring blocks are utilized.
[0228] The non-adjacent spatial Merge candidates are inserted into the Merge list after the temporal Merge candidates in the order of B1 -> A1 -> C1 -> D1 -> E1.
[0229] 2.18.2. STMVP It is proposed to use three spatial Merge candidates and one temporal Merge candidate to derive an average candidate as an STMVP candidate.
[0230] STMVP is inserted before the top-left spatial Merge candidate.
[0231] The STMVP candidate is de-duplicated with all previous Merge candidates in the Merge list.
[0232] For spatial candidates, the first three candidates in the current Merge candidate list are used.
[0233] For temporal candidates, the same position as the VTM / HEVC collocated position is used.
[0234] For spatial candidates, the first, second and third candidates inserted into the current Merge candidate list before STMVP are denoted as F, S and T.
[0235] The temporal candidate with the same position as the VTM / HEVC collocated position used in TMVP is denoted as Col.
[0236] The motion vector of the STMVP candidate in prediction direction X (denoted as mvLX) is derived as follows: 1) If the reference indices of the four Merge candidates are all valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_S + mvLX_T + mvLX_Col) » 2 2) If the reference indices of three of the four Merge candidates are valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F x 3 + mvLX_S x 3 + mvLX_Col x 2) » 3 or mvLX = (mvLX_F x 3 + mvLX_T x 3 + mvLX_Col x 2) » 3 or mvLX = (mvLX_S x 3 + mvLX_T x 3 + mvLX_Col x 2) » 3.
[0237] 3) If the reference indices of two of the four Merge candidates are valid and equal to 0 in prediction direction X (X = 0 or 1), mvLX = (mvLX_F + mvLX_Col) » 1 or mvLX = (mvLX_S + mvLX_Col) » 1 or mvLX = (mvLX_T + mvLX_Col) » 1.
[0238] Note: If the temporal candidate is not available, the STMVP mode is turned off.
[0239] 2.18.3. Merge list size If both non-adjacent Merge candidates and STMVP Merge candidates are considered, the size of the Merge list is signaled in the sequence parameter set header and the maximum allowed size of the Merge list is 8.
[0240] 2.19. Geometric partition mode (GPM) In VVC, a geometric partition mode is supported for inter prediction. The geometric partition mode is signaled using a CU-level flag as a kind of Merge mode, where other Merge modes include regular Merge mode, MMVD mode, CIIP mode and subblock Merge mode. For each possible CU size (where 8x64 and 64x8 are not included), the geometric partition mode supports 64 partitions in total.
[0241] When this mode is used, the CU is divided into two parts by a straight line positioned geometrically Figure 30 . The position of the dividing line is mathematically derived from the angle and offset parameters of the specific partition. Each part of the geometric partition in the CU is inter predicted using its own motion; only uni-prediction is allowed for each partition, i.e., each part has one motion vector and one reference index. The uni-prediction motion constraint is applied to ensure the same as regular bi-prediction, i.e., only two motion- compensated predictions are needed for each CU. The uni-prediction motion for each partition is derived using the process described in 2.19.1.
[0242] If the geometric partition mode is used for the current CU, further the geometric partition index and two Merge indices (one for each partition) are signaled to indicate the partition mode (angle and offset) of the geometric partition. The number of maximum GPM candidate size is explicitly signaled in the SPS and the syntax binarization for GPM Merge indices is specified. After each part of the predicted geometric partition, a hybrid process with adaptive weights as in 2.19.2 is used to adjust the sample values along the geometric partition edge. This is the prediction signal for the whole CU and the transform and quantization processes will be applied to the whole CU as in other prediction modes. Finally, the motion field of the CU predicted using the geometric partition mode is stored as in 2.19.3.
[0243] 2.19.1. Uni-prediction candidate list construction The uni-prediction candidate list is directly derived from the Merge candidate list constructed according to the extended Merge prediction process in 2.17. Let n denote the index of the uni-prediction motion in the geometric uni-prediction candidate list. The LX motion vector of the n-th extended Merge candidate, where X equals the parity of n, is used as the n-th uni-prediction motion vector for the geometric partition mode. These motion vectors are marked with "x" in Figure 31 If the corresponding LX motion vector of the n-th extended Merge candidate does not exist, the L(l-X) motion vector of the same candidate is used instead as the uni-prediction motion vector for the geometric partition mode.
[0244] 2.19.2. Blending along the geometric partition edge After each part of the geometric partition is predicted using its own motion, blending is applied to the two prediction signals to derive the samples around the geometric partition edge. The blending weight for each position of the CU is derived based on the distance between the single position and the partition edge.
[0245] The distance to the partition edge for a position is derived as: (2-25) (2-26) (2-27) (2-28) where is the index for the angle and offset of the geometric partition, which depends on the signaled geometric partition index. and the sign of .
[0246] The weight of each part of the geometric partition is derived as follows: (2-29) (2-30) (2-31) partldx depends on the angle index . An example of the weight is shown in Figure 32 .
[0247] 2.19.3. Motion field storage for geometric partition mode Mv1 from the first part of the geometric partition, Mv2 from the second part of the geometric partition, and a combined Mv of Mv1 and Mv2 are stored in the motion field of the geometric partition mode coded CU.
[0248] The motion vector type stored for each individual position in the motion field is determined as: (2-32) where motionldx is equal to which is re-computed from equation (2-18). partldx depends on the angle index .
[0249] If sType is equal to 0 or 1, Mv0 or Mv1 is stored in the corresponding motion field, otherwise, if sType is equal to 2, a combined Mv from Mv0 and Mv2 is stored. The combined Mv is generated using the following procedure: 1) If Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), Mv1 and Mv2 are simply combined to form a bi-predictive motion vector.
[0250] Otherwise, if Mvl and Mv2 are from the same list, only the uni-predictive motion Mv2 is stored.
[0251] 2.20. Multi-hypothesis prediction In multi-hypothesis prediction (MHP), up to two additional prediction values are signaled on top of the inter AMVP mode, regular Merge mode, affine Merge and MMVD modes. The resulting overall prediction signal is iteratively accumulated with each additional prediction signal.
[0252]
[0253] The weighting factor a is specified according to the following Table 4:
[0254] For the inter AMVP mode, MHP is applied only when non-equal weights in BCW are selected in bi-predictive mode.
[0255] The additional hypotheses can be either Merge mode or AMVP mode. In the case of Merge mode, the motion information is indicated by the Merge index and the Merge candidate list is the same as in the geometric partition mode. In the case of AMVP mode, the reference index, MVP index and MVD are signaled.
[0256] 2.21. Non-adjacent spatial candidates Non-adjacent spatial Merge candidates are inserted into the regular Merge candidate list after TMVP. The pattern of spatial Merge candidates is shown in Figure 33 Table 6. The distance between a non-adjacent spatial candidate and the current coding block is based on the width and height of the current coding block.
[0257] 2.22. Template Matching (TM) Template Matching (TM) is a decoder-side MV derivation method to refine the motion information of a current CU by finding the closest match between a template (i.e., the top and / or left neighboring blocks of the current CU) in the current picture and a block (i.e., of the same size as the template) in the reference picture. As shown in Figure 34 Table 6, a better MV is searched around the initial motion of the current CU within a search range of [-8, +8] pixels. There are two modifications in the template matching in this contribution: the search step size is determined based on the AMVR mode, and in Merge mode TM can be cascaded with the bilateral matching process.
[0258] In AMVP mode, the MVP candidates are determined based on the template matching error to select one that achieves the smallest difference between the current block template and the reference block template, and then TM is only performed for that specific MVP candidate. TM refines this MVP candidate by using an iterative diamond search, starting from the full-pel MVD precision (or 4-pel for 4-pixel AMVR mode) within a search range of [-8, +8] pixels. The AMVP candidate can be further refined by using a cross search with full-pel MVD precision (or 4-pel for 4-pixel AMVR mode), and then half-pel and quarter-pel are used in turn according to the AMVR mode specified in Table 5. This search process ensures that the MVP candidate still maintains the same MV precision as indicated by the AMVR mode after the TM process.
[0259]
[0260] In Merge mode, a similar search method is applied to the Merge candidate indicated by the Merge index. As shown in Table 5, TM can be performed up to 1 / 8-pel MVD precision, or skip those precisions beyond half-pel MVD precision, depending on whether the alternative interpolation filter is used according to the Merge motion information (i.e., when AMVR is in half-pel mode). In addition, when TM mode is enabled, template matching can work as a standalone process, or as an additional MV refinement process between the block-based and sub-block-based bilateral matching (BM) methods, depending on whether BM can be enabled according to its enabling condition check.
[0261] In template matching Merge mode, the encoder can choose for a CU from uni-prediction from list 0, uni-prediction from list 1, or bi-prediction. The choice is based on the template matching cost as follows: if costB1<= factor * min(cost0, cost1) use bi-prediction; else if cost0<= cost1 use uni-prediction from list 0; else use uni-prediction from list 1; where cost0 is the SAD of list 0 template matching, cost1 is the SAD of list 1 template matching, and costB1 is the SAD of bi-prediction template matching. The value of factor is equal to 1.125, which means the selection process is biased towards bi-prediction.
[0262] 2.23. Overlapped Block Motion Compensation (OBMC) Overlapped Block Motion Compensation (OBMC) has been used in H.263 previously. In JEM, unlike H.263, OBMC can be turned on and off using CU-level syntax. When OBMC is used in JEM, OBMC is performed for all motion compensation (MC) block boundaries except the right and bottom boundaries of a CU. In addition, it is applied to both luma and chroma components. In JEM, a MC block corresponds to a coding block. When a CU is coded in sub-CU mode (including sub-CU Merge, affine and FRUC modes), each sub-block of the CU is a MC block. To handle CU boundaries in a uniform way, OBMC is performed at the sub-block level for all MC block boundaries, with the sub-block size set to be equal to 4x4, as shown in Figure 35 .
[0263] When OBMC is applied to a current sub-block, in addition to the current motion vector, the motion vectors of four connected neighboring sub-blocks (if available and not identical to the current motion vector) are also used to derive the prediction block for the current sub-block. These multiple prediction blocks based on multiple motion vectors are combined to generate the final prediction signal for the current sub-block.
[0264] A prediction block based on motion vectors of neighboring sub-blocks is denoted as PN, where N denotes the index for the neighboring top, bottom, left and right sub-blocks, and a prediction block based on motion vectors of the current sub-block is denoted as PC. When PN is based on motion information of neighboring sub-blocks containing the same motion information as the current sub-block, no OBMC is performed from PN. Otherwise, each sample of PN is added to the same sample in PC, i.e. four rows / columns of PN are added to PC. Weighting factors {1 / 4, 1 / 8, 1 / 16, 1 / 32} are used for PN and weighting factors {3 / 4, 7 / 8, 15 / 16, 31 / 32} are used for PC. An exception is small MC blocks (i.e. when the height or width of the coded block is equal to 4 or the CU is coded in sub-CU mode), for which only two rows / columns of PN are added to PC for small MC blocks. In this case, weighting factors {1 / 4, 1 / 8} are used for PN and weighting factors {3 / 4, 7 / 8} are used for PC. For PN generated based on vertical (horizontal) neighboring sub-blocks' motion vectors, samples in the same row (column) of PN are added to PC with the same weighting factor.
[0265] In JEM, a CU-level flag is signaled to indicate whether OBMC is applied to the current CU for CUs with size smaller than or equal to 256 luma samples. For CUs with size larger than 256 luma samples or coded without AMVP mode, OBMC is applied by default. At the encoder, when OBMC is applied to a CU, its impact is considered during the motion estimation stage. The prediction signal formed by OBMC using the motion information of the top and left neighboring blocks is used to compensate the top and left boundaries of the original signal of the current CU, before the normal motion estimation process is applied.
[0266] 2.24. Multiple Transform Selection (MTS) for core transform In addition to the DCT-II already employed in HEVC, a Multiple Transform Selection (MTS) scheme is used for residual coding of both inter and intra coded blocks. It uses multiple transforms selected from DCT8 / DST7. The newly introduced transform matrices are DST-VII and DCT-VIII. Table 6 shows the basis functions of the selected DST / DCT.
[0267]
[0268] To maintain the orthogonality of the transform matrices, the transform matrices are quantized more precisely than in HEVC. To keep the mid values of the transform coefficients in the 16-bit range, all coefficients have 10 bits after the horizontal transform and after the vertical transform.
[0269] To control the MTS scheme, separate enabling flags are specified at SPS level for intra and inter respectively. When MTS is enabled at SPS, a CU level flag is signaled to indicate whether MTS is applied or not. Here, MTS is applied only to luma. MTS signaling is skipped when one of the following conditions is applied.
[0270] - The position of the last significant coefficient of the luma TB is less than 1 (i.e. only DC).
[0271] - The last significant coefficient of the luma TB is located within the MTS zero-out region.
[0272] If the MTS CU flag is equal to 0, DCT2 is applied in both directions. However, if the MTS CU flag is equal to 1, two other flags are additionally signaled to indicate the transform type for the horizontal and vertical directions respectively. The transform and signaling mapping table is shown in Table 7. A unified transform selection for ISP and implicit MTS is used by removing the intra mode and block shape dependency. If the current block is ISP mode, or if the current block is an intra block and both intra explicit MTS and inter explicit MTS are turned on, only DST7 is used for both the horizontal and vertical transform kernel. When it comes to transform matrix precision, 8-bit primary transform kernel is used. Therefore, all transform kernels used in HEVC remain unchanged, including 4-point DCT-2 and DST-7, 8-point, 16-point and 32-point DCT-2. In addition, other transform kernels including 64-point DCT-2, 4-point DCT-8, 8-point, 16-point, 32-point DST-7 and DCT-8 use 8-bit primary transform kernel.
[0273]
[0274] To reduce the complexity of large size DST-7 and DCT-8, for DST-7 and DCT-8 blocks with size (width or height, or both width and height) equal to 32, high frequency transform coefficients are zeroed out. Only coefficients within the 16x16 low frequency region are kept.
[0275] As in HEVC, the residual of a block can be coded in transform skip mode. To avoid the redundancy of syntax coding, the transform skip flag is not signaled when the CU level MTS CU flag is not equal to 0. Note that when LFNST or MIP is activated for the current CU, the implicit MTS transform is set to DCT2. And when MTS is enabled for inter coded blocks, implicit MTS can also be enabled.
[0276] 2.25. Sub-block transform (SBT) In VTM, sub-block transform is introduced for inter-predicted CUs. In this transform mode, only sub-parts of the residual block are coded for a CU. When cu_cbf is equal to 1 for an inter-predicted CU, cu_sbt_flag can be signaled to indicate whether the whole residual block or the sub-parts of the residual block are coded. In the former case, inter-MTS information is further parsed to determine the transform type of the CU. In the latter case, one part of the residual block is coded with an inferred adaptive transform and the other part of the residual block is zeroed.
[0277] When SBT is applied to an inter-coded CU, SBT type and SBT position information are signaled in the bitstream. As shown in Figure 36 , there are two SBT types and two SBT positions. For SBT-V (or SBT-H), the TU width (or height) can be equal to half or ¼ of the CU width (or height), resulting in 2:2 split or 1:3 / 3:1 split. 2:2 split is like binary tree (BT) split, while 1:3 / 3:1 split is like asymmetric binary tree (ABT) split. In ABT split, only the small region contains non-zero residual. If one dimension of the CU is 8 in luma samples, 1:3 / 3:1 split along that dimension is not allowed. There are at most 8 SBT modes for a CU.
[0278] Position-dependent transform kernel selection is applied to luma transform blocks in SBT-V and SBT-H (chroma TB always uses DCT-2). The two positions of SBT-H and SBT-V are associated with different kernel transforms. More specifically, in Figure 36 , the horizontal and vertical transforms for each SBT position are specified. For example, the horizontal and vertical transforms for SBT-V position 0 are DCT-8 and DST-7, respectively. When one side of a residual TU is larger than 32, both dimensions of transform are set to DCT-2. Thus, sub-block transform jointly specifies the TU tiling of a residual block, cbf, and the horizontal and vertical kernel transform types.
[0279] SBT is not applied to CUs coded in intra-inter joint mode.
[0280] 2.26. Template matching based adaptive Merge candidate reordering To improve coding efficiency, after constructing the Merge candidate list, the order of each Merge candidate is adjusted according to the template matching cost. Merge candidates are arranged in the list according to ascending template matching cost. It is operated in the form of subgroups.
[0281] The template matching cost is measured by the SAD (sum of absolute difference) between the neighboring samples of the current CU and their corresponding reference samples. If the Merge candidate includes motion information for bi-prediction, the corresponding reference samples are the average of the corresponding reference samples in reference list 0 and reference list 1, as shown in Figure 37 If the Merge candidate contains motion information at sub-CU level, the corresponding reference samples consist of the neighboring samples of the corresponding reference sub-block, as shown in Figure 38
[0282] As shown in Figure 39 The sorting process is operated in the form of sub-groups. The first three Merge candidates are sorted together. The last three Merge candidates are sorted together.
[0283] The template size (width of the left template or height of the top template) is 1. The sub-group size is 3.
[0284] 2.27. Adaptive Merge candidate list It can be assumed that the number of Merge candidates is 8. The first 5 Merge candidates are taken as the first sub-group and the last 3 Merge candidates are taken as the second sub-group (i.e. the last sub-group).
[0285] For the encoder, after constructing the Merge candidate list, some of the Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, as shown in Figure 40
[0286] More specifically, the template matching cost of the Merge candidates in all sub-groups except the last sub-group is calculated; then the Merge candidates in their own sub-group except the last sub-group are reordered; finally, the final Merge candidate list is obtained.
[0287] For the decoder, after constructing the Merge candidate list, some / no Merge candidates are adaptively reordered in ascending order of the Merge candidate cost, as shown in Figure 41 Figure 41 In the case, the sub-group in which the selected (signaled) Merge candidate is located is referred to as the selected sub-group.
[0288] More specifically, if the selected Merge candidate is located in the last sub-group, the Merge candidate list construction process is terminated after the selected Merge candidate is derived, no reordering is performed, and the Merge candidate list is not changed; otherwise, the process is performed as follows: After deriving all Merge candidates in the selected subgroup, the Merge candidate list construction process is terminated; the template matching cost of the Merge candidates in the selected subgroup is calculated; the Merge candidates in the selected subgroup are reordered; and finally, a new Merge candidate list is obtained.
[0289] For both the encoder and the decoder, the template matching cost is derived as a function of T and RT, where T is a set of samples in the template and RT is a set of reference samples for the template.
[0290] When deriving the reference samples of the template of a Merge candidate, the motion vector of the Merge candidate is rounded to integer pixel precision.
[0291] The reference samples (RT) of the template for bi-prediction are derived by a weighted average of the reference samples (T0) of the template in reference list 0 and the reference samples (T1) of the template in reference list 1 .
[0292] (2-33) where the weight (8-w) of the reference template in reference list 0 and the weight (w) of the reference template in reference list 1 are determined by the BCW index of the Merge candidate. The BCW index equal to {0, 1, 2, 3, 4} corresponds to w equal to {-2, 3, 4, 5, 10}, respectively.
[0293] If the local illumination compensation (LIC) flag of the Merge candidate is true, the reference samples of the template are derived using the LIC method.
[0294] The template matching cost is calculated based on the sum of absolute differences (SAD) of T and RT.
[0295] The template size is 1. This means that the width of the left template and / or the height of the above template is 1.
[0296] If the coding mode is MMVD, the Merge candidates used to derive the base Merge candidate are not reordered.
[0297] If the coding mode is GPM, the Merge candidates used to derive the uni-prediction candidate list are not reordered.
[0298] 2.28. Geometric prediction mode with motion vector difference In the geometric prediction with motion vector difference (GMVD) mode, each geometric partition in GPM can decide whether to use GMVD or not. If GMVD is selected for a geometric region, the MV of that region is calculated as the sum of the MV of the Merge candidate and the MVD. All other processing remains the same as in GPM.
[0299] With GMVD, the MVD is signaled as a pair of direction and distance. Nine candidate distances (1 / 4-pixel, 1 / 2-pixel, 1-pixel, 2-pixel, 3-pixel, 4-pixel, 6-pixel, 8-pixel, 16-pixel) and eight candidate directions (four horizontal / vertical directions and four diagonal directions) are involved. In addition, when pic_fpel_mmvd_enabled_flag is equal to 1, the MVD in GMVD is also left-shifted by 2 as in MMVD.
[0300] 2.29. Affine model inheritance based on history parameters and non-adjacent affine mode Affine model inheritance based on history parameters (HAMI) allows affine models to be inherited from previously affine coded blocks that can not be adjacent to the current block. Similar to the enhanced regular Merge mode, a non-adjacent affine mode (NA-AFF) is introduced.
[0301] A first history parameter table (HPT) is established. One entry of the first HPT stores a set of affine parameters: a, b, c, and d, each affine parameter represented by a 16-bit signed integer. The entries in the HPT are categorized by reference list and reference index. Five reference indices are supported for each reference list in the HPT. In a formal way, the category of the HPT (denoted as HPTCat) is calculated as HPTCat (RefList, RefIdx) = 5 x RefList + min (RefIdx, 4), where RefList and RefIdx represent the reference picture list (0 or 1) and the reference index, respectively. For each category, up to 7 entries can be stored, resulting in a total of 70 entries in the HPT. At the beginning of each CTU row, the number of entries for each category is initialized to 0. After decoding an affine coded CU with reference list RefList cur and reference index RefIdx cur , the affine parameters are utilized to update the entries in the category HPTCat(RefList cur , RefIdx cur ) in a similar way as the HMVP table update.
[0302] A history affine parameter candidate (HAPC) is derived from the entries in the HPT with the same reference list as the current block and the reference index being the same as the current block's reference index. Figure 41A block, denoted as A0, A1, A2, B0, B1, B2, or B3, and a set of affine parameters stored in the corresponding entry in the first HPT are derived. The MV of the neighboring 4x4 blocks is used as the base MV. Formulatically, the MV of the current block at position (x, y) is calculated as: , in( mv h base , mv v base ) represents the MV of the nearest 4x4 block, ( x base , y base () indicates the center position of the nearest 4x4 block. x , y The MV can be the top left, top right, or bottom left corner of the current block to obtain the corner position MV (CPMV) for the current block, or it can be the center of the current block to obtain the regular MV for the current block.
[0303] A second history parameter table (HPT) containing basic MV information is also attached. The second HPT contains nine entries, one of which includes the basic MV, reference index, four affine parameters for each reference list, and the base position. The attached Merge HAPC can be generated from the second HPT using the basic MV information of the corresponding affine models stored in the entries. The differences between the first and second HPTs are as follows: Figure 42 As shown.
[0304] Furthermore, paired affine merge candidates are generated from two historically derived or non-historically derived affine merge candidates. Paired affine merge candidates are generated by averaging the CPMV of existing affine merge candidates in the list.
[0305] In response to the newly introduced HAPC, the size of the sub-block-based Merge candidate list was increased from 5 to 15, all of which are involved in the ARMC process.
[0306] In NA-AFF, the pattern for obtaining the nearest neighbor in non-adjacent spatial domains is as follows: Figure 43A and Figure 43B As shown. Figure 43A Candidates for deriving inheritance are shown, and Figure 43B The candidate for the first type of construction is shown. Similar to existing non-adjacent regular Merge candidates, the distance between non-adjacent spatial neighbors in NA-AFF and the current codec block is also defined based on the width and height of the current CU.
[0307] Figure 43A and Figure 43B Motion information of non-adjacent spatial neighbors in Figure 43A is utilized to generate additional inherited and constructed affine Merge / AMVP candidates. Specifically, for inherited candidates, the same derivation process of inherited affine Merge / AMVP candidates in VVC remains unchanged except that CPMVs are inherited from non-adjacent spatial neighbors. Non-adjacent spatial neighbors are checked based on their distance to the current block (i.e., from near to far). At a certain distance, only the first available neighbor (coded in affine mode) from each side (e.g., left and above) of the current block is included for inherited candidate derivation. As shown by the red dashed arrows in
[0308] For the first type of constructed candidates, as shown in Figure 43B , the positions of one left and above non-adjacent spatial neighbor are first determined independently; afterwards, the position of the left-above neighbor can be determined accordingly, which together with the left and above non-adjacent neighbors enclose a rectangular virtual block. Then, as shown in Figure 44 , the motion information of the three non-adjacent neighbors is used to form CPMVs at the top-left (A), top-right (B), and bottom-left (C) of the virtual block, which are finally mapped to the current CU to generate the corresponding constructed candidate.
[0309] NA-AFF candidates are inserted into the existing affine Merge candidate list and affine AMVP candidate list according to the following order: Affine Merge mode: 1. SbTMVP candidate, if available.
[0310] 2. Inherited from neighboring neighbors.
[0311] 3. Inherited from non-adjacent neighbors.
[0312] 4. Constructed from neighboring neighbors.
[0313] 5. First type of constructed affine candidate from non-adjacent neighbors.
[0314] 6. Zero MV.
[0315] Affine AMVP mode: 1. Inherited from neighboring neighbors.
[0316] 2. Constructed from neighboring neighbors.
[0317] 3. Translational MV from neighboring neighbors.
[0318] 4. Translational MV from temporal neighbors.
[0319] 5. Inheritance from non-adjacent neighbors.
[0320] 6. First type of constructed affine candidates from non-adjacent neighbors.
[0321] 7. Zero MV.
[0322] Due to the inclusion of additional candidates generated by NA-AFF, the size of the affine Merge candidate list increases from 5 to 15. The sub-group size of ARMC for affine Merge mode increases from 3 to 15.
[0323] Figure 43A and Figure 43B Spatial neighbors for derivation of affine Merge / AMVP candidates are shown: Figure 43A for derivation of inherited candidates, Figure 43B for derivation of first type of constructed candidates.
[0324] In NA-AFF: 1. The region from which non-adjacent neighbors come is restricted to be within the current CTU (i.e., no additional storage requirement for row buffer).
[0325] 2. The storage granularity for affine motion information (including CPMV and reference index) is reduced from 8x8 to 16x16 (i.e., only affine motion from the top-left 8x8 block is saved). Additionally, the saved CPMV is mapped to each 16x16 block before storage, so that position and size information is not needed.
[0326] 3. Only top-left and top-right CPMV are stored (i.e., always 4-parameter affine model is used for NA-AFF).
[0327] 2.30. Affine MMVD In affine MMVD, an affine Merge candidate (referred to as base affine Merge candidate) is selected, and the MVs of the control points are further refined by MVD information signaled.
[0328] The MVD information for MVs of all control points is the same in one prediction direction.
[0329] When the starting MV is bi-predicted MV, and the two MVs point to different sides of the current picture (i.e., one reference has a POC greater than the POC of the current picture and the other reference has a POC less than the POC of the current picture), the MV offset added to the list 0 MV component of the starting MV has opposite values from the MV offset of the list 1 MV; otherwise, when the starting MV is bi-predicted MV, and both lists point to the same side of the current picture (i.e., both references have a POC greater than the POC of the current picture or both have a POC less than the POC of the current picture), the MV offset added to the list 0 MV component of the starting MV has the same values as the MV offset of the list 1 MV.
[0330] 2.31. Adaptive decoder-side motion vector refinement (ADMVR) In ECM-2.0, a multi-pass decoder-side motion vector refinement (DMVR) method is applied in regular Merge mode if the selected Merge candidate satisfies the DMVR condition. In the first pass, bilateral matching (BM) is applied to the coded block. In the second pass, BM is applied to each 16x16 sub-block within the coded block. In the third pass, the MV in each 8x8 sub-block is refined by applying bi-directional optical flow (BDOF).
[0331] The adaptive decoder-side motion vector refinement method consists of two new Merge modes that are introduced for refining the MV in only one direction (L0 or L1) of bi-prediction of the selected Merge candidate that satisfies the DMVR condition. A multi-pass DMVR process is applied to the selected Merge candidate to refine the motion vector, however, in the first pass (i.e., PU level) DMVR, MVD0 or MVD1 is set to zero.
[0332] Similar to regular Merge mode, the Merge candidates of the proposed Merge modes are derived from spatial neighboring coded blocks, TMVP, non-adjacent blocks, HMVP, and paired candidates. The difference is that only those that satisfy the DMVR condition are added to the candidate list. The two proposed Merge modes use the same Merge candidate list (i.e., ADMVR Merge list), and the Merge index is coded as in regular Merge mode.
[0333] 3. Problem In the current design of template matching, the factor (i.e., 1.125) for the determination of whether unidirectional prediction or bi-directional prediction is used after motion vector refinement is constant, which can limit the coding performance.
[0334] A refined motion vector with lower cost (e.g., MV'0) is used to further refine another refined motion vector with larger cost (e.g., MV'1) to obtain a further refined motion vector (MV''1). The refinement can be performed in an iterative manner. 4. DETAILED DESCRIPTION The following specific solutions should be considered as examples to explain the general concept. The solutions should not be interpreted in a narrow way. Furthermore, the solutions can be combined in any way.
[0336] Iterative refinement for TM 1. It is proposed that during a motion refinement process, a first motion information (MI A ) is used to refine a second MI (MI B ).
[0337] a. In one example, the motion refinement process can refer to bilateral matching.
[0338] b. In one example, the motion refinement process can refer to template matching.
[0339] c. In one example, MI A or MI B may be predefined.
[0340] d. In one example, MI A or MI B may be in the same reference list.
[0341] i. Alternatively, MI A or MI B may be in different reference lists.
[0342] e. In one example, MI A or MI B may be in the same direction, such as both MI A and MI B have a picture order count (POC) smaller or larger than the POC of the current picture.
[0343] i. Alternatively, MI A or MI B may be in different directions, such as MI A has a POC smaller than the POC of the current picture while MI B has a POC larger than the current picture.
[0344] ii. In another example, MI A has a POC larger than the POC of the current picture while MI BPOC of the POC of the current picture.
[0345] f. In one example, the MI A may be used to determine the refinement process of the MI B . i. search range ii. starting search point iii. mode shape g. In one example, the MI A may be refined before being used to refine the MI B .
[0346] h. In one example, the MI A or the MI B may be refined more than once, and the i-th refined MI of the MI A and the MI B is denoted as MI A (ri) and the MI B (ri) .
[0347] i. In one example, iterative refinement can be used.
[0348] i. In one example, the i-th refined MI of the MI A (MI A (ri) ) can be used to obtain the j-th refined MI of the MI B (MI B (rj) ).
[0349] 1) In one example, i can be less than, equal to, or greater than j.
[0350] a) In one example, i = 1 and j = 1, or i = 2 and j = 1, or i = 1 and j = 2, or i = 2 and j = 2, or i = 0 and j = 0, i = 1 and j = 0, or i = 0 and j = 1.
[0351] 2) In one example, after the i-th refined MI of the MI A (MI A (ri) ) is used to obtain the j-th refined MI of the MI B (MI B (rj) ), the j-th refined MI of the MI B (MI B (rj)) can be used to obtain the MI A the (i+1)th refined MI (MI A (r(i+1)) ).
[0352] ii. Alternatively, the MI B the jth refined MI (MI B (rj) ) can be used to obtain the MI A the ith refined MI (MI A (ri) ).
[0353] iii. In one example, the template size / shape during the iterative refinement can be different from the template size / shape without the iterative refinement.
[0354] iv. In one example, the search range and / or the pattern shape (e.g., diamond, square, cross) used for the search can be different from the search range and / or the pattern shape without the iterative refinement.
[0355] v. In one example, the iterative refinement can be applied to specific coding tools.
[0356] 1) In one example, the coding tools can refer to adaptive reorder of merge candidates (ARMC), TMMerge mode, AMVP-Merge mode, or other coding tools in which bilateral matching and / or template matching is used to refine the motion information.
[0357] 2) In one example, how to use the iterative refinement can be different for different coding tools.
[0358] j. In one example, the determination of whether to use refinement or iterative refinement and / or how to use the refinement or iterative refinement can depend on the coding information.
[0359] i. In one example, the coding information refers to the bilateral matching cost and / or the template matching cost and / or the block size / dimension and / or the template size / dimension and / or the quantization parameter (QP) and / or the POC value of the current picture and / or the reference picture.
[0360] ii. In one example, when the template matching cost of bi-prediction (costBi) is greater than or equal to S1*costUni, the iterative refinement can be used, where costUni is equal to the template matching cost of uni-prediction (cost0 or cost1, or min(cost0, cost1), or max(cost0, cost1)).
[0361] iii. In one example, the refinement can be used when the POC difference between the current picture and its reference picture is less than or equal to T P .
[0362] 1) In one example, T P = 2, or 4, or 6, or 8, or 16.
[0363] iv. In one example, the refinement can be used when the refinement is used for bi-prediction, when the POC value of the first reference picture is less than the POC value of the current picture and the POC value of the second reference picture is greater than the POC value of the current picture.
[0364] v. In one example, the refinement can be used when the bilateral matching cost and / or the template matching cost is less than or equal to T C .
[0365] 1) In one example, T C may depend on the block size / dimension.
[0366] vi. Alternatively, the iterative refinement can be used always.
[0367] 2. It is proposed that the determination of the search range and / or the pattern shape used in the search in template matching can depend on the coding information.
[0368] a. In one example, the coding information can refer to the precision of the motion vector (MV) or the MV difference (MVD), or the syntax element indicating the MV or MVD.
[0369] b. In one example, the search range for a first MV precision can be larger than the search range for a second MV precision.
[0370] i. In one example, the first MV precision is integer and the second MV precision is fractional.
[0371] ii. In one example, the first MV precision is 1 / M and the second MV precision is 1 / N, where M is smaller than N.
[0372] c. In one example, the number of search points for a pattern shape for a first MV precision can be more than the number of search points for a pattern shape for a second MV precision.
[0373] i. In one example, the pattern shapes can be the same. Examples are shown in Figure 45 .
[0374] ii. In one example, the pattern shapes can be different. Examples are shown in Figure 46 .
[0375] d.In one example, the coding information can refer to the block size / dimension. Let the width and height of the block be denoted as W and H, respectively.
[0376] i.In one example, when W*H <= T1, a first search range and / or pattern shape can be used, and when W*H > T1, a second search range and / or pattern shape can be used, where T1 is an integer greater than 0.
[0377] 1) In one example, T1 = 64, or 128, or 256, or 512, or 1024.
[0378] ii.In one example, the first search pattern shape can refer to an 8-point search pattern, and the second search pattern shape can refer to a 16-point search pattern.
[0379] 3.Instead of a constant factor (S) such as described in section 2.22, it is proposed that at least one adaptive factor can be used for the determination of whether uni-prediction or bi-prediction is used.
[0380] a.In one example, when the template matching cost of bi-prediction (costBi) is less than or equal to S*the cost of uni-prediction (costUni), uni-prediction can be used, where costUni is equal to the template matching cost of uni-prediction (cost0 or cost1, or min(cost0, cost1), or max(cost0, cost1)).
[0381] b.In one example, the determination can be used for template matching.
[0382] c.In one example, the determination can be used for the current video unit.
[0383] d.In one example, more than one factor can be predefined or transmitted by signal or derived.
[0384] e.In one example, the determination of which factor to use can depend on the coding information.
[0385] i.In one example, the coding information can refer to POC or POC difference. Let the POC of the current picture, two reference pictures be denoted as poc0, poc1, poc2, respectively.
[0386] 1) In one example, when the POC difference (d) is less than or equal to T, a first factor (S1) can be used, and when the POC difference (d) is greater than T, a second factor (S2) can be used, where T is an integer greater than 1.
[0387] a) In one example, d = abs(poc0 - poc1) + abs(poc0 - poc2).
[0388] b) In one example, T = 2, or 3, or 4, or 5, or 6, or 7, or 8, or 10, or 16.
[0389] c) In one example, S1 is smaller than S, and S2 is equal to or larger than S, e.g., S = 1.125.
[0390] i. S1 = 1.115 and S2 = 1.125.
[0391] d) In one example, S1 is equal to or smaller than S, and S2 is larger than S, e.g., S = 1.125.
[0392] i. S1 = 1.125 and S2 = 1.135 ii. In one example, the coding information can refer to whether the two reference pictures are in the same direction.
[0393] 1) In one example, when one reference picture is in the forward direction and the other reference picture is in the backward direction, a first factor (S3) can be used, and when both reference pictures are in the forward direction or both are in the backward direction, a second factor (S4) can be used, where a reference picture in the forward direction means its POC is smaller than the POC of the current picture, and a reference picture in the backward direction means its POC is larger than the POC of the current picture and.
[0394] iii. In one example, the coding information can be: 1) whether a specific coding tool is allowed.
[0395] 2) the block dimension and / or the block size.
[0396] 3) the depth of the block.
[0397] 4) the slice / picture type and / or the partition tree type (single tree, or dual tree, or local dual tree).
[0398] 5) the block position.
[0399] 6) the color component.
[0400] 4. It is proposed that the first motion refinement is used as part of the second motion refinement.
[0401] a. In one example, the first motion refinement can refer to template matching, and the second motion refinement can refer to bilateral matching (e.g., DMVR / multi-pass DMVR / adaptive DMVR).
[0402] i. In one example, the first motion refinement can refer to template matching for bi-prediction and the second motion refinement can refer to bilateral matching (e.g., DMVR / multi-pass DMVR / adaptive DMVR).
[0403] 1) In one example, the following can be used for a specific coding tool, such as ARMC or TM Merge mode, where template matching for bi-prediction is used as part of bilateral matching.
[0404] b. In one example, the coding information from the first motion refinement can be used in the second motion refinement.
[0405] i. In one example, the coding information can refer to the cost calculated in the first motion refinement.
[0406] General aspects 5. In the above examples, the video unit can refer to a color component / sub-picture / slice / tile / coding tree unit (CTU) / CTU row / CTU group / coding unit (CU) / prediction unit (PU) / transform unit (TU) / coding tree block (CTB) / coding block (CB) / prediction block (PB) / transform block (TB) / block / sub-block of a block / sub-region within a block / any other region containing more than one sample or pixel.
[0407] 6. Whether and / or how to apply the above disclosed methods can be signaled at sequence level / picture group level / picture level / slice level / tile group level, such as in sequence header / picture header / SPS / VPS / DPS / DCI / PPS / APS / slice header / tile group header.
[0408] 7. Whether and / or how to apply the above disclosed methods can be signaled at PB / TB / CB / PU / TU / CU / VPDU / CTU / CTU row / slice / tile / sub-picture / other kind of region containing more than one sample or pixel.
[0409] 8. Whether and / or how to apply the above disclosed methods can depend on coded information, such as block size, color format, mono / bi-tree partitioning, color component, slice / picture type.
[0410] As used herein, the term “video unit” or “video block” can be a sequence, a picture, a slice, a tile, a brick, a subpicture, a coding tree unit (CTU) / coding tree block (CTB), a CTU / CTB row, one or more coding units (CU) / coding blocks (CB), one or more CTU / CTB, one or more virtual pipeline data units (VPDU), a sub-region within a picture / slice / tile / brick. In the following discussion, IntraTMP can be replaced by other coding tools that rely on coded / decoded / reconstructed information within the same region, such as palette, intra block copy (IBC).
[0411] Figure 47 A flowchart of a method 4700 for video processing according to an embodiment of the disclosure is shown. The method 4700 is implemented during conversion between a video unit of a video and a bitstream of the video.
[0412] At block 4710, for conversion between a video unit of a video and a bitstream of the video unit, first motion information and second motion information of the video unit are obtained. In some embodiments, the first motion information and the second motion information are in the same reference list. In some other embodiments, the first motion information and the second motion information are in different reference lists.
[0413] At block 4720, during a refinement process of the video unit, the second motion information is refined by using the first motion information. In some embodiments, the refinement process is bilateral matching. In some other embodiments, the motion refinement process is template matching. In some embodiments, the refinement process is refinement. Alternatively, the refinement process is iterative refinement.
[0414] At block 4730, the conversion is performed based on the refined first motion information and the second motion information. In some embodiments, the conversion can include encoding the video unit into the bitstream. Alternatively or additionally, the conversion can include decoding the video unit from the bitstream. In this way, it improves coding efficiency and coding performance.
[0415] In some embodiments, the iterative refinement is applied to a coding tool. For example, the coding tool can include at least one of the following: adaptive reorder of merge candidates (ARMC), template matching (TM) merge mode, advanced motion vector prediction (AMVP)-merge mode, other coding tools in which at least one of bilateral matching or template matching is used to refine motion information. In some other embodiments, the method of using iterative refinement is different for different coding tools.
[0416] In some embodiments, the determination of whether to use refinement or iterative refinement and / or how to use refinement or iterative refinement depends on coding information. For example, the coding information includes at least one of: a bilateral matching cost, a template matching cost, a block size, a block dimension, a template size, a template dimension, a quantization parameter (QP), a picture order count (POC) value of a current picture, or a POC value of a reference picture. In some other embodiments, iterative refinement is always used.
[0417] In some embodiments, if a template matching cost of bi-directional prediction (costBi) is greater than or equal to S1*costUni, then iterative refinement is used. In this case, costUni can be equal to a template matching cost of uni-directional prediction, and S1 can be a parameter. In some embodiments, costUni is equal to one of: costO, cost1, a minimum value between costO and cost1, or a maximum value between costO and cost1.
[0418] In some embodiments, if a POC difference between a current picture and its reference pictures is less than or equal to a first threshold, then refinement is used. For example, the first threshold is one of: 2, 4, 6, 8, or 16. In some other embodiments, if refinement is used for bi-directional prediction, and a first POC value of a first reference picture is less than a POC value of the current picture, and a second POC value of a second reference picture is greater than the POC value of the current picture, then refinement is used.
[0419] In some embodiments, if at least one of a bilateral matching cost or a template matching cost is less than or equal to a second threshold, then refinement is used. In some embodiments, the second threshold depends on a block size or a block dimension.
[0420] In some embodiments, a determination of at least one of a search range or a pattern shape used for searching in template matching depends on coding information. In some embodiments, the coding information includes a block size or a block dimension.
[0421] In some embodiments, if W*H <= T1, then at least one of a first search range or a first pattern shape is used. In addition, if W*H > T1, then at least one of a second search range or a second pattern shape is used. In this case, W represents a block width, H represents a block height, and T1 is an integer greater than 0. In some embodiments, T1 is equal to 64, or 128, or 256, or 512, or 1024. In some embodiments, the first search pattern shape is an 8-point search pattern, and the second search pattern shape is a 16-point search pattern.
[0422] In some embodiments, the first motion refinement is used as part of the second motion refinement, the first motion refinement is template matching for bi-prediction, and the second motion refinement includes bi-lateral matching. In some embodiments, the following is used for a coding tool: template matching for bi-prediction is used as part of bi-lateral matching. In some embodiments, the coding tool is ARMC or TM Merge mode.
[0423] In some embodiments, the video unit comprises at least one of: a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a tile, a subpicture, a block, a subblock of a block, a subregion within a block, or a region comprising more than one sample or pixel.
[0424] In some embodiments, the indication of whether and / or how the second motion information is refined during the refinement process of the video unit by using the first motion information is indicated at one of: a sequence level, a picture group level, a picture level, a slice level, or a tile group level.
[0425] In some embodiments, the indication of whether and / or how the second motion information is refined during the refinement process of the video unit by using the first motion information is indicated in one of: a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.
[0426] In some embodiments, the indication of whether and / or how the second motion information is refined during the refinement process of the video unit by using the first motion information is in one of: a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline data unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a subpicture, or a region comprising more than one sample or pixel.
[0427] In some embodiments, the method 4700 further includes determining, based on coded information of the video unit, whether and / or how to refine the second motion information using the first motion information during a refinement process of the video unit. The coded information can include at least one of: block size, color format, single and / or dual tree partitioning, color component, slice type, or picture type.
[0428] According to further embodiments of the disclosure, a non-transitory computer-readable recording medium is provided. The non-transitory computer-readable recording medium stores a bitstream generated by a method for a video processing. The method includes obtaining first motion information and second motion information of a video unit of a video; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is a refinement or an iterative refinement; and generating the bitstream based on the refined first motion information and second motion information.
[0429] According to still further embodiments of the disclosure, a method for storing a bitstream of a video is provided. The method includes obtaining first motion information and second motion information of a video unit of a video; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is a refinement or an iterative refinement; generating the bitstream based on the refined first motion information and second motion information; and storing the bitstream in a non-transitory computer-readable recording medium.
[0430] Embodiments of the disclosure can be described in view of the following clauses, which features can be combined in any reasonable manner.
[0431] Clause 1. A method of video processing, comprising: for a conversion between a video unit of a video and a bitstream of the video unit, obtaining first motion information and second motion information of the video unit; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is a refinement or an iterative refinement; and performing the conversion based on the refined first motion information and second motion information.
[0432] Clause 2. The method of clause 1, wherein the iterative refinement is applied to a coding tool.
[0433] Item 3. The method according to item 2, wherein the coding tool comprises at least one of: adaptive reordering of merge candidates (ARMC), template matching (TM) merge mode, advanced motion vector prediction (AMVP)-merge mode, other coding tool in which at least one of bilateral matching or template matching is used to refine motion information.
[0434] Item 4. The method according to item 2, wherein the method of using the iterative refinement is different for different coding tools.
[0435] Item 5. The method according to item 1, wherein the determination of whether to use the refinement or the iterative refinement and / or how to use the refinement or the iterative refinement depends on coding information.
[0436] Item 6. The method according to item 5, wherein the coding information comprises at least one of: bilateral matching cost, template matching cost, block size, block dimension, template size, template dimension, quantization parameter (QP), picture order count (POC) value of the current picture or POC value of the reference picture.
[0437] Item 7. The method according to item 5, wherein the iterative refinement is used if template matching cost of bi-directional prediction (costBi) is greater than or equal to S1*costUni, where costUni is equal to template matching cost of uni-directional prediction and S1 is a parameter.
[0438] Item 8. The method according to item 7, wherein the costUni is equal to one of: costO, cost1, minimum between costO and cost1 or maximum between costO and cost1.
[0439] Item 9. The method according to item 5, wherein the refinement is used if POC difference between the current picture and its reference picture is less than or equal to a first threshold.
[0440] Item 10. The method according to item 9, wherein the first threshold is one of: 2, 4, 6, 8 or 16.
[0441] Item 11. The method according to item 5, wherein the refinement is used if the refinement is used for bi-directional prediction and first POC value of a first reference picture is less than POC value of the current picture and second POC value of a second reference picture is greater than POC value of the current picture.
[0442] Item 12. The method of item 5, wherein the refinement is used if at least one of the bilateral matching cost or the template matching cost is less than or equal to a second threshold.
[0443] Item 13. The method of item 12, wherein the second threshold depends on a block size or a block dimension.
[0444] Item 14. The method of item 1, wherein the iterative refinement is used.
[0445] Item 15. The method of item 1, wherein a determination of at least one of a search range or a pattern shape for searching in template matching depends on coding information.
[0446] Item 16. The method of item 15, wherein the coding information comprises a block size or a block dimension.
[0447] Item 17. The method of item 16, wherein at least one of a first search range or a first pattern shape is used if W*H <= Tl, and at least one of a second search range or a second pattern shape is used if W*H > Tl, and wherein W denotes a block width, H denotes a block height, and Tl is an integer greater than 0.
[0448] Item 18. The method of item 17, wherein Tl is equal to 64, or 128, or 256, or 512, or 1024.
[0449] Item 19. The method of item 17, wherein the first search pattern shape is an 8-point search pattern, and the second search pattern shape is a 16-point search pattern.
[0450] Item 20. The method of item 1, wherein a first motion refinement is used as part of a second motion refinement, the first motion refinement is template matching for bi-prediction, and the second motion refinement comprises bilateral matching.
[0451] Item 21. The method of item 20, wherein the following is used for a coding tool: template matching for bi-prediction is used as part of bilateral matching.
[0452] Item 22. The method of item 20, wherein the coding tool is an ARM C or TM Merge mode.
[0453] Item 23. The method according to any of items 1-22, wherein the video unit comprises at least one of a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a tile, a subpicture, a block, a subblock of a block, a subregion within a block, or a region containing more than one sample or pixel.
[0454] Item 24. The method according to any of items 1-22, wherein an indication of whether and / or how the second motion information is refined during the refinement process of the video unit by using the first motion information is indicated at one of a sequence level, a picture group level, a picture level, a slice level, or a tile group level.
[0455] Item 25. The method according to any of items 1-22, wherein an indication of whether and / or how the second motion information is refined during the refinement process of the video unit by using the first motion information is indicated in one of a sequence header, a picture header, a sequence parameter set (SPS), a video parameter set (VPS), a dependent parameter set (DPS), a decoding capability information (DCI), a picture parameter set (PPS), an adaptation parameter set (APS), a slice header, or a tile group header.
[0456] Item 26. The method according to any of items 1-22, wherein an indication of whether and / or how the second motion information is refined during the refinement process of the video unit by using the first motion information is indicated in one of a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline data unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a subpicture, or a region containing more than one sample or pixel.
[0457] Item 27. The method of any of items 1-22, further comprising determining, based on coded information of the video unit, whether and / or how to refine the second motion information using the first motion information during the refinement process of the video unit, the coded information comprising at least one of: block size, color format, single and / or dual tree partitioning, color component, slice type, or picture type.
[0458] Item 28. The method of any of items 1-27, wherein the converting comprises encoding the video unit into the bitstream.
[0459] Item 29. The method of any of items 1-27, wherein the converting comprises decoding the video unit from the bitstream.
[0460] Item 30. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of items 1-29.
[0461] Item 31. A non-transitory computer-readable storage medium storing instructions to cause a processor to perform the method of any of items 1-29.
[0462] Item 32. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: obtaining first motion information and second motion information of a video unit of the video; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is refinement or iterative refinement; and generating the bitstream based on the refined first motion information and the second motion information.
[0463] Item 33. A method for storing a bitstream of a video, comprising: obtaining first motion information and second motion information of a video unit of the video; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is refinement or iterative refinement; generating the bitstream based on the refined first motion information and the second motion information; and storing the bitstream in a non-transitory computer-readable recording medium.
[0464] Example device Figure 48A block diagram of a computing device 4800 in which various embodiments of the present disclosure can be implemented is shown. The computing device 4800 can be implemented as the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300), or can be included in the source device 110 (or video encoder 114 or 200) or the destination device 120 (or video decoder 124 or 300).
[0465] It is to be understood that Figure 48 The computing device 4800 shown in FIG. 48 is for purposes of illustration and explanation only and is not intended as any limitation on the functionality and scope of the embodiments of the present disclosure.
[0466] As Figure 48 shown, the computing device 4800 includes a general-purpose computing device 4800. The computing device 4800 can include at least one or more processors or processing units 4810, a memory 4820, a storage unit 4830, one or more communication units 4840, one or more input devices 4850, and one or more output devices 4860.
[0467] In some embodiments, the computing device 4800 can be implemented as any user terminal or server terminal having computing capability. The server terminal can be a server provided by a service provider, a mainframe computing device, or the like. The user terminal may, for example, be any type of mobile terminal, fixed terminal, or portable terminal including a mobile telephone, a station, a unit, a device, a multimedia computer, a multimedia tablet, an Internet node, a communicator, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. It is contemplated that the computing device 4800 can support any type of interface to the user (such as "wearable" circuitry, etc.).
[0468] The processing unit 4810 can be a physical or virtual processor and can implement various processing based on programs stored in the memory 4820. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to increase the processing power of the computing device 4800. The processing unit 4810 can also be referred to as a central processing unit (CPU), a microprocessor, a controller, or a microcontroller.
[0469] Computing device 4800 typically includes a variety of computer storage media. Such media can be any media that is accessible by computing device 4800 and can include, without limitation, volatile and non-volatile media, removable and non-removable media. Memory 4820 can be volatile memory (such as registers, cache, random access memory (RAM)), non-volatile memory (such as read only memory (ROM), electrically erasable programmable read only memory (EEPROM) or flash memory), or any combination thereof. Storage 4830 can be any removable or non-removable media, and can include machine readable media such as memory, flash drives, disks, or other media that can be used to store information and / or data and that can be accessed by computing device 4800.
[0470] Computing device 4800 can also include additional removable / non-removable storage media, volatile / non-volatile memory. Although not shown in Figure 48 a disk drive for reading from and / or writing to a removable, non-removable, and / or nonvolatile media such as a disk, and an optical disk drive for reading from and / or writing to a removable, non-removable, and / or nonvolatile media such as an optical disk. In such instances, each drive can be connected to the bus (not shown) by one or more data media interfaces.
[0471] Communication unit 4840 enables communications with other computing devices via a communication medium. Additionally, the functionality of the components of computing device 4800 can be implemented by a single computing cluster or a plurality of computing machines that can communicate over a communication connection. Thus, computing device 4800 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or other general-purpose network nodes.
[0472] Input device 4850 can be one or more of various input devices such as a mouse, a keyboard, a trackball, a voice input device, and / or the like. Output device 4860 can be one or more of various output devices such as a display, a speaker, a printer, and / or the like. Computing device 4800 can also include communication unit 4840, which can enable computing device 4800 to communicate with one or more external devices such as a storage device or an display device, which can enable a user to interact with computing device 4800, or any devices (e.g., network card, modem, etc.) that enable computing device 4800 to communicate with one or more other computing devices. Such communication can be via input / output (I / O) interface (not shown).
[0473] In some embodiments, some or all of the components of computing device 4800 can also be arranged in a cloud computing architecture, rather than being integrated in a single device. In a cloud computing architecture, components can be provided remotely and work together to implement the functionality described in this disclosure. In some embodiments, cloud computing provides computation, software, data access, and storage services that do not require end-user knowledge of the physical location or configuration of the system that delivers the services. In various embodiments, cloud computing delivers services via the internet using appropriate protocols. For example, a cloud computing provider delivers applications via the internet from a remote location for use by users of client devices. The software or components of the cloud computing architecture, and corresponding data, can be stored on servers at remote locations. Computing resources in a cloud computing environment can be consolidated or distributed at locations remote from users. Cloud computing infrastructure can provide services through shared data centers, although they appear as a single point of access to users. Thus, a cloud computing architecture can be used to provide the components and functionality described herein from a service provider at a remote location. Alternatively, the components and functionality described herein can be provided by a conventional server, or installed directly on a client device, either directly or in other ways.
[0474] In embodiments of the disclosure, computing device 4800 can be used to implement video encoding / decoding. Memory 4820 can include one or more video coding modules 4825 having one or more program instructions. These modules are accessible to and executable by processing unit 4810 to perform the functions of the various embodiments described herein.
[0475] In example embodiments that perform video encoding, input device 4850 can receive video data as input 4870 to be encoded. The video data can be processed, for example, by video coding module 4825, to generate an encoded bitstream. The encoded bitstream can be provided as output 4880 via output device 4860.
[0476] In example embodiments that perform video decoding, input device 4850 can receive an encoded bitstream as input 4870. The encoded bitstream can be processed, for example, by video coding module 4825, to generate decoded video data. The decoded video data can be provided as output 4880 via output device 4860.
[0477] While the present disclosure has been particularly shown and described with reference to the preferred embodiments thereof, it will be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the application as defined by the appended claims. Such variations are intended to be covered by the scope of this application. Therefore, the foregoing description of embodiments of the application is not intended to be limiting.
Claims
1. A method of video processing, comprising: for a conversion between a video unit of a video and a bitstream of the video unit, obtaining first motion information and second motion information of the video unit; during a refinement process of the video unit, refining the second motion information by using the first motion information, wherein the refinement process is a refinement or an iterative refinement; and performing the conversion based on the refined first motion information and the second motion information.
2. The method of claim 1, wherein the iterative refinement is applied to a coding tool.
3. The method of claim 2, wherein the coding tool comprises at least one of the following: adaptive reordering of merge candidates (ARMC), template matching (TM) merge mode, advanced motion vector prediction (AMVP)-merge mode, other coding tool in which at least one of bilateral matching or template matching is used to refine motion information.
4. The method of claim 2, wherein the method of using the iterative refinement is different for different coding tools.
5. The method of claim 1, wherein a determination of whether to use the refinement or the iterative refinement and / or how to use the refinement or the iterative refinement depends on coding information.
6. The method of claim 5, wherein the coding information comprises at least one of the following: a bilateral matching cost, a template matching cost, a block size, a block dimension, a template size, a template dimension, a quantization parameter (QP), a picture order count (POC) value of a current picture or a POC value of a reference picture.
7. The method of claim 5, wherein the iterative refinement is used if a template matching cost of bi-directional prediction (costBi) is greater than or equal to S1*costUni, wherein costUni is equal to a template matching cost of uni-directional prediction, and S1 is a parameter.
8. The method of claim 7, wherein the costUni is equal to one of the following: cost0, cost1, a minimum value between cost0 and cost1, or a maximum value between cost0 and cost1.
9. The method of claim 5, wherein the refinement is used if a POC difference between the current picture and its reference picture is less than or equal to a first threshold.
10. The method of claim 9, wherein the first threshold is one of the following: 2, 4, 6, 8, or 16.
11. The method of claim 5, wherein the refinement is used if the refinement is used for bi-directional prediction, and a first POC value of a first reference picture is less than a POC value of the current picture, and a second POC value of a second reference picture is greater than the POC value of current picture.
12. The method of claim 5, wherein the refinement is used if at least one of a bilateral matching cost or a template matching cost is less than or equal to a second threshold. 13. The method of claim 12, wherein the second threshold depends on a block size or a block dimension.
14. The method of claim 1, wherein the iterative refinement is used.
15. The method of claim 1, wherein a determination of at least one of a search range or a pattern shape used for searching in template matching depends on coding information.
16. The method of claim 15, wherein the coding information comprises a block size or a block dimension.
17. The method of claim 16, wherein if W*H <= Tl, at least one of a first search range or a first pattern shape is used, and if W*H > Tl, at least one of a second search range or a second pattern shape is used, wherein W denotes a block width, H denotes a block height, and Tl is an integer greater than 0.
18. The method of claim 17, wherein Tl is equal to 64, or 128, or 256, or 512, or 1024.
19. The method of claim 17, wherein the first search pattern shape is an 8-point search pattern and the second search pattern shape is a 16-point search pattern.
20. The method of claim 1, wherein a first motion refinement is used as part of a second motion refinement, the first motion refinement is a template matching for bi-prediction, and the second motion refinement comprises bi-prediction.
21. The method of claim 20, wherein the following is used for a coding tool: the template matching for bi-prediction is used as part of the bi-prediction.
22. The method of claim 20, wherein the coding tool is an ARM C or TM Merge mode.
23. The method of any of claims 1-22, wherein the video unit comprises at least one of: a color component, a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding tree block (CTB), a coding unit (CU), a coding tree unit (CTU), a CTU row, a CTU group, a slice, a tile, a sub-picture, a block, a sub-block of a block, a sub-region within a block, or a region containing more than one sample or pixel.
24. The method of any of claims 1-22, wherein an indication of whether and / or how the second motion information is refined by using the first motion information during the refinement process of the video unit is indicated at one of: a sequence level, a picture group level, a picture level, a slice level, or a tile group level.
25. The method of any of claims 1-22, wherein an indication of whether and / or how the second motion information is refined by using the first motion information during the refinement process of the video unit is indicated in one of: a sequence header, picture header, sequence parameter set (SPS), video parameter set (VPS), dependency parameter set (DPS), decoding capability information (DCI), picture parameter set (PPS), adaptation parameter set (APS), slice header, or tile group header.
26. The method of any of claims 1-22, wherein an indication of whether and / or how to refine the second motion information using the first motion information during the refinement process of the video unit is in one of: a prediction block (PB), a transform block (TB), a coding block (CB), a prediction unit (PU), a transform unit (TU), a coding unit (CU), a virtual pipeline data unit (VPDU), a coding tree unit (CTU), a CTU row, a slice, a tile, a sub-picture, or a region containing more than one sample or pixel.
27. The method of any of claims 1-22, further comprising: determining, based on coded information of the video unit, whether and / or how to refine the second motion information using the first motion information during the refinement process of the video unit, the coded information including at least one of: a block size, a color format, a single and / or dual tree partitioning, a color component, a slice type, or a picture type.
28. The method of any of claims 1-27, wherein the conversion includes encoding the video unit into the bitstream.
29. The method of any of claims 1-27, wherein the conversion includes decoding the video unit from the bitstream.
30. An apparatus for video processing, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method of any of claims 1-29.
31. A non-transitory computer-readable storage medium storing instructions that cause a processor to perform the method of any of claims 1-29.
32. A non-transitory computer-readable recording medium storing a bitstream of a video generated by a method performed by an apparatus for video processing, wherein the method comprises: obtaining first motion information and second motion information of a video unit of the video; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is a refinement or an iterative refinement; and generating the bitstream based on the refined first motion information and the second motion information.
33. A method for storing a bitstream of a video, comprising: obtaining first motion information and second motion information of a video unit of the video; refining the second motion information using the first motion information during a refinement process of the video unit, wherein the refinement process is a refinement or an iterative refinement; and generating the bitstream based on the refined first motion information and the second motion information. generating the bitstream based on the refined first motion information and the second motion information; and storing the bitstream in a non-transitory computer-readable recording medium.