Video processing method, device, medium and method for storing video bit stream

By adjusting the MVD export process in MMVD mode, inserting non-near airspace merge candidates and using STMVP candidates, optimizing the merge mode of video encoding and decoding, solving the problem of low encoding and decoding efficiency in small blocks and improving the overall encoding and decoding efficiency.

CN115136597BActive Publication Date: 2025-08-26DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080089538.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-23
Filing Date
2020-12-23
Publication Date
2025-08-26
Estimated Expiration
2040-12-23

AI Technical Summary

Technical Problem

The existing video encoding and decoding technology has problems with inefficiency in merge mode, especially in the case of small blocks in MMVD mode, and the application of non-near merge candidates and STMVP is insufficient, resulting in low encoding and decoding efficiency.

Method used

By adjusting the MVD export process according to the block dimension and prediction direction in MMVD mode, inserting non-near airspace merge candidates and using STMVP candidates, the construction process of merge candidate list is optimized, including airspace-time domain motion vector prediction, and improving encoding and decoding efficiency.

Benefits of technology

Improve the efficiency of video encoding and decoding, especially in small blocks, and enhance the effectiveness and encoding and decoding efficiency of merge mode.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115136597B_ABST
    Figure CN115136597B_ABST
Patent Text Reader

Abstract

Spatial-temporal motion vector prediction is described. An example video processing method includes: for conversion between a current block of a video and a bitstream representation of the video, constructing a merge candidate list for the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the merge candidate list, and deriving the STMVP candidate as an average candidate of M spatially neighboring motion candidates and / or N temporally neighboring motion candidates, where M and N are positive integers; and performing the conversion based on the merge candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application is intended to claim priority to and the benefit of International Patent Application No. PCT / CN2019 / 127388, filed on December 23, 2019, in a timely manner, under applicable patent law and / or under the rules of the Paris Convention. The entire disclosure of International Patent Application No. PCT / CN2019 / 127388 is incorporated by reference as a part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Art

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communications networks. As the number of connected user devices capable of receiving and displaying video increases, bandwidth demand for digital video usage is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to perform cross-component adaptive loop filtering during video encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed that includes converting between a video unit of a video and a codec representation of the video, determining a motion vector difference (MVD) calculation operation for use with a merge mode with motion vector difference (MMVD) codec based on characteristics of the video unit, and performing the conversion based on the determination.

[0007] In another example aspect, a video processing method is disclosed that includes performing a conversion between video units of a video and a codec representation of the video, wherein the conversion utilizes a motion vector scaling process that is dependent on a resolution of the video during operation.

[0008] In another example aspect, a video processing method is disclosed, comprising generating a merge candidate list for a conversion between a video unit of a video and a codec representation of the video, wherein non-adjacent spatial merge candidates for the video unit are inserted into the merge candidate list, and performing the conversion using the merge candidate list.

[0009] In another example aspect, a video processing method is disclosed, comprising: generating a candidate list for conversion between a video unit of a video and a codec representation of the video, wherein candidates in the candidate list are generated by averaging M spatially neighboring candidates and N temporally neighboring candidates, where M and N are positive integers; and performing the conversion using a merged candidate list.

[0010] In another example aspect, a video processing method is disclosed, comprising generating a merge list for conversion between a video unit of a video and a codec representation of the video, wherein a construction process for generating the merge list examines a plurality of candidates in a defined order, and performing the conversion using the merge candidate list.

[0011] In another example aspect, a video processing method is disclosed. The method includes performing conversion between a video unit of a video and a codec representation of the video using two long-term reference pictures and a motion vector scaling process.

[0012] In another example aspect, a video processing method is disclosed, comprising: for conversion between video units of a video and a bitstream of the video, deriving a motion vector difference (MVD) for use in a merge mode with motion vector difference (MMVD) codec based on characteristics of the video units; and performing the conversion based on the derived MVD.

[0013] In another example aspect, a video processing method is disclosed. The method includes: deriving a motion vector difference (MVD) using a motion vector (MV) scaling process for conversion between video units of a video and a bitstream of the video, wherein the MV scaling process depends on a resolution of the video; and performing the conversion based on the derived MVD.

[0014] In another example aspect, a video processing method is disclosed. The method includes: for conversion between a video unit of a video and a bitstream of the video, deriving a motion vector difference (MVD) using a motion vector (MV) scaling process, wherein the MV scaling process uses two long-term reference pictures; and performing the conversion based on the derived MVD.

[0015] In another example aspect, a method for storing a bitstream of a video is disclosed. The method includes: for conversion between video units of the video and the bitstream of the video, deriving a motion vector difference (MVD) for use in a merge mode with motion vector difference (MMVD) codec based on characteristics of the video units; generating a bitstream from the video units based on the derived MVD; and storing the bitstream in a non-transitory computer-readable recording medium.

[0016] In another example aspect, a video processing method is disclosed, comprising: constructing a merge candidate list for a current block of video for conversion between the current block and a bitstream representation of the video, wherein non-neighboring spatial merge candidates associated with the current block are inserted into the merge candidate list; and performing the conversion based on the merge candidate list.

[0017] In another example aspect, a video processing method is disclosed, comprising: constructing a merge candidate list for a current block of video for conversion between the current block and a bitstream representation of the video, wherein the construction of the merge candidate list examines a plurality of different types of candidates in a defined order; and performing the conversion based on the merge candidate list.

[0018] In another example aspect, a method for storing a bitstream of a video is disclosed. The method includes: constructing a merge candidate list for a current block of the video for conversion between the current block and a bitstream representation of the video, wherein non-neighboring spatial merge candidates associated with the current block are inserted into the merge candidate list; generating a bitstream from a video unit based on the merge candidate list; and storing the bitstream in a non-transitory computer-readable recording medium.

[0019] In another example aspect, a video processing method is disclosed. The method includes: constructing a merge candidate list for a conversion between a current block of a video and a bitstream representation of the video, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and / or N temporal neighboring motion candidates, where M and N are positive integers; and performing the conversion based on the merge candidate list.

[0020] In another example aspect, a method for storing a bitstream of a video is disclosed. The method includes: constructing a merge candidate list for a current block of the video for conversion between the current block of the video and a bitstream representation of the video, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and / or N temporal neighboring motion candidates, where M and N are positive integers; generating a bitstream from a video unit based on the merge candidate list; and storing the bitstream in a non-transitory computer-readable recording medium.

[0021] In yet another exemplary aspect, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the above method.

[0022] In yet another exemplary aspect, a video decoder apparatus is disclosed, wherein the video decoder includes a processor configured to implement the above method.

[0023] In yet another exemplary aspect, a computer-readable medium having code stored thereon is disclosed. The code is in the form of processor-executable code embodying one of the methods described herein.

[0024] In yet another example aspect, a computer-readable medium stores a bitstream of a video generated by the above method performed by a video processing device.

[0025] These and other features are described throughout this document. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Examples of offsets added to the horizontal or vertical component of the starting motion vector (MV) are shown.

[0027] Figure 2 The HEVC spatial neighboring blocks of the current block are shown.

[0028] Figure 3 The relationship between the virtual block and the current block is shown.

[0029] Figure 4 is a block diagram of an example video processing system in which the disclosed technology may be implemented.

[0030] Figure 5 is a block diagram of an example hardware platform for video processing.

[0031] Figure 6 is a flow chart of an example method of video processing.

[0032] Figure 7 is a block diagram illustrating a video encoding and decoding system according to some embodiments of the present disclosure.

[0033] Figure 8 is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0034] Figure 9 is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0035] Figure 10 is a flow chart of an example method of video processing.

[0036] Figure 11 is a flow chart of an example method of video processing.

[0037] Figure 12 is a flow chart of an example method of video processing.

[0038] Figure 13 is a flow chart of an example method for storing a bitstream of video.

[0039] Figure 14 is a flow chart of an example method of video processing.

[0040] Figure 15 is a flow chart of an example method of video processing.

[0041] Figure 16 is a flow chart of an example method for storing a bitstream of video.

[0042] Figure 17 is a flow chart of an example method of video processing.

[0043] Figure 18 is a flow chart of an example method for storing a bitstream of video. DETAILED DESCRIPTION

[0044] The section headings used in this document are intended to facilitate understanding and are not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is intended to facilitate understanding and is not intended to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. 1. Summary of the Invention

[0046] This patent document relates to video codec technology. Specifically, it relates to a merge mode in video codecs. It can be applied to existing video codec standards, such as HEVC, or to a future standard (Multi-Function Video Codec). It can also be applied to future video codec standards or video codecs. 2. Background Technology

[0048] Video codec standards have evolved primarily through the development of the well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, ISO / IEC produced MPEG-1 and MPEG-4 Visual, and the two organizations jointly produced the H.262 / MPEG-2 Video, H.264 / MPEG-4 Advanced Video Codec (AVC), and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture that utilizes temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1SC29 / WG11 (MPEG) was created to develop the VVC standard with the goal of a 50% bitrate reduction compared to HEVC.

[0049] 2.1. Merge Mode with MVD (MMVD)

[0050] In addition to the merge mode in which the implicitly derived motion information is directly used for the prediction sample generation of the current CU, VVC also introduces the merge mode with motion vector difference (MMVD), also known as the ultimate motion vector expression. The MMVD flag is signaled immediately after the skip flag and merge flag are sent to specify whether the MMVD mode is used for the CU.

[0051] In MMVD, a merge candidate (called a base merge candidate) is selected, which is further refined by the signaled MVD information. The relevant syntax elements include an index used to specify the MVD distance (represented by mmvd_distance_idx) and an index used to indicate the direction of motion (represented by mmvd_direction_idx). In MMVD mode, one of the first two candidates in the merge list is selected as the MV basis (or base merge candidate). The signaling merge candidate flag specifies which candidate to use.

[0052] The distance index specifies the motion magnitude information and indicates a predefined offset from the starting point. Figure 1 As shown, the offset is added to the horizontal component or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 3.

[0053] Table 3: Relationship between distance index and predefined offset

[0054]

[0055]

[0056] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 4. It should be noted that the meaning of the MVD symbol can change according to the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV in which both lists point to the same side of the current picture (i.e., the POCs of both references are greater than the POC of the current picture, or both are less than the POC of the current picture), the symbol in Table 4 specifies the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV in which the two MVs point to different sides of the current picture (i.e., the POC of one reference is greater than the POC of the current picture, and the POC of the other reference is less than the POC of the current picture), the symbol in Table 4 specifies the sign of the MV offset added to the list 0 MV component of the starting MV, while the symbol of list 1 MV has the opposite value.

[0057] Table 4: Sign of MV offset specified by direction index

[0058] Direction IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –

[0059] 2.2.1 Derivation of MVD for each reference picture list

[0060] First, an internal MVD (denoted by MmvdOffset) is derived according to the index of the decoded MVD distance (denoted by mmvd_distance_idx) and the index of the motion direction (denoted by mmvd_direction_idx).

[0061] Afterwards, if the internal MVD is determined, the final MVD of the basic merge candidate to be added to each reference picture list is further derived according to the POC distance of the reference picture relative to the current picture and the reference picture type (long-term or short-term). More specifically, the following steps are performed in sequence:

[0062] If the base merge candidate is bi-directionally predicted, the POC distances between the current picture and the reference pictures in list 0 and the POC distances between the current picture and the reference pictures in list 1 are calculated, denoted by POCDiffL0 and POCDidffL1, respectively.

[0063] If POCDiffL0 is equal to POCDidffL1, the final MVD of both reference picture lists is set to intra MVD.

[0064] —Otherwise, if Abs(POCDiffL0) is greater than or equal to Abs(POCDiffL1), the final MVD of reference picture list 0 is set to intra MVD, and the final MVD of reference picture list 1 is set to scaled MVD or to intra MVD or (zero MV minus intra MVD) using the intra MVD reference picture type of both reference pictures (neither of which is a long-term reference picture) according to the POC distance.

[0065] Otherwise, if Abs(POCDiffL0) is less than Abs(POCDiffL1), the final MVD of reference picture list 1 is set to the internal MVD, and the final MVD of reference picture list 0 is set to the scaled MVD or the internal MVD or (zero MV minus the internal MVD) using the internal MVD reference picture type of both reference pictures (neither of which is a long-term reference picture) according to the POC distance.

[0066] If the base merge candidate is unidirectionally predicted from reference picture list X, the final MVD of reference picture list X is set to the intra MVD, and the final MVD of reference picture list Y (Y=1-X) is set to 0.

[0067] 2.2.2MMVD Specification in VVC

[0068] The MMVD specification (in JVET-P2001-vE) is as follows:

[0069] 7.3.9.7merge data syntax

[0070]

[0071]

[0072] mmvd_merge_flag[x0][y0] equal to 1 specifies that merge mode with motion vector difference is used to generate inter prediction parameters for the current codec. mmvd_merge_flag[x0][y0] equal to 0 indicates that merge mode with motion vector difference is not used to generate inter prediction parameters. Array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture.

[0073] When mmvd_merge_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0074] mmvd_cand_flag[x0][y0] specifies whether the first (0) or the second (1) candidate in the merge candidate list is used with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The array index x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture.

[0075] When mmvd_cand_flag[x0][y0] is not present, it is inferred to be equal to 0.

[0076] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0] as specified in Table 17. The array index x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture.

[0077] Table 17 – Specification of MmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0].

[0078]

[0079]

[0080] mmvd_direction_idx[x0][y0] specifies the index used to derive MmvdSign[x0][y0], as specified in Table 18. The array index x0, y0 specifies the position (x0, y0) of the top-left luma sample of the considered codec block relative to the top-left luma sample of the picture.

[0081] Table 18 - Specification of MmvdSign[x0][y0] based on mmvd_direction_idx[x0][y0]

[0082]

[0083] The two components of merge plus the MVD offset MmvdOffset[x0][y0] are derived as follows:

[0084] MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][0](181)

[0085] MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSign[x0][y0][1](182)

[0086] 8.5.2.7 Merge Motion Vector Difference Derivation Process

[0087] The input to this process is:

[0088] - the luminance position (xCb, yCb) of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture,

[0089] — reference indices refIdxL0 and refIdxL1,

[0090] - The prediction list utilizes the flags predFlagL0 and predFlagL1.

[0091] The output of this process is the luminance merge motion vector differences mMvdL0 and mMvdL1 with 1 / 16 fractional sample accuracy.

[0092] The variable currPic specifies the current picture.

[0093] The luminance merge motion vector differences mMvdL0 and mMvdL1 are derived as follows:

[0094] — If both predFlagL0 and predFlagL1 are equal to 1, the following applies:

[0095] currPocDiffL0=DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0])(564)

[0096] currPocDiffL1=DiffPicOrderCnt(currPic,RefPicList[1][refIdxL1])(565)

[0097] — If currPocDiffL0 is equal to currPocDiffL1, the following applies:

[0098] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (566)

[0099] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (567)

[0100] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (568)

[0101] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (569)

[0102] — Otherwise, if Abs(currPocDiffL0) is greater than or equal to Abs(currPocDiffL1), then the following applies:

[0103] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (570)

[0104] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (571)

[0105] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, the following applies:

[0106] td=Clip3(-128,127,currPocDiffL0) (572)

[0107] tb=Clip3(-128,127,currPocDiffL1) (573)

[0108] tx=(16384+(Abs(td)>>1)) / td (574)

[0109] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (575)

[0110]

[0111] — Otherwise, the following applies:

[0112] mMvdL1[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL0[0]:-mMvdL0[0] (578)

[0113] mMvdL1[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL0[1]:-mMvdL0[1] (579)

[0114] — Otherwise (Abs(currPocDiffL0) is less than Abs(currPocDiffL1)), the following applies:

[0115] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (580)

[0116] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (581)

[0117] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, the following applies:

[0118] td=Clip3(-128,127,currPocDiffL1) (582)

[0119] tb=Clip3(-128,127,currPocDiffL0) (583)

[0120] tx=(16384+(Abs(td)>>1)) / td (584)

[0121] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (585)

[0122]

[0123] — Otherwise, the following applies:

[0124] mMvdL0[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL1[0]:-mMvdL1[0] (588)

[0125] mMvdL0[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL1[1]:-mMvdL1[1] (589)

[0126] Otherwise (predFlagL0 or predFlagL1 is equal to 1), the following applies for X equal to 0 and 1:

[0127] mMvdLX[0]=(predFlagLX==1)? MmvdOffset[xCb][yCb][0]:0(590)

[0128] mMvdLX[1]=(predFlagLX==1)? MmvdOffset[xCb][yCb][1]:0(591)

[0129] 2.2.JVET-L0323: Long-distance merge candidate

[0130] In HEVC, Figure 2 The five spatially neighboring blocks and one temporal neighbor shown in are used to derive merge candidates.

[0131] Figure 2 The HEVC spatial neighboring blocks of the current block are shown.

[0132] This contribution proposes to derive additional merge candidates from non-adjacent positions of the current block using the same pattern as in HEVC. To this end, for each search round i, a virtual block is generated based on the current block as follows:

[0133] First, the relative position of the virtual block and the current block is calculated as follows:

[0134] Offsetx=-i*gridX,Offsety=-i*gridY

[0135] Where Offsetx and Offsety represent the offset of the upper left corner of the virtual block relative to the upper left corner of the current block, and gridX and gridY are the width and height of the search grid.

[0136] Second, the width and height of the virtual block are calculated as follows:

[0137] newWidth=i*2*gridX+currWidth newHeight=i*2*gridY+currHeight.

[0138] Where currWidth and currHeight are the width and height of the current block, and newWidth and newHeight are the width and height of the new block.

[0139] gridX and gridY are currently set to currWidth and currHeight respectively.

[0140] Figure 3 The relationship between the virtual block and the current block is shown.

[0141] After generating the virtual block, block A i 、B i 、C i 、D i and E i The HEVC spatial neighboring blocks can be considered as virtual blocks, and their positions are obtained in the same way as in HEVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i 、B i 、C i 、D i and E i It is the spatial neighboring block used in HEVCmerge mode.

[0142] When constructing the merge candidate list, pruning is performed to ensure that each element in the merge candidate list is unique. As more and more blocks are examined to derive additional merge candidates, the number of pruning increases accordingly. To limit the number of pruning in the worst case, the maximum number of pruning allowed in the merge list construction is constrained to a predefined value MaxPruningNum.

[0143] In the simulation, the maximum search rounds is set to 2 and the MaxPruningNum is set to 30.

[0144] Long-distance merge candidates are also called non-adjacent merge candidates.

[0145] Figure 3 is a diagram of the virtual blocks in the i-th search round.

[0146] 2.3.JVET-M0059: Non-scaling STMVP

[0147] The proposed method uses two spatial merge candidates and one collocated merge candidate to derive the average candidate as the STMVP candidate.

[0148] The STMVP is inserted before the upper left spatial merge candidate.

[0149] For spatial candidates, the first and second candidates in the current merge candidate list are used.

[0150] For temporal candidates, the same positions as VTM / HEVC co-location are used.

[0151] If three candidates with reference equal to 0 are available, the following applies.

[0152] mvLX[0]=(mvLX_A[0]*3+mvLX_B[0]*3+mvLX_C[0]*2) / 8

[0153] mvLX[1]=(mvLX_A[1]*3+mvLX_B[1]*3+mvLX_C[1]*2) / 8

[0154] If two motion information with reference equal to zero are available, the following applies

[0155] mvLX[0]=(mvLX_A[0]+mvLX_C[0]) / 2

[0156] mvLX[1]=(mvLX_A[1]+mvLX_C[1]) / 2

[0157] or

[0158] mvLX[0]=(mvLX_B[0]+mvLX_C[0]) / 2

[0159] mvLX[1]=(mvLX_B[1]+mvLX_C[1]) / 2

[0160] NOTE: If time domain candidates are not available, STMVP mode is disabled.

[0161] MMVD is also known as Ultimate Motion Vector Expression (UMVE).

[0162] 3. Technical Problems Solved by the Technical Solutions and Examples in This Article

[0163] The current design of merge mode can be further improved.

[0164] 1. In MMVD mode, for small blocks (e.g., 4x8 / 8x4), even if only unidirectional prediction is allowed, two MVDs can still be derived if the base merge candidate is bidirectionally predicted. More specifically, if the selected MV base (or base merge candidate) is bidirectional MV, the MVD for the prediction direction from one reference list X (X=0 or 1) is directly set equal to the signaled MVD, while the MVD for the other reference list Y (Y=1–X) is derived based on the MVD for prediction direction X and the POC (Picture Order Count) distance, so scaling is required in some cases. However, in VTM-7.0, bidirectional prediction is prohibited for 4x8 / 8x4 blocks. Therefore, there is no need to derive the MVD for L1.

[0165] 2. In addition, non-adjacent merge candidates and / or STMVP can be used to improve the effectiveness of merge mode. In addition, encoding and decoding efficiency can be improved.

[0166] 4. Example Embodiments and Techniques

[0167] The following items should be considered as examples to explain the general concept. These items should not be interpreted narrowly. In addition, these items can be combined in any way.

[0168] In MMVD, the internal MVD is derived from syntax elements signaled in the bitstream (such as MVD distance and direction information), and the final MVD is the MVD used to refine the base merge candidate, that is, the MVD used to derive the final MV of the block.

[0169] In the following, currWidth and currHeight are the width and height of the current block (e.g., luma block). maxNumMergeCand represents the merge list size.

[0170] As shown in Section 2.3, after generating the virtual block, block A i 、B i 、C i 、D i and E i The HEVC spatial neighboring blocks can be considered as virtual blocks, and their positions are obtained in the same way as HEVC. Obviously, if the search round i is 0, the virtual block is the current block. In this case, block A i 、B i 、C i 、D i and E i It is the spatial neighboring block used in HEVCmerge mode.

[0171] For spatial candidates, the first, second, and third candidates in the current merge candidate list inserted before STMVP are denoted as F, S, and T, respectively.

[0172] For temporal candidates at the same position as VTM / HEVC, the co-located position used in STMVP is denoted as Col.

[0173] MVD export of MMVD

[0174] 1. How to derive the MVD used in the MMVD method may depend on the block dimensions and / or the allowed prediction directions (eg, whether only unidirectional prediction is allowed for a video unit (eg, CU / PU)).

[0175] a. In one example, if only unidirectional prediction is allowed for a video unit, only one MVD is derived from the intra MVD instead of two MVDs, regardless of the prediction direction associated with the base merge candidate in the MMVD.

[0176] i. In one example, if only prediction from reference picture list X is the prediction direction, denoted by ListX (eg, X is 0), then the final MVD of ListX is derived from the internal MVD.

[0177] (i) Alternatively, furthermore, the final MVD of ListX is set equal to the inner MVD.

[0178] (ii) Alternatively, furthermore, the final MVD of ListX is set equal to the inverse of the inner MVD.

[0179] (iii) Alternatively, furthermore, the final MVD of ListY is set to a default value, eg, zero MVD.

[0180] b. In one example, if certain condition(s) depending on the block dimension are met, only one MVD is derived from the internal MVD instead of two MVDs, regardless of the prediction direction associated with the base merge candidate in the MMVD.

[0181] i. In one example, the condition is that currWidth+currHeight is less than or equal to N (N is a positive integer). For example, N=12.

[0182] ii. In one example, the condition is that currWidth*currHeight is less than or equal to N (N is a positive integer), for example, N=32.

[0183] iii. In one example, the condition is currWidth < N1 or / and currHeight < N2 (N1, N2 are positive integers). For example, N1 = N2 = 8.

[0184] iv. In one example, the condition is currWidth < N3 * currHeight and / or currHeight < N4 * currWidth (N3, N4 are positive integers). For example, N3 = N4 = 8.

[0185] 2. It is recommended that in MMVD, when the base merge candidate is a bi - directional MV, if the block dimension or block shape satisfies one or more conditions, the internal MVD can always be directly used (e.g., without scaling) to predict the direction X (X = 0, 1).

[0186] a. In one example, the internal MVD is always directly used to predict the direction 0.

[0187] b. In one example, the condition is that currWidth + currHeight is less than or equal to N (N is a positive integer). For example, N = 12.

[0188] c. In one example, the condition is that currWidth * currHeight is less than or equal to N (N is a positive integer). For example, N = 32.

[0189] d. In one example, the condition is currWidth < N1 or / and currHeight < N2 (N1, N2 are positive integers). For example, N1 = N2 = 8.

[0190] e. In one example, the condition is currWidth < N3 * currHeight and / or currHeight < N4 * currWidth (N3, N4 are positive integers). For example, N3 = N4 = 8.

[0191] f. If the block dimension or block shape satisfies one or more conditions, the opposite value of the internal MVD (-MVD) can be used to replace the MVD to predict the direction X (X = 0, 1).

[0192] 3. The MV scaling process (e.g., those used in MMVD, TMVP, etc.) can take the picture resolution into consideration.

[0193] 4. For two reference pictures that are both long - term reference pictures, the MV scaling process can still be applied. a. In one example, the MV scaling process can be similar to the case where the two reference pictures are short - term reference pictures, i.e., depending on the POC distance.

[0194] Non-adjacent merge candidates

[0195] 5. Non-adjacent airspace merge candidates can be inserted into the merge list.

[0196] a. In one example, non-adjacent spatial domain merge candidates are inserted into the merge list after the history-based merge candidates.

[0197] b. In one example, non-adjacent spatial merge candidates are inserted into the merge list after pairwise averaging of merge candidates.

[0198] c. In one example, if the number of available merge candidates in the merge list reaches a predefined value after inserting the temporal domain merge candidate, the non-adjacent spatial domain merge candidate may not be inserted.

[0199] d. In one example, if the number of available merge candidates in the merge list reaches a predefined value when inserting non-adjacent spatial merge candidates, the insertion process will be terminated.

[0200] e. In one example, the predefined value is equal to maxNumMergeCand–N.

[0201] i. In one example, N is set equal to 1, 2, 3, or 4.

[0202] f. In one example, the maximum search round is set to be equal to 1 or 2, that is, five or ten non-adjacent spatial merge candidates can be used to build the merge list.

[0203] g. In one example, for each search round, the insertion order is A i 、B i 、C i 、D i and E i .

[0204] i. Alternatively, for each search round, the insertion order is B i 、A i 、C i 、D i and E i .

[0205] ii. Alternatively, for each search round, the insertion order is B i 、C i 、A i 、D i and E i .

[0206] iii. Alternatively, for each search round, the insertion order is A i 、D i 、B i 、C i and E i .

[0207] h. In one example, all spatial and temporal merge candidates are fully pruned against all previous merge candidates in the merge list. The pruning process for history-based merge candidates and pairwise average candidates remains unchanged.

[0208] i. Alternatively, all spatial, temporal, history-based, and pairwise average merge candidates perform full pruning on all previous merge candidates in the merge list.

[0209] j. Alternatively, for non-adjacent spatial domain merge candidates, A i Use A i-1 Perform pruning, B i Use A i Perform pruning, C i Use B i Perform pruning, D i Use A i Perform pruning, E i Use A i and Bi perform pruning. The temporal, history-based, and pairwise average candidate pruning processes remain unchanged.

[0210] k. In one example, the maximum number of prunings MaxPruningNum allowed in the merge list construction may depend on the merge list size maxNumMergeCand.

[0211] i. For example, MaxPruningNum may be set equal to maxNumMergeCand−M (M is an integer). For example, M=2.

[0212] ii. For example, MaxPruningNum may be set equal to maxNumMergeCand*M (M is an integer), for example, M=2.

[0213] iii. Alternatively, MaxPruningNum is independent of the merge list size maxNumMergeCand. For example, MaxPruningNum can be set to 30 or 35.

[0214] In one example, non-adjacent spatial merge candidate locations are constrained to be within a predefined area.

[0215] i. In one example, the region may include the current CTU row and the four sample rows above the current CTU row.

[0216] ii. In one example, the region may include the current CTU column and the four left sample columns of the current CTU.

[0217] iii. In one example, the region may include the current CTU column and the left CTU column of the current CTU.

[0218] m. In one example, non-adjacent spatial domain merge candidate positions have no constraints in the horizontal direction.

[0219] n. In one example, non-adjacent spatial domain merge candidates can be used as the basic merge candidates for MMVD.

[0220] i. Alternatively, non-adjacent spatial merge candidates are not allowed to be used as base merge candidates for MMVD.

[0221] o. In one example, non-adjacent spatial merge candidates can be used to generate inter-intra predictions.

[0222] i. Alternatively, non-neighboring spatial merge candidates are not allowed to generate inter-intra predictions.

[0223] p. In one example, non-adjacent spatial domain merge candidates can be used to generate geometric (GEO) segmentation and / or triangle segmentation merge candidates.

[0224] i. Alternatively, non-adjacent spatial merge candidates are not allowed to generate geometric (GEO) partitioning and / or triangle partitioning merge candidates.

[0225] q. In one example, non-adjacent spatial merge candidates can be used to generate affine merge candidates.

[0226] In one example, non-neighboring spatial merge candidates may be used to generate Advanced Motion Vector Prediction (AMVP) candidates.

[0227] Spatial-Temporal Motion Vector Prediction (STMVP)

[0228] 6. The STMVP candidate can be derived as the average candidate of M spatial neighboring motion candidates and / or N temporal neighboring motion candidates.

[0229] a. In one example, M>2.

[0230] b. In one example, spatial neighboring motion candidates may be derived from other neighboring blocks that are different from or the same as those used for the merge list construction process.

[0231] c. In one example, a spatially adjacent motion candidate may be selected from the spatial merge candidates included in the merge list.

[0232] d. In one example, the spatial neighboring motion candidate may be selected from the first M or last M spatial merge candidates included in the merge list before adding the STMVP.

[0233] e. In one example, spatially adjacent motion candidates may be selected from the first M or last M merge candidates included in the merge list before adding the STMVP.

[0234] f. In one example, temporal neighboring motion candidates may be selected from temporal merge candidates.

[0235] i. In one example, if the time-domain merge candidate is unavailable, the STMVP candidate is considered unavailable.

[0236] g. In one example, whether a spatial neighboring motion candidate and / or a temporal neighboring motion candidate is considered valid is based on reference picture information.

[0237] i. In one example, only if its reference index in at least one reference picture list is equal to or not greater than K (eg, K=0).

[0238] ii. In one example, only when its reference index in both reference picture lists is equal to or not greater than K (eg, K=0).

[0239] iii. Alternatively, further, when it is deemed invalid, it is not used to derive STMVP candidates.

[0240] iv. Alternatively, further, if at least one candidate among the first M spatial merge candidates and a co-located merge candidate is valid, the STMVP candidate is valid.

[0241] h. In one example, M is set equal to 3 and N is set equal to 1.

[0242] i. In one example, if the reference indexes of the four merge candidates are all valid and equal to 0 in the prediction direction X (X=0 or 1), the motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:

[0243] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e

[0244] (i) In one example, a, b, c, d, and e are set equal to 1, 1, 1, 1, and 2.

[0245] ii. In one example, if the reference indexes of three of the four merge candidates are valid and equal to 0 in the prediction direction X (X=0 or 1), the motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:

[0246] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_Col*c)>>d or

[0247] mvLX=(mvLX_F*a+mvLX_T*b+mvLX_Col*c)>>d or

[0248] mvLX=(mvLX_S*a+mvLX_T*b+mvLX_Col*c)>>d

[0249] (i) In one example, a, b, c, and d are set equal to 3, 3, 2, and 3.

[0250] (ii) In one example, a, b, c, and d are set equal to 2, 2, 4, and 3.

[0251] (iii) In one example, a, b, c, and d are set equal to 1, 1, 6, and 3.

[0252] iii. In one example, if the reference indexes of two of the four merge candidates are valid and equal to 0 in the prediction direction X (X=0 or 1), the motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:

[0253] mvLX=(mvLX_F*a+mvLX_Col*b)>>c or

[0254] mvLX=(mvLX_S*a+mvLX_Col*b)>>c or

[0255] mvLX=(mvLX_T*a+mvLX_Col*b)>>c

[0256] (i) In one example, a, b, and c are set equal to 1, 1, and 1

[0257] i. In one example, the STMVP candidates may be pruned with all previous merge candidates in the merge list.

[0258] j. In one example, STMVP candidates may not be pruned with other merge candidates.

[0259] k. In one example, the STMVP candidate can be pruned with only the merge candidates above and to the left.

[0260] 1. In one example, the STMVP candidate refers to one or two specific reference pictures.

[0261] i. For example, a specific reference picture is a reference picture with a reference index equal to 0 in the reference list.

[0262] ii. For example, the specific reference picture is a reference picture of the M spatial neighboring motion candidates and / or the N temporal neighboring motion candidates having the smallest reference index in the reference list.

[0263] Merge list construction process

[0264] 7. The merge list building process may include checking the following candidates in order.

[0265] a. The first set of spatial merge candidates (e.g., derived from B, A, C, D), STMVP, the second set of spatial merge candidates (e.g., derived from E), TMVP, HMVP, pairwise average merge candidates, and zero motion vector merge candidates.

[0266] b. Spatial merge candidates derived from neighboring blocks (for example, derived from B, A, C, D, E), TMVP, the first set of spatial merge candidates derived from non-neighboring blocks (for example, derived from B1, A1, C1, D1, E1), HMVP, pairwise average merge candidates, zero motion vector merge candidates.

[0267] c. Spatial merge candidates derived from neighboring blocks (for example, derived from B, A, C, D, E), TMVP, spatial merge candidates derived from non-neighboring blocks (for example, derived from B1, A1, C1, D1, E1, B2, A2, C2, D2, E2), HMVP, pairwise average merge candidate, zero motion vector merge candidate.

[0268] d. The first group of spatial merge candidates (for example, derived from B, A, C, D), STMVP, the second group of spatial merge candidates (for example, derived from E), TMVP, the first group of spatial merge candidates derived from non-adjacent blocks (for example, derived from B1, A1, C1, D1, E1), HMVP, pairwise average merge candidate, zero motion vector merge candidate.

[0269] e. In the above example, if the corresponding candidate is unavailable, invalid, or identical or similar to an existing candidate (added before the corresponding candidate), the corresponding candidate is not included in the motion candidate list.

[0270] 8. The above method can be applied to other types of motion candidate lists besides the merge candidate list.

[0271] Alternatively, the above method can be applied to the block vector candidate list construction process for IBC coded blocks. In this case, the check whether the block is coded in IBC mode can be used instead of checking whether the reference picture index is equal to K.

[0272] 5. Examples

[0273] The deleted parts are gray The newly added parts are highlighted in grey.

[0274] 5.1. Example 1 of MMVD

[0275] If the MV basis (or base merge candidate) selected in MMVD mode is a bidirectional MV, and the sum of the width and height of the block is less than or equal to 12, the MVD for prediction direction 0 (L0) is directly set equal to the signaled MVD, and the MMVDmerge candidate is converted to an L0 unidirectional prediction candidate.

[0276] 8.5.2.7 Merge Motion Vector Difference Derivation Process

[0277] The input to this process is:

[0278] - the luminance position (xCb, yCb) of the upper left sample of the current luminance codec block relative to the upper left luminance sample of the current picture,

[0279] — reference indices refIdxL0 and refIdxL1,

[0280] - The prediction list utilizes the flags predFlagL0 and predFlagL1.

[0281] The output of this process is the luminance merge motion vector differences mMvdL0 and mMvdL1 with 1 / 16 fractional sample accuracy.

[0282] The variable currPic specifies the current picture.

[0283] The luminance merge motion vector differences mMvdL0 and mMvdL1 are derived as follows:

[0284] - If both predFlagL0 and predFlagL1 are equal to 1 and the current luma codec block

[0285] If the sum of width and height is greater than 12, the following applies:

[0286] currPocDiffL0=DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0])(564)

[0287] currPocDiffL1=DiffPicOrderCnt(currPic,RefPicList[1][refIdxL1])(565)

[0288] — If currPocDiffL0 is equal to currPocDiffL1, the following applies:

[0289] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (566)

[0290] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (567)

[0291] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (568)

[0292] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (569)

[0293] — Otherwise, if Abs(currPocDiffL0) is greater than or equal to Abs(currPocDiffL1), then the following applies:

[0294] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (570)

[0295] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (571)

[0296] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, the following applies:

[0297] td=Clip3(-128,127,currPocDiffL0) (572)

[0298] tb=Clip3(-128,127,currPocDiffL1) (573)

[0299] tx=(16384+(Abs(td)>>1)) / td (574)

[0300] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (575)

[0301]

[0302] — Otherwise, the following applies:

[0303] mMvdL1[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL0[0]:-mMvdL0[0] (578)

[0304] mMvdL1[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL0[1]:-mMvdL0[1] (579)

[0305] — Otherwise (Abs(currPocDiffL0) is less than Abs(currPocDiffL1)), the following applies:

[0306] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (580)

[0307] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (581)

[0308] If RefPicList[0][refIdxL0] is not a long-term reference picture and RefPicList[1][refIdxL1] is not a long-term reference picture, the following applies:

[0309] td=Clip3(-128,127,currPocDiffL1) (582)

[0310] tb=Clip3(-128,127,currPocDiffL0) (583)

[0311] tx=(16384+(Abs(td)>>1)) / td (584)

[0312] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (585)

[0313]

[0314]

[0315] — Otherwise, the following applies:

[0316] mMvdL0[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL1[0]:-mMvdL1[0] (588)

[0317] mMvdL0[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)? mMvdL1[1]:-mMvdL1[1] (589)

[0318] Otherwise (predFlagL0 or predFlagL1 is equal to 1 or the sum of the width and height of the current luma codec block is less than or equal to 12), then for X 0 and 1 the following applies:

[0319] mMvdLX[0]=(predFlagLX==1)? MmvdOffset[xCb][yCb][0]:0(590)

[0320] mMvdLX[1]=(predFlagLX==1)? MmvdOffset[xCb][yCb][1]:0(591)

[0321] Figure 4 1 is a block diagram illustrating an example video processing system 1900 in which the various techniques disclosed herein may be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, a passive optical network (PON), and wireless interfaces such as Wi-Fi or a cellular interface.

[0322] System 1900 may include a codec component 1904 that can implement the various codecs or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via a connected communication, as represented by component 1906. The bitstream (or codec) representation of the stored or transmitted video received at input 1902 can be used by component 1908 to generate pixel values ​​or a displayable video that is sent to display interface 1910. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. In addition, although some video processing operations are referred to as "codec" operations or tools, it should be understood that the codec tools or operations are used at the encoder, while the corresponding decoding tools or operations that reverse the codec results will be performed by the decoder.

[0323] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High-Definition Multimedia Interface (HDMI), or DisplayPort, etc. Examples of storage interfaces include SATA (Serial Advanced Technology Attachment), PCI, IDE interfaces, etc. The technology described in this document may be embodied in various electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0324] Figure 5 is a block diagram of a video processing device 3600. Device 3600 can be used to implement one or more methods described herein. Device 3600 can be embodied in a smartphone, tablet, computer, Internet of Things (IoT) receiver, etc. Device 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. Processor(s) 3602 can be configured to implement one or more methods described herein. Memory(s) 3604 can be used to store data and code used to implement the methods and techniques described herein. Video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0325] Figure 7 is a block diagram illustrating an example video coding system 100 that may utilize the techniques of this disclosure.

[0326] like Figure 7As shown, the video encoding and decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 may decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0327] Source device 110 may include a video source 112 , a video encoder 114 , and an input / output (I / O) interface 116 .

[0328] The video source 112 may include, for example, a video capture device, an interface for receiving video data from a video content provider, and / or a source of a computer graphics system for generating video data, or a combination of such sources. The video data may include one or more pictures. The video encoder 114 encodes the video data from the video source 112 to generate a bitstream. The bitstream may include a sequence of bits that form a codec representation of the video data. The bitstream may include a codec picture and associated data. The codec picture is a codec representation of the picture. The associated data may include a sequence parameter set, a picture parameter set, and other syntax structures. The I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. The encoded video data may be transmitted directly to the target device 120 via the network 130a via the I / O interface 116. The encoded video data may also be stored on a storage medium / server 130b for access by the target device 120.

[0329] Target device 120 may include an I / O interface 126 , a video decoder 124 , and a display device 122 .

[0330] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may obtain encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120, or may be external to target device 120 configured to interface with an external display device.

[0331] The video encoder 114 and the video decoder 124 may operate according to a video compression standard, such as the High Efficiency Video Codec (HEVC) standard, the Versatile Video Codec (VVC) standard, and other current and / or additional standards.

[0332] Figure 8 is a block diagram illustrating an example of a video encoder 200, which may be Figure 7 The video encoder 114 in the system 100 is shown in FIG.

[0333] Video encoder 200 may be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example of , video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video encoder 200. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0334] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213 and an entropy coding unit 214, and the prediction unit may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205 and an intra-frame prediction unit 206.

[0335] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra block copy (IBC) unit. The IBC unit may perform prediction in IBC mode, where at least one reference picture is the picture in which the current video block is located.

[0336] Furthermore, some components such as the motion estimation unit 204 and the motion compensation unit 205 may be highly integrated but are shown for the purpose of explanation. Figure 8 In the example, they are represented separately.

[0337] The partitioning unit 201 may partition a picture into one or more video blocks. The video encoder 200 and the video decoder 300 may support various video block sizes.

[0338] The mode selection unit 203 can, for example, select one of the codec modes (intra or inter) based on the error result, and provide the resulting intra or inter codec block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the codec block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra and inter prediction (CIIP) mode, in which prediction is based on an inter prediction signal and an intra prediction signal. In the case of inter prediction, the mode selection unit 203 can also select a resolution of motion vectors for the block (e.g., sub-pixel or integer pixel precision).

[0339] To perform inter-frame prediction on the current video block, the motion estimation unit 204 may generate motion information of the current video block by comparing the current video block with one or more reference frames from the buffer 213. The motion compensation unit 205 may determine a predicted video block for the current video block based on the motion information and decoded samples of pictures other than the picture associated with the current video block from the buffer 213.

[0340] Motion estimation unit 204 and motion compensation unit 205 may perform different operations on the current video block, eg, depending on whether the current video block is in an I slice, a P slice, or a B slice.

[0341] In some examples, motion estimation unit 204 may perform unidirectional prediction for the current video block, and motion estimation unit 204 may search the reference pictures in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 may then generate a reference index indicating the reference picture in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 may output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 may generate a predicted video block for the current block based on the reference video block indicated by the motion information for the current video block.

[0342] In other examples, the motion estimation unit 204 may perform bidirectional prediction for the current video block. The motion estimation unit 204 may search for a reference video block for the current video block in the reference pictures in list 0 and may also search for another reference video block for the current video block in the reference pictures in list 1. The motion estimation unit 204 may then generate reference indexes indicating the reference pictures containing the reference video blocks in list 0 and list 1, and a motion vector indicating the spatial displacement between the reference video blocks and the current video block. The motion estimation unit 204 may output the reference index and motion vector of the current video block as motion information for the current video block. The motion compensation unit 205 may generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0343] In some examples, motion estimation unit 204 may output a complete set of motion information for use in a decoding process by a decoder.

[0344] In some examples, motion estimation unit 204 may not output a complete set of motion information for the current video. Instead, motion estimation unit 204 may reference motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of the neighboring video block.

[0345] In one example, motion estimation unit 204 may indicate a value in a syntax structure associated with the current video block that indicates to video decoder 300 that the current video block has the same motion information as another video block.

[0346] In another example, the motion estimation unit 204 may identify another video block and a motion vector difference (MVD) in a syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. The video decoder 300 may use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0347] As discussed above, the video encoder 200 may predictively signal motion vectors.Two examples of predictive signaling techniques that may be implemented by the video encoder 200 include advanced motion vector prediction (AMVP) and merge mode signaling.

[0348] The intra-frame prediction unit 206 can perform intra-frame prediction on the current video block. When the intra-frame prediction unit 206 performs intra-frame prediction on the current video block, the intra-frame prediction unit 206 can generate prediction data for the current video block based on decoded samples of other video blocks in the same picture. The prediction data for the current video block may include a predicted video block and various syntax elements.

[0349] The residual generation unit 207 may generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) the predicted video block(s) of the current video block from the current video block. The residual data for the current video block may include residual video blocks corresponding to different sample components of the samples in the current video block.

[0350] In other examples, there may be no residual data for the current video block, such as in skip mode, and the residual generation unit 207 may not perform the subtraction operation.

[0351] Transform processing unit 208 may generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0352] After transform processing unit 208 generates a transform coefficient video block associated with the current video block, quantization unit 209 may quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0353] The inverse quantization unit 210 and the inverse transform unit 211 may apply inverse quantization and inverse transform, respectively, to the transform coefficient video block to reconstruct a residual video block from the transform coefficient video block. The reconstruction unit 212 may add the reconstructed residual video block to corresponding samples from one or more predicted video blocks generated by the prediction unit 202 to generate a reconstructed video block associated with the current block for storage in the buffer 213.

[0354] After the reconstruction unit 212 reconstructs the video block, a loop filtering operation may be performed to reduce video block artifacts in the video block.

[0355] The entropy coding unit 214 may receive data from other functional components of the video encoder 200. When the entropy coding unit 214 receives the data, the entropy coding unit 214 may perform one or more entropy coding operations to generate entropy-coded data and output a bitstream including the entropy-coded data.

[0356] Figure 9 is a block diagram illustrating an example of a video decoder 300, which may be Figure 7 The video decoder 114 in the system 100 is shown in FIG.

[0357] Video decoder 300 may be configured to perform any or all of the techniques of this disclosure. Figure 8 In the example of FIG, video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of video decoder 300. In some examples, a processor can be configured to perform any or all of the techniques described in this disclosure.

[0358] exist Figure 9 In the example of FIG, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform the same operations as those performed with respect to the video encoder 200 ( Figure 8 ) The encoding pass described is essentially the opposite of the decoding pass.

[0359] The entropy decoding unit 301 can retrieve an encoded bitstream. The encoded bitstream may include entropy-encoded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-encoded video data, and based on the entropy-encoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference picture list index, and other motion information. For example, the motion compensation unit 302 can determine such information by performing AMVP and merge modes.

[0360] The motion compensation unit 302 may generate a motion compensated block, possibly performing interpolation based on an interpolation filter. An identifier of the interpolation filter used with sub-pixel precision may be included in a syntax element.

[0361] The motion compensation unit 302 may calculate interpolated values ​​of sub-integer pixels of a reference block using interpolation filters used by the video encoder 200 during encoding of the video block. The motion compensation unit 302 may determine the interpolation filters used by the video encoder 200 based on received syntax information and use the interpolation filters to generate a prediction block.

[0362] The motion compensation unit 302 can use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, partitioning information describing how each macroblock of the pictures of the encoded video sequence is partitioned, a mode indicating how to encode each partition, one or more reference frames (and reference frame lists) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0363] The intra prediction unit 303 can form a prediction block from spatially neighboring blocks using, for example, an intra prediction mode received in the bitstream. The inverse quantization unit 303 inversely quantizes, i.e., dequantizes, the quantized video block coefficients provided in the bitstream and decoded by the entropy decoding unit 301. The inverse transform unit 303 applies an inverse transform.

[0364] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to the decoded block to remove blocking artifacts. The decoded video block is then stored in a buffer 307, which provides reference blocks for subsequent motion compensation / intra prediction and also produces decoded video for presentation on a display device.

[0365] The following provides a list of solutions preferred by some embodiments.

[0366] The following solution shows an example embodiment of the technique discussed in the previous section (eg, item 1).

[0367] 1. A video processing method (e.g., Figure 6 ), comprising: for converting between a video unit of a video and a codec representation of the video, determining (602) a motion vector difference (MVD) calculation operation for use with a merge mode with motion vector difference (MMVD) codec tool based on characteristics of the video unit, and performing (604) the conversion based on the determination.

[0368] 2. The method according to solution 1, wherein the characteristics of the video unit include dimensions of the video unit or allowed prediction directions.

[0369] 3. The method according to solution 1-2, wherein the characteristic of the video unit is that only unidirectional prediction is allowed, and based on this characteristic, the calculation operation determines a single MVD regardless of the prediction direction of the base merge candidate associated with the MMVD codec.

[0370] 4. The method of solution 2, wherein the computation operation determines a single MVD regardless of the prediction direction of the base merge candidate associated with the MMVD codec due to the dimensions of the video unit satisfying the condition.

[0371] The following solution shows an example embodiment of the technique discussed in the previous section (eg, item 2).

[0372] 5. A method according to solution 1, wherein the characteristic of the video unit is that the basic merge candidate associated with the MMVD codec is a bidirectional motion vector, and the conversion includes directly using the internal motion vector difference for prediction direction X, where X=0 or 1, if the video unit meets the block size condition.

[0373] 6. The method according to solution 5, wherein the internal motion vector difference is directly used for prediction direction 0.

[0374] 7. The method according to solution 5, wherein the block size condition is that the height plus the width of the video unit is less than or equal to N, where N is a positive integer less than the total picture pixel width.

[0375] The following solution shows an example embodiment of the technique discussed in the previous section (eg, item 3).

[0376] 8. A video processing method, comprising:

[0377] Conversion between video units of a video and a codec representation of the video is performed, wherein the conversion uses a motion vector scaling process that depends on a resolution of the video during operation.

[0378] 9. The method of solution 8, wherein the operation comprises encoding or decoding using a merge tool with motion vector differences.

[0379] 10. The method of solution 8, wherein the operations include encoding or decoding using a temporal motion vector prediction process.

[0380] The following solution shows an example embodiment of the technique discussed in the previous section (eg, item 4).

[0381] 11. A video processing method, comprising:

[0382] The conversion is performed using two long-term reference pictures and a motion vector scaling process during conversion between video units of a video and a codec representation of the video.

[0383] 12. The method of solution 11, wherein the motion vector scaling process is based on picture order counts of two long-term reference pictures.

[0384] The following solution shows an example embodiment of the technique discussed in the previous section (e.g., item 5).

[0385] 13. A video processing method, comprising:

[0386] generating a merge candidate list for conversion between a video unit of the video and a codec representation of the video, wherein non-adjacent spatial merge candidates of the video unit are inserted into the merge candidate list; and

[0387] Perform the transformation using the merge candidate list.

[0388] 14. The method of solution 13, wherein non-adjacent spatial domain merge candidates are inserted into the merge list after the history-based merge candidates.

[0389] 15. The method according to solution 13, wherein non-adjacent spatial domain merge candidates are inserted into the merge list after pairwise averaging of merge candidates.

[0390] 16. The method according to solution 13, wherein after inserting the time domain merge candidate, after the number of available merge candidates in the merge list reaches a predefined value, no non-adjacent spatial domain merge candidate is inserted.

[0391] 17. The method of solution 13, wherein non-adjacent spatial merge candidates are inserted in a defined order.

[0392] The following solution shows an example embodiment of the technique discussed in the previous section (eg, item 6).

[0393] 18. A video processing method, comprising:

[0394] generating a candidate list for conversion between a video unit of a video and a codec representation of the video, wherein candidates in the candidate list are generated by averaging M spatial neighboring candidates and N temporal neighboring candidates, where M and N are positive integers; and

[0395] Perform the transformation using the merge candidate list.

[0396] 19. The method according to solution 18, wherein M>2.

[0397] 20. The method according to solutions 18-19, wherein M spatial neighbor candidates can be derived from the spatial merge candidates included in the merge list.

[0398] 21. The method according to solutions 18-19, wherein the N time-domain neighboring candidates can be derived from the time-domain merge candidates included in the merge list.

[0399] 22. The method of solution 18, wherein M=3 and N=1.

[0400] The following solution shows an example embodiment of the technique discussed in the previous section (e.g., item 7).

[0401] 23. A video processing method, comprising:

[0402] generating a merge list for conversion between a video unit of a video and a codec representation of the video, wherein a construction process for generating the merge list examines a plurality of candidates in a defined order; and

[0403] Perform the transformation using the merge candidate list.

[0404] 24. A method according to solution 23, wherein the defined order includes: a first group of spatial merge candidates, spatial-temporal merge candidates, a second group of merge candidates, a temporal motion vector predictor, a history-based motion vector predictor, a pairwise average merge candidate vector, and a zero motion vector merge candidate.

[0405] 25. A method according to solution 23, wherein the defined order includes: spatial merge candidates derived from neighboring blocks (e.g., derived from B, A, C, D, E), TMVP, a first group of spatial merge candidates derived from non-neighboring blocks (e.g., derived from B1, A1, C1, D1, E1), HMVP, pairwise average merge candidates, and zero motion vector merge candidates.

[0406] 26. The method according to any one of solutions 1-25, wherein the video unit comprises a video block or a video codec tree unit or a video transform unit or a video codec unit.

[0407] 27. The method of any of solutions 1-26, wherein performing the conversion comprises encoding the video to generate a codec representation.

[0408] 28. The method of any of solutions 1-26, wherein performing the conversion comprises parsing and decoding the codec representation to generate the video.

[0409] 29. A video decoding device comprising a processor configured to implement the method according to one or more of solutions 1 to 28.

[0410] 30. A video encoding apparatus comprising a processor configured to implement the method according to one or more of solutions 1 to 28.

[0411] 31. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method according to any one of solutions 1 to 28.

[0412] 32. The methods, apparatus, or systems described in this document.

[0413] Figure 10 A flow chart of an example method for video processing is shown. The method includes: for conversion between video units of a video and a bitstream of the video, deriving (1002) a motion vector difference (MVD) for use in a mode with motion vector differences (MMVD) codec based on characteristics of the video units; and performing (1004) conversion based on the derived MVD.

[0414] In some examples, the characteristics of the video unit include dimensions or shape of the video unit and / or allowed prediction directions of the video unit, and the dimensions of the video unit include currWidth and currHeight parameters indicating the width and height of the video unit, respectively.

[0415] In some examples, when characteristics of a video unit indicate that only unidirectional prediction is allowed for the video unit, a single MVD is derived from an internal MVD associated with an MMVD codec, regardless of the prediction direction associated with a base merge candidate in the MMVD codec, where the internal MVD is derived from syntax elements signaled in the bitstream.

[0416] In some examples, if only prediction from reference picture list X is an allowed prediction direction, denoted by ListX, then the MVD of ListX is derived from the internal MVD, and X is 0 or 1.

[0417] In some examples, the MVD of ListX is set equal to the internal MVD.

[0418] In some examples, the MVD of ListX is set equal to the inverse of the inner MVD.

[0419] In some examples, the MVD used for prediction from the reference picture list Y represented by ListY is set to a default value, where Y is 1 or 0.

[0420] In some examples, when the characteristics of a video unit indicate that the dimensions of the video unit meet a condition, a single MVD is derived from the internal MVD associated with the MMVD codec tool, regardless of the prediction direction associated with the base merge candidate in the MMVD codec tool, where the internal MVD is an MVD derived from a syntax element signaled in the bitstream.

[0421] In some examples, the condition is that currWidth + currHeight is less than or equal to N, where N is a positive integer.

[0422] In some examples, N = 12, or N = 32.

[0423] In some examples, the condition is that currWidth < N1 or / and currHeight < N2, where N1 and N2 are positive integers.

[0424] In some examples, N1 = N2 = 8.

[0425] In some examples, the condition is that currWidth < N3 * currHeight and / or currHeight < N4 * currWidth, where N3 and N4 are positive integers.

[0426] In some examples, N3 = N4 = 8.

[0427] In some examples, if the characteristics of a video unit indicate that the dimensions or shape of the video unit meet one or more conditions, then when the base merge candidate is a bi - directional MV, the internal MVD is always directly used for prediction direction X without scaling, where X = 0 or 1.

[0428] In some examples, the internal MVD is always directly used for prediction direction 0.

[0429] In some examples, the condition is that currWidth + currHeight is less than or equal to N, where N is a positive integer.

[0430] In some examples, N = 12 or N = 32.

[0431] In some examples, the condition is that currWidth < N3 * currHeight and / or currHeight < N4 * currWidth, where N3 and N4 are positive integers.

[0432] In some examples, N3 = N4 = 8.

[0433] In some examples, the opposite value of the inner MVD is used to predict direction X without scaling, where X=0 or 1.

[0434] Figure 11 A flow chart of an example method for video processing is shown. The method includes: deriving (1102) a motion vector difference (MVD) for conversion between video units of a video and a bitstream of the video using a motion vector (MV) scaling process, wherein the MV scaling process depends on the resolution of the video; and performing (1104) the conversion based on the derived MVD.

[0435] In some examples, MVD is used in a merge mode with motion vector difference (MMVD) codec or a temporal motion vector prediction (TMVP) codec.

[0436] Figure 12 A flow chart of an example method for video processing is shown. The method includes: deriving (1202) a motion vector difference (MVD) for conversion between a video unit of a video and a bitstream of the video using a motion vector (MV) scaling process, wherein the MV scaling process uses two long-term reference pictures; and performing (1204) the conversion based on the derived MVD.

[0437] In some examples, a video unit includes at least one of a codec unit (CU), a prediction unit (PU), or a block of video.

[0438] In some examples, converting includes encoding video units of the video into a bitstream.

[0439] In some examples, converting includes decoding video units of the video from the bitstream.

[0440] Figure 13 A flowchart of an example method for storing a bitstream of a video is shown. The method includes: for conversion between video units of a video and a bitstream of the video, deriving (1302) a motion vector difference (MVD) for use in a merge mode with motion vector difference (MMVD) codec based on characteristics of the video units; generating (1304) a bitstream from the video units based on the derived MVD; and storing (1306) the bitstream in a non-transitory computer-readable recording medium.

[0441] Figure 14 A flowchart of an example method for video processing is shown. The method includes: for conversion between a current block of video and a bitstream representation of the video, constructing (1402) a merge candidate list for the current block, wherein non-adjacent spatial merge candidates associated with the current block are inserted into the merge candidate list; and performing (1404) the conversion based on the merge candidate list.

[0442] In some examples, non-adjacent spatial domain merge candidates are inserted into the merge list after the history-based merge candidates.

[0443] In some examples, non-adjacent spatial merge candidates are inserted into the merge list after pairwise averaging of merge candidates.

[0444] In some examples, if the number of available merge candidates in the merge candidate list reaches a predefined value after inserting the temporal merge candidate, the non-adjacent spatial merge candidate is not inserted.

[0445] In some examples, when inserting non-adjacent spatial merge candidates, if the number of available merge candidates in the merge candidate list reaches a predefined value, the process is terminated.

[0446] In some examples, the predefined value is equal to maxNumMergeCand-N, where maxNumMergeCand represents the size of the merge candidate list and N is a positive integer.

[0447] In some examples, N is set equal to 2, 3, or 4.

[0448] In some examples, the maximum number of search rounds during construction of the merge candidate list is set equal to 1 or 2.

[0449] In some examples, for each round of search, non-adjacent spatial merge candidates are inserted in a predefined insertion order, and the non-adjacent spatial merge candidates include candidate Ai derived from the non-adjacent left block of the current block, candidate Bi derived from the non-adjacent above block of the current block, candidate Ci derived from the non-adjacent upper right block of the current block, candidate Di derived from the non-adjacent lower left block n of the current block, and candidate Ei derived from the non-adjacent upper left block to the left of the current block, where i is the search round.

[0450] In some examples, the insertion order is Ai, Bi, Ci, Di, and Ei, the insertion order is Bi, Ai, Ci, Di, and Ei, the insertion order is Bi, Ci, Ai, Di, and Ei, or the insertion order is Ai, Di, Bi, Ci, and Ei.

[0451] In some examples, all spatial and temporal merge candidates undergo a full pruning process with all previous merge candidates in the merge candidate list, and the pruning process for history-based merge candidates and pairwise average candidates is unchanged.

[0452] In some examples, all spatial, temporal, history-based, and pairwise average merge candidates perform a full pruning process with all previous merge candidates in the merge candidate list.

[0453] In some examples, for non-adjacent spatial merge candidates, Ai is pruned with Ai-1, Bi is pruned with Ai, Ci is pruned with Bi, Di is pruned with Ai, Ei is pruned with Ai and Bi, and the pruning process for time domain, history-based and pairwise average candidates remains unchanged.

[0454] In some examples, the maximum number of prunings allowed during construction of the merge candidate list is denoted as MaxPruningNum, which depends on the size of the merge candidate list, denoted as maxNumMergeCand.

[0455] In some examples, MaxPruningNum is set equal to maxNumMergeCand−M, where M is an integer.

[0456] In some examples, M=2.

[0457] In some examples, MaxPruningNum is set equal to maxNumMergeCand*M, where M is an integer.

[0458] In some examples, M=2.

[0459] In some examples, the maximum number of prunings allowed during construction of the merge candidate list is denoted as MaxPruningNum, which is independent of the size of the merge candidate list, denoted as maxNumMergeCand.

[0460] In some examples, MaxPruningNum is set equal to 30 or 35.

[0461] In some examples, the locations of non-adjacent spatial merge candidates are constrained to be within a predefined area.

[0462] In some examples, the predefined area includes the current codec tree unit (CTU) row and four sample rows above the current CTU row.

[0463] In some examples, the predefined area includes the current CTU column and the four left sample columns of the current CTU column.

[0464] In some examples, the predefined area includes the current CTU column and the CTU column to the left of the current CTU column.

[0465] In some examples, the positions of non-adjacent spatial merge candidates are not constrained in the horizontal direction.

[0466] In some examples, non-neighboring spatial merge candidates are allowed to be used as base merge candidates for the merge mode with motion vector difference (MMVD) codec.

[0467] In some examples, non-neighboring spatial merge candidates are not allowed to be used as base merge candidates for the merge mode with motion vector difference (MMVD) codec.

[0468] In some examples, it is allowed to use non-neighboring spatial merge candidates to generate inter-intra predictions.

[0469] In some examples, the use of non-neighboring spatial merge candidates to generate inter-intra predictions is not allowed.

[0470] In some examples, it is allowed to use non-adjacent spatial merge candidates to generate geometric (GEO) partitioning and / or triangle partitioning merge candidates.

[0471] In some examples, it is not allowed to use non-adjacent spatial merge candidates to generate geometric (GEO) partitioning and / or triangle partitioning merge candidates.

[0472] In some examples, it is allowed to use non-neighboring spatial merge candidates to generate affine merge candidates.

[0473] In some examples, it is allowed to use non-neighboring spatial merge candidates to generate Advanced Motion Vector Prediction (AMVP) candidates.

[0474] Figure 15 A flowchart of an example method for video processing is shown. The method includes: for conversion between a current block of video and a bitstream representation of the video, constructing (1502) a merge candidate list for the current block, wherein the construction of the merge candidate list examines a plurality of different categories of candidates in a defined order; and performing (1504) the conversion based on the merge candidate list.

[0475] In some examples, the defined order includes: a first set of spatial merge candidates, spatial temporal motion vector predictor (STMVP) candidates, a second set of spatial merge candidates, temporal motion vector predictor (TMVP) candidates, history-based motion vector predictor (HMVP) candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0476] In some examples, the first group of spatial merge candidates includes candidate B derived from the block above the current block, candidate A derived from the block to the left of the current block, candidate C derived from the upper right block of the current block, and a fourth candidate D derived from the lower left block of the current block, and the second group of spatial merge candidates includes candidate E derived from the upper left block of the current block.

[0477] In some examples, the defined order includes: spatial merge candidates derived from neighboring blocks of the current block, TMVP candidates, a first set of spatial merge candidates derived from non-neighboring blocks of the current block, HMVP candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0478] In some examples, the spatial merge candidates derived from the neighboring blocks of the current block include candidate B derived from the upper block of the current block, candidate A derived from the left block of the current block, candidate C derived from the upper right block of the current block, and a fourth candidate D derived from the lower left block of the current block, as well as candidate E derived from the upper left block of the current block, and the first group of spatial merge candidates derived from the non-neighboring blocks of the current block include candidate B1 derived from the non-neighboring upper block of the current block, candidate A1 derived from the non-neighboring left block of the current block, candidate C1 derived from the non-neighboring upper right block of the current block, candidate D1 derived from the non-neighboring lower left block n of the current block, and candidate E1 derived from the non-neighboring upper left block on the left of the current block, which are derived in the first search round.

[0479] In some examples, the defined order includes: spatial merge candidates derived from neighboring blocks of the current block, TMVP candidates, spatial merge candidates derived from non-neighboring blocks of the current block, HMVP candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0480] In some examples, the spatial merge candidates derived from the neighboring blocks of the current block include candidate B derived from the upper block of the current block, candidate A derived from the left block of the current block, candidate C derived from the upper right block of the current block, and a fourth candidate D derived from the lower left block of the current block, as well as candidate E derived from the upper left block of the current block, and the spatial merge candidates derived from the non-neighboring blocks of the current block include candidate B1 derived from the non-neighboring upper block of the current block, candidate A1 derived from the non-neighboring left block of the current block, and candidate B2 derived from the non-neighboring right block of the current block. Candidate C1 derived from the upper block, candidate D1 derived from the non-adjacent lower left block n of the current block, and candidate E1 derived from the non-adjacent upper left block on the left of the current block, which are derived in the first search round; and candidate B2 derived from the non-adjacent upper block of the current block, candidate A2 derived from the non-adjacent left block of the current block, candidate C2 derived from the non-adjacent upper right block of the current block, candidate D2 derived from the non-adjacent lower left block n of the current block, and candidate E2 derived from the non-adjacent upper left block on the left of the current block, which are derived in the second search round.

[0481] In some examples, the defined order includes: a first set of spatial merge candidates, STMVP candidates, a second set of spatial merge candidates, temporal motion vector predictor (TMVP) candidates, a first set of spatial merge candidates derived from non-neighboring blocks of the current block, HMVP candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0482] In some examples, the first group of spatial merge candidates includes candidate B derived from the block above the current block, candidate A derived from the block to the left of the current block, candidate C derived from the upper right block of the current block, and a fourth candidate D derived from the lower left block of the current block, and the first group of spatial merge candidates derived from the non-adjacent blocks of the current block includes candidate B1 derived from the non-adjacent upper block of the current block, candidate A1 derived from the non-adjacent left block of the current block, candidate C1 derived from the non-adjacent upper right block of the current block, candidate D1 derived from the non-adjacent lower left block n of the current block, and candidate E1 derived from the non-adjacent upper left block on the left of the current block, which are derived in the first search round.

[0483] In some examples, if the corresponding candidate is unavailable, invalid, or the same as or similar to an existing candidate in the merge candidate list, the corresponding candidate is not included in the merge candidate list.

[0484] In some examples, converting includes encoding video units of the video into a bitstream.

[0485] In some examples, converting includes decoding video units of the video from the bitstream.

[0486] Figure 16A flowchart of an example method for storing a bitstream of a video is shown. The method includes: for conversion between a current block of video and a bitstream representation of the video, constructing (1602) a merge candidate list for the current block, wherein non-adjacent spatial merge candidates associated with the current block are inserted into the merge candidate list; generating (1604) a bitstream from a video unit based on the merge candidate list; and storing (1606) the bitstream in a non-transitory computer-readable recording medium.

[0487] Figure 17 A flowchart of an example method for video processing is shown. The method includes: for conversion between a current block of video and a bitstream representation of the video, constructing (1702) a merge candidate list for the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and / or N temporal neighboring motion candidates, where M and N are positive integers; and performing (1704) the conversion based on the merge candidate list.

[0488] In some examples, M>2.

[0489] In some examples, spatial neighboring motion candidates are derived from other neighboring blocks that are different from or the same as those used during the construction process of the merge candidate list.

[0490] In some examples, the spatially adjacent motion candidate is selected from the spatial merge candidates included in the merge candidate list.

[0491] In some examples, before adding the STMVP candidate, a spatially adjacent motion candidate is selected from the first M or last M spatial merge candidates included in the merge candidate list.

[0492] In some examples, before adding the STMVP candidate, a spatially adjacent motion candidate is selected from the first M or last M merge candidates included in the merge candidate list.

[0493] In some examples, the temporal neighboring motion candidate is selected from the temporal merge candidates included in the merge candidate list.

[0494] In some examples, if the time-domain merge candidate is not available, the STMVP candidate is considered unavailable.

[0495] In some examples, whether a spatial neighboring motion candidate and / or a temporal neighboring motion candidate is considered valid is based on reference picture information associated with the current block.

[0496] In some examples, a picture is considered valid only if its reference index in at least one reference picture list is equal to or not greater than K, where K is an integer.

[0497] In some examples, a picture is considered valid only if its reference index in both reference picture lists is equal to or no greater than K, where K is an integer.

[0498] In some examples, K=0.

[0499] In some examples, when it is deemed invalid, it is not used to derive STMVP candidates.

[0500] In some examples, the STMVP candidate is valid if at least one of the first M spatial merge candidates and a co-located merge candidate are valid.

[0501] In some examples, M=3 and N=1, and the STMVP candidate is derived as the average candidate of four merge candidates.

[0502] In some examples, if the reference indices of the four merge candidates are all valid and equal to 0 in the prediction direction X, where X is 0 or 1, the motion vector of the STMVP candidate in the prediction direction X, denoted as mvLX, is derived as follows:

[0503] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e,

[0504] Where a, b, c, d, and e are integers.

[0505] In some examples, a, b, c, d, and e are set equal to 1, 1, 1, 1, and 2.

[0506] In some examples, if the reference indexes of three of the four merge candidates are valid and equal to 0, X=0 or 1 in the prediction direction X, the motion vector of the STMVP candidate in the prediction direction X, denoted as mvLX, is derived as follows:

[0507] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_Col*c)>>d; or

[0508] mvLX=(mvLX_F*a+mvLX_T*b+mvLX_Col*c)>>d; or

[0509] mvLX=(mvLX_S*a+mvLX_T*b+mvLX_Col*c)>>d,

[0510] Where a, b, c, and d are integers.

[0511] In some examples, a, b, c, and d are set equal to 3, 3, 2, and 3, or a, b, c, and d are set equal to 2, 2, 4, and 3, or a, b, c, and d are set equal to 1, 1, 6, and 3.

[0512] In some examples, if the reference indexes of two of the four merge candidates are valid and equal to 0, X=0 or 1 in the prediction direction X, the motion vector of the STMVP candidate in the prediction direction X, denoted as mvLX, is derived as follows:

[0513] mvLX = (mvLX_F*a+mvLX_Col*b)>>c; or

[0514] mvLX=(mvLX_S*a+mvLX_Col*b)>>c; or

[0515] mvLX=(mvLX_T*a+mvLX_Col*b)>>c,

[0516] Where a, b, and c are integers.

[0517] In some examples, a, b, and c are set equal to 1, 1, and 1.

[0518] In some examples, the STMVP candidate is pruned with all previous merge candidates in the merge candidate list.

[0519] In some examples, STMVP candidates are not pruned with other merge candidates.

[0520] In some examples, STMVP candidates are pruned only with merge candidates above and to the left.

[0521] In some examples, the STMVP candidate references one or two specific reference pictures.

[0522] In some examples, the specific reference picture is the reference picture with a reference index equal to 0 in the reference list.

[0523] In some examples, the specific reference picture is a reference picture of the M spatial neighboring motion candidates and / or the N temporal neighboring motion candidates with the smallest reference index in the reference list.

[0524] In some examples, converting includes encoding video units of the video into a bitstream.

[0525] In some examples, converting includes decoding video units of the video from the bitstream.

[0526] Figure 18A flowchart of an example method for storing a bitstream of a video is shown. The method includes: for conversion between a current block of the video and a bitstream representation of the video, constructing (1802) a merge candidate list for the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and / or N temporal neighboring motion candidates, where M and N are positive integers; generating (1804) a bitstream from a video unit based on the merge candidate list; and storing (1806) the bitstream in a non-transitory computer-readable recording medium.

[0527] In this document, the term "video processing" may refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, or vice versa. The bitstream representation of the current video block may, for example, correspond to bits located at the same position within the bitstream or distributed across different positions, as defined by the syntax. For example, a macroblock may be encoded based on the error residual values ​​of a transform and codec, and may also use bits in the header and other fields in the bitstream.

[0528] The disclosed solutions, examples, embodiments, modules, and functional operations, as well as other solutions, examples, embodiments, modules, and functional operations described in this document, can be implemented in digital electronic circuitry or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of matter that effects a machine-readable propagated signal, or a combination of one or more thereof. The term "data processing apparatus" encompasses all apparatus, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program in question, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more thereof. A propagated signal is an artificially generated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to a suitable receiver device.

[0529] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordinated files (e.g., files storing one or more modules, subroutines, or portions of code). A computer program can be deployed to execute on one computer or on multiple computers located in one location or distributed across multiple locations and interconnected by a communications network.

[0530] The processes and logic flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can also be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).

[0531] By way of example, processors suitable for executing a computer program include both general-purpose and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Typically, a processor will receive instructions and data from a read-only memory or a random-access memory, or both. The essential elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer may also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic, magneto-optical, or optical disks) for storing data, receiving data from them or transferring data to them, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of nonvolatile memory, media, and storage devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.

[0532] Although this patent document contains many details, these should not be construed as limitations on the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be unique to particular embodiments of particular technologies. Certain features described in this patent document in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple implementations or in any suitable subcombination. Furthermore, although features may be described above as working in certain combinations and even initially claimed as such, in some cases one or more features in the claimed combination may be cut out of the combination, and the claimed combination may be directed to a subcombination or variation of the subcombination.

[0533] Similarly, while operations are depicted in a particular order in the drawings, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve desired results. Furthermore, the separation of various system components in the embodiments described in this patent document should not be understood as requiring such separation in all embodiments.

[0534] Only a few implementations and examples are described, and other implementations, enhancements, and variations can be made based on what is described and illustrated in this patent document.

Claims

1. A video processing method, comprising: For conversion between a current block of a video and a bitstream of the video, construct a Merge candidate list of the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the Merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and N temporal neighboring motion candidates, where M and N are positive integers; as well as Performing the conversion based on the Merge candidate list; Wherein M=3 and N=1, and the STMVP candidate is derived as an average candidate of four neighboring motion candidates, wherein the spatial domain neighboring motion candidate is selected from the spatial domain Merge candidates included in the Merge candidate list, and the temporal domain neighboring motion candidate is selected from the temporal domain Merge candidates included in the Merge candidate list; If the reference indexes of the four neighboring motion candidates are all valid and equal to 0 in the prediction direction X, where X is 0 or 1, the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX: mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e, Where a, b, c, d and e are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate.

2. The method of claim 1, wherein the spatial neighboring motion candidates are derived from other neighboring blocks that are different from or the same as those used during the construction process of the Merge candidate list.

3. The method according to claim 1, wherein before adding the STMVP candidate, the spatial neighboring motion candidate is selected from the first M or last M spatial Merge candidates included in the Merge candidate list.

4. The method according to claim 1, wherein before adding the STMVP candidate, the spatially neighboring motion candidate is selected from the first M or last M Merge candidates included in the Merge candidate list. The method according to claim 1 , wherein if the time-domain Merge candidate is unavailable, the STMVP candidate is considered unavailable. The method according to claim 1 , wherein whether a spatial neighboring motion candidate and / or a temporal neighboring motion candidate is considered valid is based on reference picture information associated with the current block.

7. The method of claim 6, wherein a picture is considered valid only if its reference index in at least one reference picture list is equal to or not greater than K, where K is an integer.

8. The method of claim 6, wherein a picture is considered valid only if its reference index in both reference picture lists is equal to or not greater than K, where K is an integer.

9. The method according to claim 7 or 8, wherein K=0.

10. The method of claim 6, wherein when it is deemed invalid, it is not used to derive the STMVP candidate.

11. The method according to claim 6, wherein the STMVP candidate is valid if at least one candidate among the first M spatial domain Merge candidates and one temporal domain Merge candidate are valid.

12. The method of claim 11, wherein a, b, c, d, and e are set equal to 1, 1, 1, 1, and 2.

13. The method of claim 10 , wherein if the reference indexes of three of the four neighboring motion candidates are valid and equal to 0, X=0 or 1 in the prediction direction X, the motion vector of the STMVP candidate in the prediction direction X, denoted as mvLX, is derived as follows: mvLX=(mvLX_F*a+mvLX_S*b+mvLX_Col*c)>>d; or mvLX=(mvLX_F*a+mvLX_T*b+mvLX_Col*c)>>d; or mvLX=(mvLX_S*a+mvLX_T*b+mvLX_Col*c)>>d, Where a, b, c and d are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate.

14. The method of claim 13, wherein a, b, c and d are set equal to 3, 3, 2 and 3, or a, b, c and d are set equal to 2, 2, 4 and 3, or a, b, c and d are set equal to 1, 1, 6 and 3.

15. The method of claim 14 , wherein if the reference indexes of two of the four neighboring motion candidates are valid and equal to 0, X=0 or 1 in the prediction direction X, the motion vector of the STMVP candidate in the prediction direction X, denoted as mvLX, is derived as follows: mvLX = (mvLX_F*a+mvLX_Col*b)>>c; or mvLX=(mvLX_S*a+mvLX_Col*b)>>c; or mvLX=(mvLX_T*a+mvLX_Col*b)>>c, Where a, b and c are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate. The method of claim 15 , wherein a, b, and c are set equal to 1, 1, and 1.

17. The method according to any one of claims 1-8 and 10-16, wherein the STMVP candidate is deduplicated using all previous Merge candidates in the Merge candidate list.

18. The method according to any one of claims 1-8 and 10-16, wherein the STMVP candidate is not deduplicated with other Merge candidates.

19. The method according to any one of claims 1-8 and 10-16, wherein the STMVP candidates are deduplicated using only the upper spatial domain Merge candidates and the left spatial domain Merge candidates.

20. The method according to any one of claims 1-8 and 10-16, wherein the STMVP candidate refers to one or two specific reference pictures. The method of claim 20 , wherein the specific reference picture is a reference picture with a reference index equal to 0 in a reference list. 22 . The method according to claim 20 , wherein the specific reference picture is a reference picture of M spatial neighboring motion candidates and N temporal neighboring motion candidates having a smallest reference index in a reference list.

23. The method of any one of claims 1-8, 10-16, 21 and 22, wherein the converting comprises encoding video units of the video into the bitstream.

24. The method of any one of claims 1-8, 10-16, 21, and 22, wherein the converting comprises decoding video units of the video from the bitstream.

25. An apparatus for processing video data, comprising a processor and non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: For conversion between a current block of a video and a bitstream of the video, construct a Merge candidate list of the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the Merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and N temporal neighboring motion candidates, where M and N are positive integers; as well as Performing the conversion based on the Merge candidate list; Wherein M=3 and N=1, and the STMVP candidate is derived as an average candidate of four neighboring motion candidates, wherein the spatial domain neighboring motion candidate is selected from the spatial domain Merge candidates included in the Merge candidate list, and the temporal domain neighboring motion candidate is selected from the temporal domain Merge candidates included in the Merge candidate list; If the reference indexes of the four neighboring motion candidates are all valid and equal to 0 in the prediction direction X, where X is 0 or 1, the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX: mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e, Where a, b, c, d and e are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate.

26. A non-transitory computer-readable medium storing instructions that cause a processor to: For conversion between a current block of a video and a bitstream of the video, construct a Merge candidate list of the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the Merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and N temporal neighboring motion candidates, where M and N are positive integers; as well as Performing the conversion based on the Merge candidate list; Wherein M=3 and N=1, and the STMVP candidate is derived as an average candidate of four neighboring motion candidates, wherein the spatial domain neighboring motion candidate is selected from the spatial domain Merge candidates included in the Merge candidate list, and the temporal domain neighboring motion candidate is selected from the temporal domain Merge candidates included in the Merge candidate list; If the reference indexes of the four neighboring motion candidates are all valid and equal to 0 in the prediction direction X, where X is 0 or 1, the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX: mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e, Where a, b, c, d and e are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate.

27. A non-transitory computer-readable medium storing a bitstream of a video generated by a method performed by a video processing device, wherein the method comprises: For conversion between a current block of a video and a bitstream of the video, construct a Merge candidate list of the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the Merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and N temporal neighboring motion candidates, where M and N are positive integers; as well as generating the bitstream from the video unit of the video based on the Merge candidate list; Wherein M=3 and N=1, and the STMVP candidate is derived as an average candidate of four neighboring motion candidates, wherein the spatial domain neighboring motion candidate is selected from the spatial domain Merge candidates included in the Merge candidate list, and the temporal domain neighboring motion candidate is selected from the temporal domain Merge candidates included in the Merge candidate list; If the reference indexes of the four neighboring motion candidates are all valid and equal to 0 in the prediction direction X, where X is 0 or 1, the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX: mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e, Where a, b, c, d and e are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate.

28. A method for storing a bitstream of a video, comprising: For conversion between a current block of a video and a bitstream of the video, construct a Merge candidate list of the current block, wherein a spatial-temporal motion vector prediction (STMVP) candidate associated with the current block is added to the Merge candidate list, and the STMVP candidate is derived as an average candidate of M spatial neighboring motion candidates and N temporal neighboring motion candidates, where M and N are positive integers; generating the bitstream from the video unit of the video based on the Merge candidate list; as well as storing the bitstream in a non-transitory computer-readable recording medium; Wherein M=3 and N=1, and the STMVP candidate is derived as an average candidate of four neighboring motion candidates, wherein the spatial domain neighboring motion candidate is selected from the spatial domain Merge candidates included in the Merge candidate list, and the temporal domain neighboring motion candidate is selected from the temporal domain Merge candidates included in the Merge candidate list; If the reference indexes of the four neighboring motion candidates are all valid and equal to 0 in the prediction direction X, where X is 0 or 1, the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX: mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e, Where a, b, c, d and e are integers, mvLX_F is the first spatial neighboring motion candidate, mvLX_S is the second spatial neighboring motion candidate, mvLX_T is the third spatial neighboring motion candidate, and mvLX_Col is the temporal neighboring motion candidate.

Citation Information

Patent Citations

  • Merge candidates for motion vector prediction for video coding

    CN109076236A

  • Asymmetric weighted bidirectional prediction Merge

    CN110572645A