Improvements to merge mode

By adjusting the motion vector difference derivation process based on block dimension and prediction direction in MMVD mode, and inserting non-nearby spatial merge candidates and STMVP candidates, the video encoding and decoding in merge mode is optimized, solving the problem of low efficiency in the existing technology and improving the encoding and decoding efficiency in small block cases.

CN115104309BActive Publication Date: 2026-04-17DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2020-12-23
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing video codec standards suffer from inefficiency in the design of merge mode, especially when bidirectional prediction is prohibited in small block cases, and the use of non-nearest merge candidates and STMVP is insufficient, resulting in low codec efficiency.

Method used

By adjusting the derivation process of motion vector difference based on block dimension and prediction direction in MMVD mode, inserting non-nearby spatial merge candidates and spatial-temporal motion vector prediction candidates, optimizing the construction process of the merge candidate list, and directly using internal MVD for prediction under specific conditions.

Benefits of technology

It improves the efficiency of video encoding and decoding, especially in small-block scenarios, reducing unnecessary scaling operations and enhancing both encoding and decoding efficiency and quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115104309B_ABST
    Figure CN115104309B_ABST
Patent Text Reader

Abstract

An improvement to the merge mode is described. An example video processing method includes: constructing a merge candidate list for the current block of the video and the bitstream of the video for a transformation, wherein non-nearest spatial merge candidates associated with the current block are inserted into the merge candidate list; and performing the transformation based on the merge candidate list.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] In accordance with applicable patent law and / or the rules of the Paris Convention, this application aims to promptly claim priority and benefit to International Patent Application No. PCT / CN2019 / 127388, filed on December 23, 2019. The entire disclosure of International Patent Application No. PCT / CN2019 / 127388 is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] The patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video consumes the largest share of bandwidth on the internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to perform cross-component adaptive loop filtering during video encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed. The method includes: a conversion between video units of a video and a codec representation of the video; determining computational operations for motion vector difference (MVD) used with a merge mode (MMVD) codec tool with motion vector difference based on the characteristics of the video units; and performing the conversion based on the determination.

[0007] In another example, a video processing method is disclosed. This method includes performing a conversion between video units of a video and a codec representation of the video, wherein the conversion is performed during operation using a motion vector scaling process that depends on the video's resolution.

[0008] In another example, a video processing method is disclosed. The method includes: generating a merge candidate list for a transformation between video units and the codec representation of the video, wherein non-nearest spatial merge candidates of the video units are inserted into the merge list; and performing a transformation using the merge candidate list.

[0009] In another example, a video processing method is disclosed. The method includes: generating a candidate list for a transformation between video units and the codec representation of the video, the candidates in the candidate list being generated by an average of M spatially adjacent candidates and N temporally adjacent candidates, where M and N are positive integers; and performing the transformation using the merged candidate list.

[0010] In another example, a video processing method is disclosed. The method includes: generating a merge list for the transformation between video units and the codec representation of the video, wherein a construction process for generating the merge list checks multiple candidates in a defined order, and performing the transformation using the merge candidate list.

[0011] In another example, a video processing method is disclosed. This method includes performing a conversion using two long-term reference images and a motion vector scaling process during the conversion between video units of a video and the codec representation of the video.

[0012] In another example, a video processing method is disclosed. The method includes: a conversion between video units and a bitstream of a video; deriving motion vector difference (MVD) for use in a merge mode (MMVD) codec tool with motion vector difference based on the characteristics of the video units; and performing a conversion based on the derived MVD.

[0013] In another example, a video processing method is disclosed. This method includes: for the conversion between video units and the bitstream of the video, deriving a motion vector difference (MVD) using a motion vector (MV) scaling process, wherein the MV scaling process depends on the resolution of the video; and performing a conversion based on the derived MVD.

[0014] In another example, a video processing method is disclosed. This method includes: for the conversion between video units and the bitstream of the video, deriving a motion vector difference (MVD) using a motion vector (MV) scaling process, wherein the MV scaling process uses two long-term reference images; and performing a conversion based on the derived MVD.

[0015] In another example, a method for storing a bitstream of video is disclosed. The method includes: a conversion between video units and a bitstream of video; deriving motion vector difference (MVD) for use in a merge mode (MMVD) codec tool with motion vector difference based on the characteristics of the video units; generating a bitstream from the video units based on the derived MVD; and storing the bitstream in a non-transitory computer-readable recording medium.

[0016] In another example, a video processing method is disclosed. The method includes: constructing a merge candidate list for the current block of a video and a bitstream representation of the video, wherein non-nearest spatial merge candidates associated with the current block are inserted into the merge candidate list; and performing a transformation based on the merge candidate list.

[0017] In another example, a video processing method is disclosed. The method includes: constructing a merge candidate list for the current block of video for a transformation between the current block and the bitstream representation of the video, wherein the construction process of the merge candidate list examines multiple different kinds of candidates in a defined order; and performing a transformation based on the merge candidate list.

[0018] In another example, a method for storing a bitstream of video is disclosed. The method includes: constructing a merge candidate list for the current block of video based on a transformation between the current block and a bitstream representation of the video, wherein non-nearby spatial merge candidates associated with the current block are inserted into the merge candidate list; generating a bitstream from video units based on the merge candidate list; and storing the bitstream in a non-transitory computer-readable recording medium.

[0019] In another example, a video processing method is disclosed. The method includes: constructing a merge candidate list for the current block of a video for a transformation between the current block and the bitstream representation of the video, wherein spatial-temporal motion vector prediction (STMVP) candidates associated with the current block are added to the merge candidate list, and deriving the STMVP candidates as an average candidate of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates, where M and N are positive integers; and performing a transformation based on the merge candidate list.

[0020] In another example, a method for storing a bitstream of video is disclosed. The method includes: constructing a merge candidate list for the current block of video based on a transformation between the current block and the bitstream representation of the video, wherein spatial-temporal motion vector prediction (STMVP) candidates associated with the current block are added to the merge candidate list, and deriving the STMVP candidates as an average of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates, where M and N are positive integers; generating a bitstream from video units based on the merge candidate list; and storing the bitstream in a non-transitory computer-readable recording medium.

[0021] In yet another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0022] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0023] In yet another example, a computer-readable medium on which code is stored is disclosed. This code embodies one of the methods described herein in the form of processor-executable code.

[0024] In yet another example, a computer-readable medium stores a bitstream of video generated by the method described above, performed by a video processing apparatus.

[0025] These and other features are described throughout this document. Attached Figure Description

[0026] Figure 1 An example of an offset added to the horizontal or vertical component of the initial motion vector (MV) is shown.

[0027] Figure 2 The HEVC spatial neighboring blocks of the current block are shown.

[0028] Figure 3 This illustrates the relationship between the virtual block and the current block.

[0029] Figure 4 This is a block diagram of an example video processing system in which the disclosed technology can be implemented.

[0030] Figure 5 This is a block diagram of an example hardware platform used for video processing.

[0031] Figure 6 This is a flowchart of an example method for video processing.

[0032] Figure 7 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0033] Figure 8 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0034] Figure 9 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0035] Figure 10 This is a flowchart of an example method for video processing.

[0036] Figure 11 This is a flowchart of an example method for video processing.

[0037] Figure 12 This is a flowchart of an example method for video processing.

[0038] Figure 13 This is a flowchart of an example method for storing video bitstreams.

[0039] Figure 14 This is a flowchart of an example method for video processing.

[0040] Figure 15 This is a flowchart of an example method for video processing.

[0041] Figure 16 This is a flowchart of an example method for storing video bitstreams.

[0042] Figure 17 This is a flowchart of an example method for video processing.

[0043] Figure 18 This is a flowchart of an example method for storing video bitstreams. Detailed Implementation

[0044] The use of section headings in this document is for ease of understanding and is not intended to limit the applicability of the techniques and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not to limit the scope of the disclosed techniques. Therefore, the techniques described herein are also applicable to other video codec protocols and designs. 1. Summary of the Invention

[0046] This patent document relates to video codec technology. Specifically, it relates to the merge mode in video codecs. It can be applied to existing video codec standards, such as HEVC, or to the finalized standard (Multi-Functional Video Codec). It can also be applied to future video codec standards or video codecs. 2. Background Technology

[0048] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T produced H.261 and H.263, while ISO / IEC produced MPEG-1 and MPEG-4 Visual. The two organizations jointly developed the H.262 / MPEG-2 Video and H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Since H.262, video codec standards have been based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the Joint Exploration Model (JEM). In April 2018, the Joint Video Experts Group (JVET) between VCEG (Q6 / 16) and ISO / IEC JTC1SC29 / WG11 (MPEG) was created to develop the VVC standard, with the goal of reducing the bit rate by 50% compared to HEVC.

[0049] 2.1. Merge pattern with MVD (MMVD)

[0050] In addition to the merge mode, which directly uses implicitly derived motion information to generate prediction samples for the current CU, VVC also introduces a merge mode with motion vector difference (MMVD), also known as the ultimate motion vector representation. Immediately after sending the skip flag and merge flag, signaling is used to notify the MMVD flag to specify whether the MMVD mode is used for the CU.

[0051] In MMVD, a merge candidate (called the base merge candidate) is selected, and the MVD information notified via signaling is further refined. Relevant syntax elements include an index specifying the MVD distance (represented by `mmvd_distance_idx`) and an index indicating the direction of motion (represented by `mmvd_direction_idx`). In MMVD mode, one of the top two candidates in the merge list is selected as the MV base (or base merge candidate). Signaling notifies the merge candidate flags to specify which candidate to use.

[0052] The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. For example... Figure 1 As shown, the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 3.

[0053] Table 3: Relationship between Distance Index and Predefined Offset

[0054]

[0055]

[0056] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 4. It is important to note that the meaning of the MVD symbol can vary depending on the information of the starting MV. When the starting MV is a unidirectional prediction MV or a bidirectional prediction MV where both lists point to the same side of the current image (i.e., both reference POCs are greater than or less than the current image's POC), the symbols in Table 4 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV where two MVs point to different sides of the current image (i.e., one reference POC is greater than the current image's POC, and the other reference POC is less than the current image's POC), the symbols in Table 4 specify the sign of the MV offset added to the list 0 MV component of the starting MV, while the signs for list 1 MV have the opposite value.

[0057] Table 4: Sign of MV Offset Specified by Direction Index

[0058] Directional IDX 00 01 10 11 x-axis + – N / A N / A y-axis N / A N / A + –

[0059] 2.2.1 Exporting MVD for each reference image list

[0060] First, an internal MVD (represented by MmvdOffset) is derived based on the index of the decoded MVD distance (represented by mmvd_distance_idx) and the index of the motion direction (represented by mmvd_direction_idx).

[0061] Next, if the internal MVD is determined, the final MVD of the base merge candidates to be added to each reference image list is further derived based on the POC distance of the reference image relative to the current image and the type of the reference image (long-term or short-term). More specifically, the following steps are performed in sequence:

[0062] —If the base merge candidate is bidirectional prediction, then calculate the POC distance between the current image and the reference image in list 0, and the POC distance between the current image and the reference image in list 1, which are represented by POCDiffL0 and POCDidffL1, respectively.

[0063] —If POCDiffL0 equals POCDidffL1, then the final MVD of both reference image lists is set to the inner MVD.

[0064] —Otherwise, if Abs(POCDiffL0) is greater than or equal to Abs(POCDiffL1), the final MVD of reference picture list 0 is set to the inner MVD, while the final MVD of reference picture list 1 uses the inner MVD of both reference pictures. The reference picture type (neither of which is a long-term reference picture) is set to the scaled MVD or is set to the inner MVD based on the POC distance or (zero MV minus the inner MVD).

[0065] —Otherwise, if Abs(POCDiffL0) is less than Abs(POCDiffL1), the final MVD of reference picture list 1 is set to the inner MVD, while the final MVD of reference picture list 0 uses the inner MVD of both reference pictures. The reference picture type (neither of which is a long-term reference picture) is set to the scaled MVD or is set to the inner MVD based on the POC distance or (zero MV minus the inner MVD).

[0066] —If the base merge candidate is a one-way prediction from the reference image list X, then the final MVD of the reference image list X is set to the internal MVD, while the final MVD of the reference image list Y (Y = 1-X) is set to 0.

[0067] 2.2.2 MMVD Specification in VVC

[0068] The MMVD specification (in JVET-P2001-vE) is as follows:

[0069] 7.3.9.7 Merge Data Syntax

[0070]

[0071]

[0072] `mmvd_merge_flag[x0][y0]` equal to 1 indicates that the merge mode with motion vector difference is used to generate the inter-frame prediction parameters for the current codec unit. `mmvd_merge_flag[x0][y0]` equal to 0 indicates that the merge mode with motion vector difference is not used to generate the inter-frame prediction parameters. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the considered codec block relative to the top-left luminance sample of the image.

[0073] When mmvd_merge_flag[x0][y0] does not exist, it is inferred to be equal to 0.

[0074] mmvd_cand_flag[x0][y0] specifies whether to merge the first (0) or second (1) candidate in the candidate list with the motion vector difference derived from mmvd_distance_idx[x0][y0] and mmvd_direction_idx[x0][y0]. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0075] When mmvd_cand_flag[x0][y0] does not exist, it is inferred to be equal to 0.

[0076] mmvd_distance_idx[x0][y0] specifies the index used to derive MmvdDistance[x0][y0], as specified in Table 17. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0077] Table 17 – Specification of MmvdDistance[x0][y0] based on mmvd_distance_idx[x0][y0].

[0078]

[0079]

[0080] mmvd_direction_idx[x0][y0] specifies the index used to export MmvdSign[x0][y0], as specified in Table 18. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0081] Table 18 - Specification of MmvdSign[x0][y0] based on mmvd_direction_idx[x0][y0]

[0082]

[0083] The merged components, plus the MVD offset MmvdOffset[x0][y0], are exported as follows:

[0084] MmvdOffset[x0][y0][0]=(MmvdDistance[x0][y0]<<2)*MmvdSig n[x0][y0][0](181)

[0085] MmvdOffset[x0][y0][1]=(MmvdDistance[x0][y0]<<2)*MmvdSig n[x0][y0][1](182)

[0086] 8.5.2.7 Derivation process of merging motion vector differences

[0087] The input to this process is:

[0088] —The brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.

[0089] —Refer to indices refIdxL0 and refIdxL1,

[0090] —The prediction list uses the flags predFlagL0 and predFlagL1.

[0091] The output of this process is the brightness merge motion vector difference mMvdL0 and mMvdL1 with a precision of 1 / 16 fractional sample point.

[0092] The variable currPic specifies the current image.

[0093] The brightness merge motion vector differences mMvdL0 and mMvdL1 are derived as follows:

[0094] —If both predFlagL0 and predFlagL1 are equal to 1, then the following applies:

[0095] currPocDiffL0=DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) (564)

[0096] currPocDiffL1=DiffPicOrderCnt(currPic,RefPicList[1][refIdxL1]) (565)

[0097] —If currPocDiffL0 equals currPocDiffL1, then the following applies:

[0098] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (566)

[0099] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (567)

[0100] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (568)

[0101] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (569)

[0102] —Otherwise, if Abs(currPocDiffL0) is greater than or equal to Abs(currPocDiffL1), then the following applies:

[0103] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (570)

[0104] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (571)

[0105] —If RefPicList[0][refIdxL0] is not a long-term reference image and RefPicList[1][refIdxL1] is not a long-term reference image, then the following applies:

[0106] td=Clip3(-128,127,currPocDiffL0) (572)

[0107] tb=Clip3(-128,127,currPocDiffL1) (573)

[0108] tx=(16384+(Abs(td)>>1)) / td (574)

[0109] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (575)

[0110] mMvdL1[0]=Clip3(-2 17 ,2 17 -1,(distScaleFactor*mMvdL0[0]+ (576)128-(distScaleFactor*mMvdL0[0]>=0))>>8)

[0111] mMvdL1[1]=Clip3(-2 17 ,2 17 -1,(distScaleFactor*mMvdL0[1]+ (577)128-(distScaleFactor*mMvdL0[1]>=0))>>8)

[0112] —Otherwise, the following applies:

[0113] mMvdL1[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0114] mMvdL0[0]:-mMvdL0[0] (578)

[0115] mMvdL1[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0116] mMvdL0[1]:-mMvdL0[1] (579)

[0117] —Otherwise (Abs(currPocDiffL0) is less than Abs(currPocDiffL1)), the following applies:

[0118] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (580)

[0119] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (581)

[0120] —If RefPicList[0][refIdxL0] is not a long-term reference image and RefPicList[1][refIdxL1] is not a long-term reference image, then the following applies:

[0121] td=Clip3(-128,127,currPocDiffL1) (582)

[0122] tb=Clip3(-128,127,currPocDiffL0) (583)

[0123] tx=(16384+(Abs(td)>>1)) / td (584)

[0124] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (585)

[0125] mMvdL0[0]=Clip3(-2 17 ,2 17 -1,(distScaleFactor*mMvdL1[0]+ (586)128-(distScaleFactor*mMvdL1[0]>=0))>>8)

[0126] mMvdL0[1]=Clip3(-2 17 ,2 17 -1,,(distScaleFactor*mMvdL1[1]+ (587)128-(distScaleFactor*mMvdL1[1]>=0))>>8))

[0127] —Otherwise, the following applies:

[0128] mMvdL0[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0129] mMvdL1[0]:-mMvdL1[0] (588)

[0130] mMvdL0[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0131] mMvdL1[1]:-mMvdL1[1] (589)

[0132] —Otherwise (predFlagL0 or predFlagL1 equals 1), the following applies for X equal to 0 and 1:

[0133] The following situations:

[0134] mMvdLX[0]=(predFlagLX==1)? MmvdOffset[xCb][yCb][0]:0(590)

[0135] mMvdLX[1]=(predFlagLX==1)? MmvdOffset[xCb][yCb][1]:0(591)

[0136] 2.2. JVET-L0323: Long-range merge candidate

[0137] In HEVC, Figure 2 The five spatially neighboring blocks and one temporal neighbor shown are used to derive merge candidates.

[0138] Figure 2 The HEVC spatial neighboring blocks of the current block are shown.

[0139] This contribution proposes deriving additional merge candidates from non-adjacent positions of the current block using the same pattern as in HEVC. To this end, for each search round i, a virtual block is generated based on the current block, as described below:

[0140] First, the relative position of the virtual block to the current block is calculated using the following formula:

[0141] Offsetx=-i*gridX,Offsety=-i*gridY

[0142] Offsetx and Offsety represent the offset of the top-left corner of the virtual block relative to the top-left corner of the current block, while gridX and gridY are the width and height of the search grid.

[0143] Secondly, the width and height of the virtual block are calculated using the following formula:

[0144] newWidth=i*2*gridX+currWidth newHeight=i*2*gridY+currHeight.

[0145] Where currWidth and currHeight are the width and height of the current block, and newWidth and newHeight are the width and height of the new block.

[0146] gridX and gridY are currently set to currWidth and currHeight, respectively.

[0147] Figure 3 This illustrates the relationship between the virtual block and the current block.

[0148] After the virtual block is generated, block A i B i C i D i and E i These can be considered as HEVC spatially adjacent blocks, and their positions are obtained using the same pattern as in HEVC. Clearly, if the search round i is 0, then the virtual block is the current block. In this case, block A... i B i C i D i and E i It is a spatially adjacent block used in HEVCmerge mode.

[0149] When constructing the merge candidate list, deduplication is performed to ensure that each element in the merge candidate list is unique. As more and more blocks are examined to derive additional merge candidates, the deduplication count increases accordingly. To limit the deduplication count in the worst case, the maximum deduplication count allowed in the merge list construction is constrained to a predefined value, MaxPruningNum.

[0150] In the simulation, the maximum number of search rounds was set to 2, and MaxPruningNum was set to 30.

[0151] Long-distance merge candidates are also called non-nearest merge candidates.

[0152] Figure 3 It is a diagram of the virtual block in the i-th search round.

[0153] 2.3.JVET-M0059: Non-scaling STMVP

[0154] The proposed method uses two spatial merge candidates and one collocated merge candidate to derive an average candidate as an STMVP candidate.

[0155] STMVP is inserted before the merge candidate in the upper left space.

[0156] For airspace candidates, the first and second candidates in the current merge candidate list are used.

[0157] For time-domain candidates, use the same position as the VTM / HEVC co-position.

[0158] If three candidates with a reference value of 0 are available, the following applies.

[0159] mvLX[0]=(mvLX_A[0]*3+mvLX_B[0]*3+mvLX_C[0]*2) / 8

[0160] mvLX[1]=(mvLX_A[1]*3+mvLX_B[1]*3+mvLX_C[1]*2) / 8

[0161] If two motion data points with a reference value of zero are available, the following applies.

[0162] mvLX[0]=(mvLX_A[0]+mvLX_C[0]) / 2

[0163] mvLX[1]=(mvLX_A[1]+mvLX_C[1]) / 2

[0164] or

[0165] mvLX[0]=(mvLX_B[0]+mvLX_C[0]) / 2

[0166] mvLX[1]=(mvLX_B[1]+mvLX_C[1]) / 2

[0167] Note: STMVP mode is turned off if time-domain candidates are not available.

[0168] MMVD is also known as Ultimate Motion Vector Expression (UMVE).

[0169] 3. Technical problems solved by the technical solutions and embodiments in this paper

[0170] The current design of the merge pattern can be further improved.

[0171] 1. In MMVD mode, for small blocks (e.g., 4x8 / 8x4), even if only unidirectional prediction is allowed, two MVDs can still be derived if the base merge candidate is bidirectional prediction. More specifically, if the selected MV base (or base merge candidate) is bidirectional MV, the MVD of the prediction direction from one reference list X (X = 0 or 1) is directly set to equal the MVD of the signaling notification, while the MVD of the other reference list Y (Y = 1 – X) is derived based on the MVD of the prediction direction X and the POC (Picture Order Count) distance, thus requiring scaling in some cases. However, in VTM-7.0, bidirectional prediction is disabled for 4x8 / 8x4 blocks. Therefore, there is no need to derive the L1 MVD.

[0172] 2. Furthermore, non-nearest neighbor merge candidates and / or STMVPs can be used to improve the effectiveness of merge patterns. Additionally, encoding / decoding efficiency can be improved.

[0173] 4. Example embodiments and techniques

[0174] The following items should be considered as examples for explaining general concepts. These items should not be interpreted narrowly. Furthermore, these items can be combined in any way.

[0175] In MMVD, the internal MVD is derived from the syntax elements of signaling notifications in the bitstream (such as MVD distance and direction information). Furthermore, the final MVD is the MVD used to refine the basic merge candidates, i.e., the MVD used to derive the final MV of the block.

[0176] In the following text, currWidth and currHeight are the width and height of the current block (e.g., the brightness block). maxNumMergeCand represents the size of the merge list.

[0177] As shown in Chapter 2.3, after the virtual block is generated, block A i B i C i D i and E i These can be considered HEVC spatially adjacent blocks, and their positions are obtained using the same pattern as HEVC. Clearly, if the search round i is 0, then the virtual block is the current block. In this case, block A... i B i C i D i and E i It is a spatially adjacent block used in HEVCmerge mode.

[0178] For airspace candidates, the first, second, and third candidates in the current merge candidate list inserted before STMVP are denoted as F, S, and T, respectively.

[0179] For time-domain candidates at the same position as VTM / HEVC, the corresponding position used in STMVP is denoted as Col.

[0180] MVD Export

[0181] 1. How to derive the MVD used in the MMVD method can depend on the block dimension and / or the allowed prediction direction (e.g., whether only unidirectional prediction is allowed for video units (e.g., CU / PU)).

[0182] a. In one example, if only unidirectional prediction is allowed for a video unit, then only one MVD is derived from the internal MVD instead of two MVDs, regardless of the prediction direction associated with the underlying merge candidate in the MMVD.

[0183] (i) In one example, if the prediction from only the reference picture list X, denoted by ListX (e.g., X = 0), is the prediction direction, the final MVD of ListX is derived from the internal MVD.

[0184] (i) Alternatively, in addition, the final MVD of ListX is set equal to the internal MVD.

[0185] (ii) Alternatively, in addition, the final MVD of ListX is set equal to the opposite value of the internal MVD.

[0186] (iii) Alternatively, in addition, the final MVD of ListY is set to a default value, e.g., zero MVD.

[0187] (b) In one example, if one or more specific conditions depending on the block dimension are satisfied, only one MVD instead of two MVDs is derived from the internal MVD, regardless of the prediction direction associated with the base merge candidate in the MMVD.

[0188] (i) In one example, the condition is that currWidth + currHeight is less than or equal to N (N is a positive integer). For example, N = 12.

[0189] (ii) In one example, the condition is that currWidth * currHeight is less than or equal to N (N is a positive integer). For example, N = 32.

[0190] (iii) In one example, the condition is that currWidth < N1 or / and currHeight < N2 (N1, N2 are positive integers). For example, N1 = N2 = 8.

[0191] (iv) In one example, the condition is that currWidth < N3 * currHeight and / or currHeight < N4 * currWidth (N3, N4 are positive integers). For example, N3 = N4 = 8.

[0192] 2. It is proposed that in the MMVD, when the base merge candidate is a bi - directional MV, if the block dimension or block shape satisfies one or more conditions, the internal MVD can always be directly used (e.g., without scaling) for the prediction direction X (X = 0, 1).

[0193] (a) In one example, the internal MVD is always directly used for the prediction direction 0.

[0194] (b) In one example, the condition is that currWidth + currHeight is less than or equal to N (N is a positive integer). For example, N = 12.

[0195] c. In one example, the condition is that currWidth * currHeight is less than or equal to N (N is a positive integer). For example, N = 32.

[0196] d. In one example, the condition is that currWidth < N1 or / and currHeight < N2 (N1, N2 are positive integers). For example, N1 = N2 = 8.

[0197] e. In one example, the condition is that currWidth < N3 * currHeight and / or currHeight < N4 * currWidth (N3, N4 are positive integers). For example, N3 = N4 = 8.

[0198] f. If the block dimension or block shape satisfies one or more conditions, the opposite value (-MVD) of the internal MVD can be used to replace the MVD prediction direction X (X = 0, 1).

[0199] 3. The MV scaling process (such as those used in MMVD, TMVP, etc.) can take the picture resolution into consideration.

[0200] 4. For two reference pictures that are both long-term reference pictures, the MV scaling process can still be applied.

[0201] a. In one example, the MV scaling process can be similar to the case where the two reference pictures are short-term reference pictures, that is, depending on the POC distance.

[0202] Non-nearest merge candidates

[0203] 5. Non-adjacent spatial domain merge candidates can be inserted into the merge list.

[0204] a. In one example, the non-adjacent spatial domain merge candidates are inserted into the merge list after the history-based merge candidates.

[0205] b. In one example, the non-adjacent spatial domain merge candidates are inserted into the merge list after the pairwise average merge candidates.

[0206] c. In one example, if the number of available merge candidates in the merge list reaches a predefined value after inserting the temporal domain merge candidates, the non-adjacent spatial domain merge candidates may not be inserted. d. In one example, if the number of available merge candidates in the merge list reaches a predefined value when inserting the non-adjacent spatial domain merge candidates, the insertion process will be terminated.

[0207] e. In one example, the predefined value is equal to maxNumMergeCand–N.

[0208] i. In one example, N is set to be equal to 1, 2, 3 or 4.

[0209] f. In one example, the maximum number of search rounds is set to 1 or 2, meaning that five or ten non-nearest airspace merge candidates can be used to build the merge list.

[0210] g. In one example, for each search round, the insertion order is A. i B i C i D i and E i .

[0211] i. Alternatively, for each search round, the insertion order is B. i A i C i D i and E i .

[0212] ii. Alternatively, for each search round, the insertion order is B. i C i A i D i and E i .

[0213] iii. Alternatively, for each search round, the insertion order is A i D i B i C i and E i .

[0214] h. In one example, all spatial and temporal merge candidates undergo complete deduplication against all previous merge candidates in the merge list. The deduplication process for historical merge candidates and pairwise averaged candidates remains unchanged.

[0215] i. Alternatively, all spatial, temporal, historical, and pairwise average merge candidates are completely deduplicated against all previous merge candidates in the merge list.

[0216] j. Alternatively, for non-adjacent airspace merge candidates, A i Use A i-1 Perform deduplication, B i Use A i Perform deduplication, C i Use B i Perform deduplication, Di Use A i Perform deduplication, E i Use A i Deduplication is performed using Bi. The deduplication processes for the time domain, history-based, and pairwise averaged candidates remain unchanged.

[0217] k. In one example, the maximum number of deduplications allowed in the merge list construction, MaxPruningNum, can depend on the merge list size, maxNumMergeCand.

[0218] i. For example, MaxPruningNum can be set to be equal to maxNumMergeCand – M (where M is an integer). For example, M = 2.

[0219] ii. For example, MaxPruningNum can be set to equal maxNumMergeCand*M (where M is an integer). For example, M = 2.

[0220] iii. Alternatively, MaxPruningNum is independent of the merge list size maxNumMergeCand. For example, MaxPruningNum can be set to equal 30 or 35.

[0221] l. In one example, non-adjacent airspace merge candidate locations are constrained to a predefined region.

[0222] i. In one example, the region could contain the current CTU row and four sample rows above it.

[0223] ii. In one example, the region may contain the current CTU column and the four left sample point columns of the current CTU.

[0224] iii. In one example, the area may contain the current CTU column and the left CTU column of the current CTU.

[0225] m. In one example, there are no constraints on the horizontal direction for merging candidate locations in non-nearest spatial domains.

[0226] In one example, non-nearest neighbor space merge candidates can be used as the base merge candidates for MMVD.

[0227] i. Alternatively, non-adjacent airspace merge candidates are not permitted to be used as base merge candidates for MMVD.

[0228] o. In one example, non-nearest spatial domain merge candidates can be used to generate inter-frame-intra-frame predictions.

[0229] i. Alternatively, non-adjacent spatial domain merge candidates are not allowed to generate inter-frame-intra-frame predictions.

[0230] p. In one example, non-nearest spatial domain merge candidates can be used to generate geometric (GEO) segmentation and / or triangular segmentation merge candidates.

[0231] i. Alternatively, non-adjacent spatial domain merge candidates are not allowed to generate geometric (GEO) segmentation and / or triangular segmentation merge candidates.

[0232] q. In one example, non-nearest spatial domain merge candidates can be used to generate affine merge candidates.

[0233] r. In one example, non-nearby spatial domain merge candidates can be used to generate advanced motion vector prediction (AMVP) candidates.

[0234] Spatial-Time Motion Vector Prediction (STMVP)

[0235] 6. STMVP candidates can be derived as the average candidate of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates.

[0236] a. In one example, M>2.

[0237] b. In one example, spatial neighbor motion candidates can be derived from other neighboring blocks that are different from or the same as those neighboring blocks used in the merge list construction process.

[0238] c. In one example, spatially adjacent motion candidates can be selected from the spatially adjacent motion candidates included in the merge list.

[0239] d. In one example, spatially adjacent motion candidates can be selected from the first M or last M spatially adjacent motion candidates included in the merge list before adding STMVP.

[0240] e. In one example, spatially adjacent motion candidates can be selected from the first M or last M merge candidates included in the merge list before adding STMVP.

[0241] f. In one example, temporally adjacent motion candidates can be selected from temporally merged candidates.

[0242] i. In one example, if the temporal merge candidate is unavailable, then the STMVP candidate is considered unavailable.

[0243] g. In one example, whether spatially adjacent motion candidates and / or temporally adjacent motion candidates are considered valid is based on reference image information.

[0244] i. In one example, only if its reference index in at least one list of reference images is equal to or not greater than K (e.g., K = 0).

[0245] ii. In one example, only if its reference index in both reference image lists is equal to or not greater than K (e.g., K = 0).

[0246] iii. Alternatively, it is not used to derive STMVP candidates when it is considered invalid.

[0247] iv. Alternatively, the STMVP candidate is valid if at least one of the first M spatial merge candidates and one co-located merge candidate are valid.

[0248] h. In one example, M is set to equal 3 and N is set to equal 1.

[0249] i. In one example, if the reference indices of all four merge candidates are valid and equal to 0 (X = 0 or 1) in the prediction direction X, then the motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:

[0250] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e

[0251] (i) In one example, a, b, c, d, and e are set to equal 1, 1, 1, 1, and 2.

[0252] ii. In one example, if the reference indices of three of the four merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1), then the motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:

[0253] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_Col*c)>>d or

[0254] mvLX=(mvLX_F*a+mvLX_T*b+mvLX_Col*c)>>d or

[0255] mvLX=(mvLX_S*a+mvLX_T*b+mvLX_Col*c)>>d

[0256] (i) In one example, a, b, c and d are set to equal 3, 3, 2 and 3.

[0257] (ii) In one example, a, b, c, and d are set to equal 2, 2, 4, and 3.

[0258] (iii) In one example, a, b, c and d are set to equal 1, 1, 6 and 3.

[0259] iii. In one example, if the reference indices of two of the four merge candidates are valid and equal to 0 in the prediction direction X (X = 0 or 1), then the motion vector of the STMVP candidate in the prediction direction X (denoted as mvLX) is derived as follows:

[0260] mvLX=(mvLX_F*a+mvLX_Col*b)>>c or

[0261] mvLX = (mvLX_S*a + mvLX_Col*b) >> c or

[0262] mvLX=(mvLX_T*a+mvLX_Col*b)>>c

[0263] (i) In one example, a, b, and c are set to equal to 1, 1, and 1, respectively.

[0264] i. In one example, STMVP candidates can be deduplicated using all previous merge candidates in the merge list.

[0265] j. In one example, the STMVP candidate can be deduplicated without other merge candidates.

[0266] k. In one example, the STMVP candidate can be deduplicated using only the merge candidates above and to the left.

[0267] l. In one example, STMVP candidates refer to one or two specific reference images.

[0268] i. For example, a specific reference image is the reference image with a reference index of 0 in the reference list.

[0269] ii. For example, a specific reference image is a reference image of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates with the minimum reference index in the reference list.

[0270] Merge list construction process

[0271] 7. The merge list construction process may include checking the following candidates in sequence.

[0272] a. First group of spatial merge candidates (e.g., derived from B, A, C, D), STMVP, second group of spatial merge candidates (e.g., derived from E), TMVP, HMVP, pairwise average merge candidates, zero motion vector merge candidates.

[0273] b. Spatial merge candidates derived from neighboring blocks (e.g., derived from B, A, C, D, E), TMVP, the first set of spatial merge candidates derived from non-neighboring blocks (e.g., derived from B1, A1, C1, D1, E1), HMVP, pairwise average merge candidates, and zero motion vector merge candidates.

[0274] c. Spatial merge candidates derived from neighboring blocks (e.g., derived from B, A, C, D, E), TMVP, spatial merge candidates derived from non-neighboring blocks (e.g., derived from B1, A1, C1, D1, E1, B2, A2, C2, D2, E2), HMVP, pairwise average merge candidates, and zero motion vector merge candidates.

[0275] d. First group of spatial merge candidates (e.g., derived from B, A, C, D), STMVP, second group of spatial merge candidates (e.g., derived from E), TMVP, first group of spatial merge candidates derived from non-neighboring blocks (e.g., derived from B1, A1, C1, D1, E1), HMVP, pairwise average merge candidates, zero motion vector merge candidates.

[0276] e. In the example above, if the corresponding candidate is unavailable, invalid, or the same as or similar to an existing candidate (added before the corresponding candidate), then the corresponding candidate will not be included in the motion candidate list.

[0277] 8. The above method can be applied to other types of motion candidate lists besides merging candidate lists.

[0278] a. Alternatively, the above method can be applied to the block vector candidate list construction process for IBC codec blocks. In this case, checking whether the block is encoded and decoded in IBC mode can be used instead of checking whether the reference image index is equal to K.

[0279] 5. Examples

[0280] The deleted parts are in gray. The newly added section is highlighted. Highlight.

[0281] 5.1. Example #1 of MMVD

[0282] If the MV base (or base merge candidate) selected in MMVD mode is a bidirectional MV, and the sum of the width and height of the block is less than or equal to 12, then the MVD of prediction direction 0 (L0) is directly set to be equal to the MVD of the signaling notification, and the MMVD merge candidate is converted to the L0 unidirectional prediction candidate.

[0283] 8.5.2.7 Derivation process of merging motion vector differences

[0284] The input to this process is:

[0285] —The brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.

[0286] —Refer to indices refIdxL0 and refIdxL1,

[0287] —The prediction list uses the flags predFlagL0 and predFlagL1.

[0288] The output of this process is the brightness merge motion vector difference mMvdL0 and mMvdL1 with a precision of 1 / 16 fractional sample point.

[0289] The variable currPic specifies the current image.

[0290] The brightness merge motion vector differences mMvdL0 and mMvdL1 are derived as follows:

[0291] —If both predFlagL0 and predFlagL1 are equal to 1 and The following situations apply:

[0292] currPocDiffL0=DiffPicOrderCnt(currPic,RefPicList[0][refIdxL0]) (564)

[0293] currPocDiffL1=DiffPicOrderCnt(currPic,RefPicList[1][refIdxL1]) (565)

[0294] —If currPocDiffL0 equals currPocDiffL1, then the following applies:

[0295] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (566)

[0296] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (567)

[0297] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (568)

[0298] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (569)

[0299] —Otherwise, if Abs(currPocDiffL0) is greater than or equal to Abs(currPocDiffL1), then the following applies:

[0300] mMvdL0[0]=MmvdOffset[xCb][yCb][0] (570)

[0301] mMvdL0[1]=MmvdOffset[xCb][yCb][1] (571)

[0302] —If RefPicList[0][refIdxL0] is not a long-term reference image and RefPicList[1][refIdxL1] is not a long-term reference image, then the following applies:

[0303] td=Clip3(-128,127,currPocDiffL0) (572)

[0304] tb=Clip3(-128,127,currPocDiffL1) (573)

[0305] tx=(16384+(Abs(td)>>1)) / td (574)

[0306] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (575)

[0307] mMvdL1[0]=Clip3(-2 17 ,2 17 -1,(distScaleFactor*mMvdL0[0]+ (576)128-(distScaleFactor*mMvdL0[0]>=0))>>8)

[0308] mMvdL1[1]=Clip3(-2 17 ,2 17 -1,(distScaleFactor*mMvdL0[1]+ (577)128-(distScaleFactor*mMvdL0[1]>=0))>>8)

[0309] —Otherwise, the following applies:

[0310] mMvdL1[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0311] mMvdL0[0]:-mMvdL0[0] (578)

[0312] mMvdL1[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0313] mMvdL0[1]:-mMvdL0[1] (579)

[0314] —Otherwise (Abs(currPocDiffL0) is less than Abs(currPocDiffL1)), the following applies:

[0315] mMvdL1[0]=MmvdOffset[xCb][yCb][0] (580)

[0316] mMvdL1[1]=MmvdOffset[xCb][yCb][1] (581)

[0317] —If RefPicList[0][refIdxL0] is not a long-term reference image and RefPicList[1][refIdxL1] is not a long-term reference image, then the following applies:

[0318] td=Clip3(-128,127,currPocDiffL1) (582)

[0319] tb=Clip3(-128,127,currPocDiffL0) (583)

[0320] tx=(16384+(Abs(td)>>1)) / td (584)

[0321] distScaleFactor=Clip3(-4096,4095,(tb*tx+32)>>6) (585)

[0322] mMvdL0[0]=Clip3(-2 17 ,2 17 -1,(distScaleFactor*mMvdL1[0]+ (586)128-(distScaleFactor*mMvdL1[0]>=0))>>8)

[0323] mMvdL0[1]=Clip3(-217 ,2 17 -1,,(distScaleFactor*mMvdL1[1]+ (587)128-(distScaleFactor*mMvdL1[1]>=0))>>8))

[0324] —Otherwise, the following applies:

[0325] mMvdL0[0]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0326] mMvdL1[0]:-mMvdL1[0] (588)

[0327] mMvdL0[1]=Sign(currPocDiffL0)==Sign(currPocDiffL1)?

[0328] mMvdL1[1]:-mMvdL1[1] (589)

[0329] —Otherwise (predFlagL0 or predFlagL1 equals 1) If X is 0 and 1, then the following cases apply:

[0330] mMvdLX[0]=(predFlagLX==1)? MmvdOffset[xCb][yCb][0]:0(590)

[0331] mMvdLX[1]=(predFlagLX==1)? MmvdOffset[xCb][yCb][1]:0(591)

[0332] Figure 4 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein may be implemented. Various specific implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8-bit or 10-bit multi-component pixel values), or it may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet and Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0333] System 1900 may include a codec component 1904 capable of implementing the various codec or encoding methods described in this document. Codec component 1904 reduces the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 may be stored or transmitted via connected communication, as represented by component 1906. The bitstream (or codec) representation of the stored or transmitted video received at input 1902 can be used by component 1908 to generate pixel values ​​or to send displayable video to display interface 1910. The process of generating a user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, while the corresponding decoding tools or operations that reverse the codec results are performed by the decoder.

[0334] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be found in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0335] Figure 5 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors 3602 can be configured to implement one or more methods described herein. The memories 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0336] Figure 7 This is a block diagram illustrating an example video encoding / decoding system 100 that can utilize the technology of the present invention.

[0337] like Figure 7As shown, the video encoding / decoding system 100 may include a source device 110 and a target device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The target device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0338] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0339] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems used to generate video data, or combinations of such sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to target device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by target device 120.

[0340] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0341] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120 or may be external to target device 120 configured to interface with external display devices.

[0342] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Multi-Functional Video Codec (VVC) standard, and other current and / or additional standards.

[0343] Figure 8 This is a block diagram illustrating an example of a video encoder 200, which may be... Figure 7 The video encoder 114 in the system 100 shown.

[0344] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 8 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0345] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0346] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, in which at least one reference picture is the picture containing the current video block.

[0347] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for interpretable purposes... Figure 8 The example is shown separately.

[0348] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0349] The mode selection unit 203 can select one of the encoding / decoding modes (intra-frame or inter-frame) based, for example, on the error result, and provide the obtained intra-frame or inter-frame codec block to the residual generation unit 207 to generate residual block data and provide it to the reconstruction unit 212 to reconstruct the codec block for use as a reference picture. In some examples, the mode selection unit 203 can select a combined intra-frame and inter-frame prediction (CIIP) mode, in which prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0350] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.

[0351] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block, for example, depending on whether the current video block is in an I-slice, a P-slice or a B-slice.

[0352] In some examples, motion estimation unit 204 can perform unidirectional prediction for the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0353] In other examples, motion estimation unit 204 can perform bidirectional prediction for the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images in list 0 and also search for another reference video block for the current video block in the reference images in list 1. Then, motion estimation unit 204 can generate reference indices indicating the reference images containing the reference video blocks in lists 0 and 1, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0354] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder to use in the decoding process.

[0355] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block by referencing the motion information of another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0356] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0357] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0358] As discussed above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and merge pattern signaling.

[0359] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0360] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) one or more predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0361] In other examples, residual data for the current video block may not exist, for example in skip mode, and the residual generation unit 207 may not perform subtraction operations.

[0362] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0363] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0364] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to generate a reconstructed video block associated with the current block and store it in buffer 213.

[0365] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0366] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0367] Figure 9 This is a block diagram illustrating an example of a video decoder 300, which may be... Figure 7 The video decoder 114 in the system 100 shown.

[0368] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 8 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0369] exist Figure 9 In the example, the video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, the video decoder 300 can perform operations related to the video encoder 200 ( Figure 8 The encoding pass described is roughly the opposite of the decoding pass.

[0370] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded video data blocks). The entropy decoding unit 301 can decode the entropy-coded video data, and based on the entropy-coded video data, the motion compensation unit 302 can determine motion information, including motion vectors, motion vector precision, reference image list indexes, and other motion information. For example, the motion compensation unit 302 can determine such information by executing AMVP and merge modes.

[0371] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. Identifiers for the interpolation filter used at sub-pixel precision can be included in the syntax elements.

[0372] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of the video block to calculate the interpolation of the sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information and use the interpolation filter to generate the prediction block.

[0373] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) slices of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame encoded block, and other information for decoding the encoded video sequence.

[0374] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially neighboring blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0375] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on a display device.

[0376] The following provides a list of preferred solutions for some of the embodiments.

[0377] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 1).

[0378] 1. A video processing method (e.g., Figure 6 The method described in the text (600) includes: conversion between video units and codec representations of a video, determining (602) motion vector difference (MVD) calculation operations for use with merge mode (MMVD) codec tools with motion vector difference based on the characteristics of the video units, and performing (604) conversion based on the determination.

[0379] 2. The method according to Solution 1, wherein the characteristics of the video unit include the dimension of the video unit or the allowed prediction direction.

[0380] 3. The method according to solutions 1-2, wherein the characteristic of the video unit is that only unidirectional prediction is allowed, and based on this characteristic, the computation operation determines a single MVD, regardless of the prediction direction of the underlying merge candidate associated with the MMVD encoding / decoding tool.

[0381] 4. The method according to Solution 2, wherein the computation operation determines a single MVD because the dimension of the video unit satisfies the condition, regardless of the prediction direction of the underlying merge candidate associated with the MMVD encoding / decoding tool.

[0382] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 2).

[0383] 5. The method according to Solution 1, wherein the characteristics of the video unit are that the underlying merge candidate associated with the MMVD encoding / decoding tool is a bidirectional motion vector, and the transformation includes using the internal motion vector difference directly to predict the direction X, where X = 0 or 1, provided that the video unit satisfies the block size condition.

[0384] 6. The method according to Solution 5, wherein the internal motion vector difference is directly used to predict direction 0.

[0385] 7. According to the method described in Solution 5, the block size condition is that the height plus width of the video unit is less than or equal to N, where N is a positive integer less than the total image pixel width.

[0386] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 3).

[0387] 8. A video processing method, comprising:

[0388] Performs a conversion between video units and the encoded / decoded representation of the video, where the conversion uses a motion vector scaling process that depends on the video resolution during operation.

[0389] 9. The method according to Solution 8, wherein the operation includes encoding or decoding using a merge tool with motion vector difference.

[0390] 10. The method according to Solution 8, wherein the operation includes encoding or decoding using a temporal motion vector prediction process.

[0391] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 4).

[0392] 11. A video processing method, comprising:

[0393] The conversion between the video unit and the codec representation of the video is performed using two long-term reference images and a motion vector scaling process.

[0394] 12. The method according to solution 11, wherein the motion vector scaling process is based on the image sequence count of two long-term reference images.

[0395] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 5).

[0396] 13. A video processing method, comprising:

[0397] A merge candidate list is generated for the conversion between video units and the video's codec representation, where non-nearest spatial domain merge candidates of video units are inserted into the merge list; and

[0398] Perform the transformation using the merge candidate list.

[0399] 14. The method according to solution 13, wherein non-nearest spatial domain merge candidates are inserted into the merge list after the history-based merge candidates.

[0400] 15. The method according to solution 13, wherein non-nearest spatial domain merge candidates are inserted into the merge list after pairwise average merge candidates.

[0401] 16. The method according to solution 13, wherein after inserting a temporal merge candidate, and after the number of available merge candidates in the merge list reaches a predefined value, no non-adjacent spatial domain merge candidate is inserted.

[0402] 17. The method according to solution 13, wherein non-adjacent spatial domain merge candidates are inserted in a defined order.

[0403] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 6).

[0404] 18. A method for video processing, comprising:

[0405] A candidate list is generated for the conversion between video units and the video's codec representation. The candidates in this list are generated by averaging M spatially adjacent candidates and N temporally adjacent candidates, where M and N are positive integers.

[0406] Perform the transformation using the merge candidate list.

[0407] 19. The method according to solution 18, wherein M>2.

[0408] 20. The method according to solutions 18-19, wherein M spatial adjacency candidates can be derived from spatial merge candidates included in the merge list.

[0409] 21. The method according to solutions 18-19, wherein N temporal adjacent candidates can be derived from temporal merge candidates included in the merge list.

[0410] 22. The method according to solution 18, wherein M = 3 and N = 1.

[0411] The following solutions show example embodiments of the techniques discussed in the previous chapter (e.g., item 7).

[0412] 23. A video processing method, comprising:

[0413] A merge list is generated for the conversion between video units and the codec representation of the video, wherein the construction process for generating the merge list checks multiple candidates in a defined order; and

[0414] Perform the transformation using the merge candidate list.

[0415] 24. The method according to solution 23, wherein the defined order includes: a first group of spatial merge candidates, spatial-temporal merge candidates, a second group of merge candidates, temporal motion vector predictors, history-based motion vector predictors, pairwise average merge candidate vectors, and zero motion vector merge candidates.

[0416] 25. The method according to solution 23, wherein the defined order includes: spatial merge candidates derived from neighboring blocks (e.g., derived from B, A, C, D, E), TMVP, a first set of spatial merge candidates derived from non-neighboring blocks (e.g., derived from B1, A1, C1, D1, E1), HMVP, pairwise average merge candidates, and zero motion vector merge candidates.

[0417] 26. The method according to any one of solutions 1-25, wherein the video unit comprises a video block or a video codec tree unit or a video transform unit or a video codec unit.

[0418] 27. The method according to any one of solutions 1-26, wherein performing the conversion includes encoding the video to generate a codec representation.

[0419] 28. The method according to any one of solutions 1-26, wherein performing the conversion includes parsing and decoding the codec representation to generate video.

[0420] 29. A video decoding apparatus comprising a processor configured to implement the method according to one or more of solutions 1 to 28.

[0421] 30. A video encoding apparatus comprising a processor configured to implement the method according to one or more of solutions 1 to 28.

[0422] 31. A computer program product having computer code stored thereon, said code causing the processor, when executed by a processor, to perform the method according to any one of solutions 1 to 28.

[0423] 32. The methods, apparatus or systems described in this document.

[0424] Figure 10 A flowchart of an example method for video processing is shown. The method includes: conversion between video units and bitstream of a video; deriving (1002) motion vector difference (MVD) for use in a mode with motion vector difference (MMVD) codec tool based on the characteristics of the video units; and performing (1004) conversion based on the derived MVD.

[0425] In some examples, the characteristics of a video cell include the dimensions or shape of the video cell and / or the allowed prediction orientation of the video cell, and the dimensions of the video cell include the currWidth and currHeight parameters, which respectively indicate the width and height of the video cell.

[0426] In some examples, when the characteristics of a video unit indicate that only unidirectional prediction is allowed for the video unit, a single MVD is derived from the internal MVD associated with the MMVD codec tool, regardless of the prediction direction associated with the underlying merge candidate in the MMVD codec tool, where the internal MVD is derived from the syntax elements of the signaling notification in the bitstream.

[0427] In some examples, if the predictions from only the reference image list X are allowed to predict the direction, denoted by ListX, then the MVD of ListX is derived from the internal MVD, where X is 0 or 1.

[0428] In some examples, the MVD of ListX is set to be equal to the inner MVD.

[0429] In some examples, the MVD of ListX is set to the opposite of the internal MVD.

[0430] In some examples, the MVD for prediction from the reference picture list Y represented by ListY is set to a default value, where Y is 1 or 0.

[0431] In some examples, when the characteristics of a video unit indicate that the dimensions of the video unit satisfy a condition, a single MVD is derived from the internal MVD associated with the MMVD codec tool, regardless of the prediction direction associated with the base merge candidate in the MMVD codec tool, where the internal MVD is an MVD derived from a syntax element signaled in the bitstream.

[0432] In some examples, the condition is that currWidth + currHeight is less than or equal to N, where N is a positive integer.

[0433] In some examples, N = 12, or N = 32.

[0434] In some examples, the condition is that currWidth < N1 or / and currHeight < N2, where N1 and N2 are positive integers.

[0435] In some examples, N1 = N2 = 8.

[0436] In some examples, the condition is that currWidth < N3 * currHeight and / or currHeight < N4 * currWidth, where N3 and N4 are positive integers.

[0437] In some examples, N3 = N4 = 8.

[0438] In some examples, if the characteristics of a video unit indicate that the dimensions or shape of the video unit satisfy one or more conditions, then when the base merge candidate is a bi - directional MV, the internal MVD is always directly used for prediction direction X without scaling, where X = 0 or 1.

[0439] In some examples, the internal MVD is always directly used for prediction direction 0.

[0440] In some examples, the condition is that currWidth + currHeight is less than or equal to N, where N is a positive integer.

[0441] In some examples, N = 12 or N = 32.

[0442] In some examples, the condition is that currWidth < N3 * currHeight and / or currHeight < N4 * currWidth, where N3 and N4 are positive integers.

[0443] In some examples, N3 = N4 = 8.

[0444] In some examples, the inverse value of the internal MVD is used to predict the direction X without scaling, where X = 0 or 1.

[0445] Figure 11 A flowchart of an example method for video processing is shown. The method includes: for the conversion between video units of a video and the bitstream of the video, deriving (1102) a motion vector difference (MVD) using a motion vector (MV) scaling process, wherein the MV scaling process depends on the resolution of the video; and performing (1104) a conversion based on the derived MVD.

[0446] In some examples, MVD is used in merge mode (MMVD) codecs with motion vector difference or temporal motion vector prediction (TMVP) codecs.

[0447] Figure 12 A flowchart of an example method for video processing is shown. The method includes: for the conversion between video units and the bitstream of the video, deriving (1202) motion vector difference (MVD) using a motion vector (MV) scaling process, wherein the MV scaling process uses two long-term reference pictures; and performing (1204) conversion based on the derived MVD.

[0448] In some examples, a video unit includes at least one of a video codec unit (CU), a prediction unit (PU), or a block.

[0449] In some examples, the conversion involves encoding video units of the video into a bitstream.

[0450] In some examples, the conversion involves decoding video units from a bitstream.

[0451] Figure 13 A flowchart of an example method for storing a bitstream of video is shown. The method includes: conversion between video units and a bitstream of video for the purpose of video; deriving (1302) a motion vector difference (MVD) for use in a merge mode (MMVD) codec tool with motion vector difference based on the characteristics of the video units; generating (1304) a bitstream from the video units based on the derived MVD; and storing (1306) the bitstream in a non-transitory computer-readable recording medium.

[0452] Figure 14 A flowchart of an example method for video processing is shown. The method includes: constructing (1402) a merge candidate list for the current block of the video and a bitstream representation of the video for a conversion between the current block and the bitstream representation of the video; wherein non-nearest spatial merge candidates associated with the current block are inserted into the merge candidate list; and performing (1404) a conversion based on the merge candidate list.

[0453] In some examples, non-nearest airspace merge candidates are inserted into the merge list after historical merge candidates.

[0454] In some examples, non-nearest neighbor space merge candidates are inserted into the merge list after pairwise average merge candidates.

[0455] In some examples, if the number of available merge candidates in the merge candidate list reaches a predefined value after inserting a temporal merge candidate, then a non-nearest spatial merge candidate is not inserted.

[0456] In some examples, when inserting a non-adjacent space merge candidate, the process terminates if the number of available merge candidates in the merge candidate list reaches a predefined value.

[0457] In some examples, the predefined value is equal to maxNumMergeCand-N, where maxNumMergeCand represents the size of the merge candidate list and N is a positive integer.

[0458] In some examples, N is set to equal to 2, 3, or 4.

[0459] In some examples, the maximum number of search rounds during the construction of the merge candidate list is set to 1 or 2.

[0460] In some examples, for each round of search, non-adjacent spatial merge candidates are inserted in a predefined insertion order, and the non-adjacent spatial merge candidates include candidate Ai derived from the non-adjacent left block of the current block, candidate Bi derived from the non-adjacent upper block of the current block, candidate Ci derived from the non-adjacent upper right block of the current block, candidate Di derived from the non-adjacent lower left block n of the current block, and candidate Ei derived from the non-adjacent upper left block to the left of the current block, where i is the search round.

[0461] In some examples, the insertion order is Ai, Bi, Ci, Di and Ei, the insertion order is Bi, Ai, Ci, Di and Ei, the insertion order is Bi, Ci, Ai, Di and Ei, or the insertion order is Ai, Di, Bi, Ci and Ei.

[0462] In some examples, all spatial and temporal merge candidates undergo a complete deduplication process using all previous merge candidates in the merge candidate list, and the deduplication process based on historical merge candidates and pairwise average candidates remains unchanged.

[0463] In some examples, all spatial, temporal, historical, and pairwise average merge candidates undergo a complete deduplication process using all previous merge candidates in the merge candidate list.

[0464] In some examples, for non-adjacent spatial domain merge candidates, Ai performs deduplication with Ai-1, Bi performs deduplication with Ai, Ci performs deduplication with Bi, Di performs deduplication with Ai, and Ei performs deduplication with both Ai and Bi. The deduplication process for candidates in the time domain, based on history, and with pairwise averages remains unchanged.

[0465] In some examples, the maximum number of deduplications allowed during the construction of the merge candidate list is denoted as MaxPruningNum, which depends on the size of the merge candidate list and is denoted as maxNumMergeCand.

[0466] In some examples, MaxPruningNum is set to be equal to maxNumMergeCand – M, where M is an integer.

[0467] In some examples, M = 2.

[0468] In some examples, MaxPruningNum is set to be equal to maxNumMergeCand*M, where M is an integer.

[0469] In some examples, M = 2.

[0470] In some examples, the maximum number of deduplications allowed during the construction of the merge candidate list is denoted as MaxPruningNum, which is independent of the size of the merge candidate list and is denoted as maxNumMergeCand.

[0471] In some examples, MaxPruningNum is set to equal to 30 or 35.

[0472] In some examples, the locations of non-adjacent spatial domain merge candidates are constrained to a predefined region.

[0473] In some examples, the predefined region contains the current codec tree unit (CTU) row and four sample rows above the current CTU row.

[0474] In some examples, the predefined region contains the current CTU column and four left sample columns of the current CTU column.

[0475] In some examples, the predefined region includes the current CTU column and the CTU column to the left of the current CTU column.

[0476] In some examples, the location of non-adjacent spatial merge candidates is not constrained in the horizontal direction.

[0477] In some examples, non-adjacent spatial merge candidates are allowed to be used as base merge candidates for merge mode (MMVD) codec tools with motion vector difference.

[0478] In some examples, non-adjacent spatial merge candidates are not allowed to be used as base merge candidates for merge mode (MMVD) codec tools with motion vector difference.

[0479] In some examples, non-nearby spatial domain merge candidates are allowed to generate inter-frame-intra-frame predictions.

[0480] In some examples, non-nearby spatial merge candidates are not allowed to generate inter-frame-intra-frame predictions.

[0481] In some examples, non-nearest spatial domain merge candidates are allowed to generate geometric (GEO) segmentation and / or triangular segmentation merge candidates.

[0482] In some examples, non-nearest spatial domain merge candidates are not allowed to be used to generate geometric (GEO) segmentation and / or triangular segmentation merge candidates.

[0483] In some examples, non-adjacent spatial merge candidates are allowed to be used to generate affine merge candidates.

[0484] In some examples, non-nearby spatial domain merge candidates are allowed to generate advanced motion vector prediction (AMVP) candidates.

[0485] Figure 15 A flowchart of an example method for video processing is shown. The method includes: constructing (1502) a merge candidate list for the current block of the video for a conversion between the current block of the video and the bitstream representation of the video, wherein the process of constructing the merge candidate list checks multiple different kinds of candidates in a defined order; and performing (1504) a conversion based on the merge candidate list.

[0486] In some examples, the defined order includes: a first group of spatial merge candidates, spatial temporal motion vector predictor (STMVP) candidates, a second group of spatial merge candidates, temporal motion vector predictor (TMVP) candidates, history-based motion vector predictor (HMVP) candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0487] In some examples, the first group of spatial merge candidates includes candidate B derived from the block above the current block, candidate A derived from the block to the left of the current block, candidate C derived from the block to the right of the current block, and a fourth candidate D derived from the block to the left of the current block, and the second group of spatial merge candidates includes candidate E derived from the block to the left of the current block.

[0488] In some examples, the defined order includes: spatial merge candidates derived from neighboring blocks of the current block, TMVP candidates, the first set of spatial merge candidates derived from non-neighboring blocks of the current block, HMVP candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0489] In some examples, the spatial merge candidates derived from the neighboring blocks of the current block include candidate B derived from the block above the current block, candidate A derived from the block to the left of the current block, candidate C derived from the block to the right of the current block, and a fourth candidate D derived from the block to the left of the current block, as well as candidate E derived from the block to the left of the current block. The first set of spatial merge candidates derived from the non-neighboring blocks of the current block includes candidate B1 derived from the non-neighboring block above the current block, candidate A1 derived from the non-neighboring block to the left of the current block, candidate C1 derived from the non-neighboring block to the right of the current block, candidate D1 derived from the non-neighboring block to the left of the current block, and candidate E1 derived from the non-neighboring block to the left of the current block. These candidates were derived in the first search round.

[0490] In some examples, the defined order includes: spatial merge candidates derived from neighboring blocks of the current block, TMVP candidates, spatial merge candidates derived from non-neighboring blocks of the current block, HMVP candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0491] In some examples, space merge candidates derived from neighboring blocks of the current block include candidate B derived from the block above the current block, candidate A derived from the block to the left of the current block, candidate C derived from the block to the right of the current block, and a fourth candidate D derived from the block to the left of the current block, as well as candidate E derived from the block to the left of the current block. Space merge candidates derived from non-neighboring blocks of the current block include: candidate B1 derived from the non-neighboring block above the current block, candidate A1 derived from the non-neighboring block to the left of the current block, and candidate E derived from the non-neighboring block to the right of the current block. Candidate C1 derived from the upper block, candidate D1 derived from the non-adjacent lower-left block n of the current block, and candidate E1 derived from the non-adjacent upper-left block to the left of the current block are derived in the first search round; and candidate B2 derived from the non-adjacent upper block of the current block, candidate A2 derived from the non-adjacent left block of the current block, candidate C2 derived from the non-adjacent upper-right block of the current block, candidate D2 derived from the non-adjacent lower-left block n of the current block, and candidate E2 derived from the non-adjacent upper-left block to the left of the current block are derived in the second search round.

[0492] In some examples, the defined order includes: the first set of spatial merge candidates, STMVP candidates, the second set of spatial merge candidates, temporal motion vector predictor (TMVP) candidates, the first set of spatial merge candidates derived from the non-neighboring blocks of the current block, HMVP candidates, pairwise average merge candidates, and zero motion vector merge candidates.

[0493] In some examples, the first set of spatial merge candidates includes candidate B derived from the block above the current block, candidate A derived from the block to the left of the current block, candidate C derived from the block to the upper right of the current block, and a fourth candidate D derived from the block to the lower left of the current block. The first set of spatial merge candidates derived from non-adjacent blocks of the current block includes candidate B1 derived from the block above the current block, candidate A1 derived from the block to the left of the current block, candidate C1 derived from the block to the upper right of the current block, candidate D1 derived from the block to the lower left of the current block, and candidate E1 derived from the block to the upper left of the current block. These candidates were derived in the first search round.

[0494] In some examples, if the corresponding candidate is unavailable, invalid, or the same as or similar to an existing candidate in the merge candidate list, then the corresponding candidate will not be included in the merge candidate list.

[0495] In some examples, the conversion involves encoding video units of the video into a bitstream.

[0496] In some examples, the conversion involves decoding video units from a bitstream.

[0497] Figure 16A flowchart of an example method for storing a bitstream of video is shown. The method includes: constructing (1602) a merge candidate list for the current block of video and a bitstream representation of video for a conversion between the current block and the bitstream representation of video, wherein non-nearest spatial merge candidates associated with the current block are inserted into the merge candidate list; generating (1604) a bitstream from video units based on the merge candidate list; and storing (1606) the bitstream in a non-transitory computer-readable recording medium.

[0498] Figure 17 A flowchart of an example method for video processing is shown. The method includes: constructing (1702) a merge candidate list for the current block of the video and a bitstream representation of the video for a transformation, wherein spatial-temporal motion vector prediction (STMVP) candidates associated with the current block are added to the merge candidate list, and deriving the STMVP candidates as the average of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates, where M and N are positive integers; and performing (1704) a transformation based on the merge candidate list.

[0499] In some examples, M > 2.

[0500] In some examples, spatial neighbor motion candidates are derived from other neighboring blocks that are different from or the same as those used during the construction of the merge candidate list.

[0501] In some examples, spatially adjacent motion candidates are selected from the spatially adjacent motion candidates included in the merge candidate list.

[0502] In some examples, before adding STMVP candidates, spatially adjacent motion candidates are selected from the first M or last M spatially adjacent motion candidates included in the merge candidate list.

[0503] In some examples, before adding STMVP candidates, spatially adjacent motion candidates are selected from the first M or last M merge candidates included in the merge candidate list.

[0504] In some examples, temporally adjacent motion candidates are selected from the temporally adjacent merge candidates included in the merge candidate list.

[0505] In some examples, if the temporal merge candidate is unavailable, then the STMVP candidate is considered unavailable.

[0506] In some examples, whether spatially adjacent motion candidates and / or temporally adjacent motion candidates are considered valid is based on reference image information associated with the current block.

[0507] In some examples, it is considered valid only if its reference index in at least one list of reference images is equal to or not greater than K, where K is an integer.

[0508] In some examples, it is considered valid only if its reference index in both reference image lists is equal to or not greater than K, where K is an integer.

[0509] In some examples, K = 0.

[0510] In some examples, it is not used to derive STMVP candidates when it is considered invalid.

[0511] In some examples, the STMVP candidate is valid if at least one of the first M spatial merge candidates and a corresponding merge candidate are valid.

[0512] In some examples, M=3 and N=1, and the STMVP candidate is derived as the average candidate of the four merge candidates.

[0513] In some examples, if the reference indices of all four merge candidates are valid and equal to 0 in the prediction direction X (where X is 0 or 1), then the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX:

[0514] mvLX=(mvLX_F*a+mvLX_S*b+mvLX_T*c+mvLX_Col*d)>>e,

[0515] Where a, b, c, d, and e are integers.

[0516] In some examples, a, b, c, d, and e are set to equal 1, 1, 1, 1, and 2.

[0517] In some examples, if the reference indices of three of the four merge candidates are valid and equal to 0, X = 0 or 1 in the prediction direction X, then the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX:

[0518] mvLX = (mvLX_F*a + mvLX_S*b + mvLX_Col*c) >> d; or

[0519] mvLX=(mvLX_F*a+mvLX_T*b+mvLX_Col*c)>>d; or

[0520] mvLX=(mvLX_S*a+mvLX_T*b+mvLX_Col*c)>>d,

[0521] Where a, b, c, and d are integers.

[0522] In some examples, a, b, c, and d are set to equal 3, 3, 2, and 3, or a, b, c, and d are set to equal 2, 2, 4, and 3, or a, b, c, and d are set to equal 1, 1, 6, and 3.

[0523] In some examples, if the reference indices of two of the four merge candidates are valid and equal to 0, X = 0 or 1 in the prediction direction X, then the motion vector of the STMVP candidate in the prediction direction X is derived as follows, denoted as mvLX:

[0524] mvLX = (mvLX_F*a + mvLX_Col*b) >> c; or

[0525] mvLX = (mvLX_S*a + mvLX_Col*b) >> c; or

[0526] mvLX=(mvLX_T*a+mvLX_Col*b)>>c,

[0527] Where a, b, and c are integers.

[0528] In some examples, a, b, and c are set to equal 1, 1, and 1, respectively.

[0529] In some examples, the STMVP candidate is deduplicated using all previous merge candidates from the merge candidate list.

[0530] In some examples, the STMVP candidate is not deduplicated by other merge candidates.

[0531] In some examples, the STMVP candidate is deduplicated using only the merge candidate above and to the left.

[0532] In some examples, STMVP candidates reference one or two specific reference images.

[0533] In some examples, a specific reference image is the reference image with a reference index of 0 in the reference list.

[0534] In some examples, a specific reference image is a reference image of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates with the smallest reference index in the reference list.

[0535] In some examples, the conversion involves encoding video units of the video into a bitstream.

[0536] In some examples, the conversion involves decoding video units from a bitstream.

[0537] Figure 18A flowchart of an example method for storing a bitstream of video is shown. The method includes: constructing (1802) a merge candidate list for the current block of video and a bitstream representation of video for a conversion between the current block and the bitstream representation of video, wherein spatial-temporal motion vector prediction (STMVP) candidates associated with the current block are added to the merge candidate list, and deriving the STMVP candidates as an average candidate of M spatially adjacent motion candidates and / or N temporally adjacent motion candidates, where M and N are positive integers; generating (1804) a bitstream from video units based on the merge candidate list; and storing (1806) the bitstream in a non-transitory computer-readable recording medium.

[0538] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm may be applied during the conversion from a pixel representation of a video to its corresponding bitstream representation, and vice versa. The bitstream representation of the current video block may, for example, correspond to bits located at the same position or distributed at different positions within the bitstream, as defined in the syntax. For example, macroblocks may be encoded based on the error residuals from the transform and encoding / decoding, and may also utilize bits in the header and other fields in the bitstream.

[0539] The solutions, examples, embodiments, modules, and functional operations disclosed in this document, as well as other solutions, examples, embodiments, modules, and functional operations, can be implemented in digital electronic circuits or computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or one or more combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by or control of the operation of a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a composition of substances influencing machine-readable propagation signals, or one or more combinations thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer program under consideration, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or one or more combinations thereof. Propagation signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0540] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language, including compiled or interpreted languages, and can be deployed in any form, including as standalone programs or modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as a portion of a file containing other programs or data (e.g., one or more scripts stored in a markup language document), a single file dedicated to the program in question, or multiple coordinating files (e.g., files storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on a single computer, at a single location, or distributed across multiple locations on multiple computers interconnected by a communication network.

[0541] The processes and logic flows described in this document can be executed by one or more programmable processors executing one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by dedicated logic circuitry, and the devices can be implemented as dedicated logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0542] By way of example, a processor suitable for executing computer programs includes both general-purpose and special-purpose microprocessors, as well as any one or more processors of any kind of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer may also include, or be operatively coupled to, one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, receiving data from them, transferring data to them, or both. However, a computer does not need to have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or integrated into dedicated logic circuitry.

[0543] While this patent document contains numerous details, these should not be construed as limiting the scope of any subject matter or what may be claimed, but rather as descriptions of features that may be characteristic of specific embodiments of a particular technology. Certain features described in this patent document within the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple embodiments or in any suitable sub-combination. Furthermore, although features may be described above as operating in certain combinations and even initially claimed so, in some cases, one or more features from a claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0544] Similarly, although operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring such operations to be performed in the specific order shown or in a sequential order, or as requiring all shown operations to be performed to achieve the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0545] Only a few specific implementations and examples are described, and other specific implementations, enhancements and variations may be made based on the content described and illustrated in this patent document.

Claims

1. A video processing method, comprising: For the conversion between the current block of the video and the bitstream of the video, a merge candidate list for the current block is constructed, wherein non-nearest spatial domain merge candidates associated with the current block are inserted into the merge candidate list; as well as The transformation is performed based on the merge candidate list; Specifically, after averaging the candidate merges in pairs, the candidate merges in the non-nearest spatial domains are inserted into the merge list.

2. The method of claim 1, wherein the non-nearest spatial domain merge candidate is inserted into the merge list after the history-based merge candidate.

3. The method of claim 1, wherein if the number of available merge candidates in the merge candidate list reaches a predefined value after inserting the temporal merge candidate, then the non-adjacent spatial domain merge candidate is not inserted.

4. The method according to claim 1, wherein when inserting the non-adjacent space merge candidate, the process terminates if the number of available merge candidates in the merge candidate list reaches a predefined value.

5. The method according to claim 3 or 4, wherein the predefined value is equal to maxNumMergeCand-N, where maxNumMergeCand represents the size of the merge candidate list and N is a positive integer.

6. The method of claim 5, wherein N is set to be equal to 2, 3 or 4.

7. The method of claim 1, wherein the maximum number of search rounds during the construction of the merge candidate list is set to be equal to 1 or 2.

8. The method according to claim 7, wherein for each search round, the non-adjacent spatial domain merge candidates are inserted in a predefined insertion order, and the non-adjacent spatial domain merge candidates include candidate Ai derived from the non-adjacent left block of the current block, candidate Bi derived from the non-adjacent upper block of the current block, candidate Ci derived from the non-adjacent upper right block of the current block, candidate Di derived from the non-adjacent lower left block of the current block, and candidate Ei derived from the non-adjacent upper left block of the current block, where i is the search round.

9. The method according to claim 8, wherein the insertion order is Ai, Bi, Ci, Di, and Ei. The insertion order is Bi, Ai, Ci, Di, and Ei. The insertion order is Bi, Ci, Ai, Di, and Ei, or The insertion order is Ai, Di, Bi, Ci, and Ei.

10. The method of claim 1, wherein all spatial and temporal merge candidates perform a complete deduplication process using all previous merge candidates in the merge candidate list, and the deduplication process based on historical merge candidates and pairwise average candidates remains unchanged.

11. The method of claim 1, wherein all spatial, temporal, historical, and pairwise average merge candidates perform a complete deduplication process using all previous merge candidates in the merge candidate list.

12. The method of claim 8, wherein for the non-adjacent spatial domain merge candidates, Ai performs deduplication using Ai-1, Bi performs deduplication using Ai, Ci performs deduplication using Bi, Di performs deduplication using Ai, Ei performs deduplication using both Ai and Bi, and the deduplication process for candidates in the time domain, based on history, and with pairwise averages remains unchanged.

13. The method of claim 1, wherein the maximum number of deduplications allowed during the construction of the merge candidate list is denoted as MaxPruningNum, which depends on the size of the merge candidate list and is denoted as maxNumMergeCand.

14. The method of claim 13, wherein MaxPruningNum is set to be equal to maxNumMergeCand-M, where M is an integer.

15. The method of claim 14, wherein M=2.

16. The method of claim 13, wherein MaxPruningNum is set to be equal to maxNumMergeCand*M, where M is an integer.

17. The method of claim 16, wherein M=2.

18. The method of claim 1, wherein the maximum number of deduplications allowed during the construction of the merge candidate list is denoted as MaxPruningNum, which is independent of the size of the merge candidate list and is denoted as maxNumMergeCand.

19. The method of claim 18, wherein MaxPruningNum is set to be equal to 30 or 35.

20. The method of claim 1, wherein the location of the non-adjacent spatial domain merge candidate is constrained within a predefined region.

21. The method of claim 20, wherein the predefined region comprises the current codec tree unit (CTU) row and four sample rows above the current CTU row.

22. The method of claim 20, wherein the predefined region comprises the current CTU column and four left sample point columns of the current CTU column.

23. The method of claim 20, wherein the predefined region comprises the current CTU column and the CTU column to the left of the current CTU column.

24. The method of claim 1, wherein the position of the non-adjacent spatial domain merge candidate is not constrained in the horizontal direction.

25. The method of claim 1, wherein the non-adjacent spatial domain merge candidate is allowed to be used as a base merge candidate for a merge mode (MMVD) codec tool with motion vector difference.

26. The method of claim 1, wherein the non-adjacent spatial domain merge candidate is not permitted to be used as the base merge candidate for a merge mode (MMVD) codec tool with motion vector difference.

27. The method of claim 1, wherein the non-nearest spatial domain merge candidate is allowed to generate inter-frame-intra-frame predictions.

28. The method of claim 1, wherein the non-nearest spatial domain merge candidate is not permitted to be used to generate inter-frame-intra-frame predictions.

29. The method of claim 1, wherein the non-nearest spatial domain merge candidates are allowed to be used to generate geometric (GEO) segmentation and / or triangular segmentation merge candidates.

30. The method of claim 1, wherein the non-nearest spatial domain merge candidates are not permitted to be used to generate geometric (GEO) segmentation and / or triangular segmentation merge candidates.

31. The method of claim 1, wherein the non-adjacent spatial domain merge candidate is allowed to be used to generate affine merge candidate.

32. The method of claim 1, wherein the non-nearest spatial domain merge candidate is allowed to generate advanced motion vector prediction (AMVP) candidates.

33. The method of claim 1, wherein the conversion comprises encoding the video units of the video into the bitstream.

34. The method of claim 1, wherein the conversion comprises decoding the video unit of the video from the bitstream.

35. An apparatus for processing video data, comprising a processor and a non-transitory memory thereon having instructions, wherein the instructions, when executed by the processor, cause the processor to: For the conversion between the current block of the video and the bitstream of the video, a merge candidate list for the current block is constructed, wherein non-nearest spatial domain merge candidates associated with the current block are inserted into the merge candidate list; and The transformation is performed based on the merge candidate list; in, After pairwise average merge candidates, the non-adjacent spatial domain merge candidates are inserted into the merge list.

36. A non-transitory computer-readable medium storing instructions that cause a processor to: For the conversion between the current block of the video and the bitstream of the video, a merge candidate list for the current block is constructed, wherein non-nearest spatial domain merge candidates associated with the current block are inserted into the merge candidate list; and The transformation is performed based on the merge candidate list; in, After pairwise average merge candidates, the non-adjacent spatial domain merge candidates are inserted into the merge list.

37. A non-transitory computer-readable medium having a computer program and a bit stream stored thereon, the computer program, when executed by a processor, implementing the following method to generate the bit stream: For the current block of the video, a merge candidate list is constructed, wherein non-nearest spatial domain merge candidates associated with the current block are inserted into the merge candidate list; and The bitstream is generated from the video unit based on the merge candidate list; in, After pairwise average merge candidates, the non-adjacent spatial domain merge candidates are inserted into the merge list.

38. A method for storing a bitstream of video, comprising: For the current block of the video, a merge candidate list for the current block is constructed, wherein non-nearest spatial domain merge candidates associated with the current block are inserted into the merge candidate list; as well as The bitstream is generated from the video unit based on the merge candidate list; as well as The bitstream is stored in a non-transitory computer-readable recording medium; Specifically, after averaging the candidate merges in pairs, the candidate merges in the non-nearest spatial domains are inserted into the merge list.

Citation Information

Patent Citations

  • Encoding / decoding method and device of motion information

    CN109495738A

  • Combining history-based motion vector prediction and non-adjacent merge prediction

    US10362330B1