Motion vector difference of geometrically segmented blocks

By introducing the Geometric Motion Vector Differential (GMVD) encoding and decoding method into video encoding and decoding, and by notifying or deriving multiple motion vector differences (MVDs) for video block signaling, the problem of inaccurate direct inheritance of motion vectors in the prior art is solved, thereby improving the accuracy and efficiency of encoding and decoding.

CN115428452BActive Publication Date: 2025-12-02DOUYIN VISION CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180026287.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-09
Filing Date
2021-04-08
Publication Date
2025-12-02
Estimated Expiration
2041-04-08

AI Technical Summary

Technical Problem

In existing video codec standards, directly inheriting Merge candidates from segmented motion vectors may be inaccurate, leading to insufficient codec efficiency and quality.

Method used

The geometric motion vector differential encoding and decoding (GMVD) method is introduced to refine the motion compensation process by signaling or deriving multiple motion vector differences (MVDs) for video blocks and adding them to the motion vectors of the merge candidate.

Benefits of technology

It improves the accuracy and efficiency of video encoding and decoding, especially when using geometric segmentation mode, it enhances the refinement of motion vectors and improves the quality of encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115428452B_ABST
    Figure CN115428452B_ABST
Patent Text Reader

Abstract

The motion vector difference (MVD) of a block with geometric segmentation is described. An example of a video processing method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded and decoded using a geometric segmentation mode; deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived from a merge candidate associated with the current video block to a motion vector (MV) derived from a merge candidate associated with the current video block, the MV being related to an offset distance and / or offset direction; and performing the conversion based on the refined MV.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application is based on International Patent Application No. PCT / CN2021 / 085914, filed on April 8, 2021, which claims priority and interest in International Patent Application No. PCT / CN2020 / 083916, filed on April 9, 2020. All of the aforementioned patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] This patent document relates to the encoding and decoding of images and videos. Background Technology

[0004] Digital video consumes the most bandwidth on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used in video encoders and decoders to perform cross-component adaptive loop filtering during video encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed. The method includes: for a conversion between video units of a video and a codec representation of the video, segmenting the video into multiple segments, and for at least some of the multiple segments, performing the conversion based on the multiple segments using a Merge pattern with motion vector differences.

[0007] In another example, a different video processing method is disclosed. This method includes: for the conversion between video units of a video and a codec representation of the video, segmenting the video into multiple segments, and performing the conversion based on the multiple segments, wherein the codec representation includes a field indicating motion vector differences, the field indicating motion vector differences using one or more of two variables corresponding to the direction and magnitude of the motion vector differences.

[0008] In another example, a different video processing method is disclosed. This method includes: segmenting the video into multiple segments for a conversion between video units and a codec representation of the video, and performing the conversion based on the multiple segments, wherein the codec representation contains fields arranged in a sequential order.

[0009] In another example, a different video processing method is disclosed. The method includes: for a conversion between a current video block and the bitstream of the current video, determining that the current video block is encoded and decoded using a geometric segmentation mode; deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to motion vectors (MVs) derived from Merge candidates associated with the current video block; and performing the conversion based on the refined MV.

[0010] In another example, a method for storing a bitstream of video is disclosed. The method includes: determining, for a conversion between a current video block and the bitstream of the current video, that the current video block is encoded and decoded using a geometric segmentation mode; deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to motion vectors (MVs) derived from Merge candidates associated with the current video block; generating the bitstream from the current video block based on the refined MV; and storing the bitstream in a non-transitory computer-readable recording medium.

[0011] In another example, a different video processing method is disclosed. The method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded and decoded using a geometric segmentation mode; deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block, the MV being related to an offset distance and / or offset direction; and performing the conversion based on the refined MV.

[0012] In another example, a different video processing method is disclosed. The method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded and decoded using a geometric segmentation mode; determining whether a geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD from a plurality of motion vector differences (MVDs) signaled or derived for the current video block to motion vectors (MVs) derived from Merge candidates associated with the current video block; and performing the conversion based on the refined MV.

[0013] In another example, a method for storing a bitstream of video is disclosed. The method includes: determining, for a conversion between a current video block and the bitstream of the video, that the current video block is encoded using a geometric segmentation mode; determining whether a geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD from a plurality of motion vector differences (MVDs) signaled or derived for the current video block to motion vectors (MVs) derived from Merge candidates associated with the current video block; generating the bitstream from the current video block based on the refined MV; and storing the bitstream in a non-transitory computer-readable recording medium.

[0014] In another example aspect, a video encoding apparatus is disclosed. This video encoding apparatus includes a processor configured to implement the method described above.

[0015] In another example aspect, a video decoding apparatus is disclosed, which includes a processor configured to implement the method described above.

[0016] In another example, a computer-readable medium is disclosed that stores code. This code embodies one of the methods described herein in the form of processor-executable code.

[0017] This article describes the above and other features throughout. Attached Figure Description

[0018] Figure 1 An example of the location of the spatial merge candidate is shown.

[0019] Figure 2 An example of candidate pairs considered for redundancy checks in the spatial merge candidate is shown.

[0020] Figure 3 This is a diagram illustrating motion vector scaling for temporal Merge candidates.

[0021] Figure 4 Examples of candidate positions for time-domain Merge candidates C0 and C1 are shown.

[0022] Figure 5 An example of inter-frame prediction based on triangulation is shown.

[0023] Figure 6 An example of unidirectional prediction MV selection for triangular segmentation pattern is shown.

[0024] Figure 7 An example of weights used in hybrid processing is shown.

[0025] Figure 8 This section explains the recommendations and TPM design in VTM-6.0.

[0026] Figure 9 An example of GEO boundary description is shown.

[0027] Figure 10A The edges supported in GEO are shown.

[0028] Figure 10B The geometric relationship between a given sample point location (x, y) and the two edges is shown.

[0029] Figure 11 An example of UMVE search processing is shown.

[0030] Figure 12 An example of a UMVE search point is shown.

[0031] Figure 13 This is a block diagram of an example video processing system.

[0032] Figure 14 This is a block diagram of an example hardware platform used for video processing.

[0033] Figure 15 This is a flowchart illustrating an example of a video processing method.

[0034] Figure 16 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0035] Figure 17 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0036] Figure 18 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0037] Figure 19 This is a flowchart illustrating an example of a video processing method.

[0038] Figure 20 This is a flowchart illustrating an example of a method for storing video bitstreams.

[0039] Figure 21 This is a flowchart illustrating an example of a video processing method.

[0040] Figure 22 This is a flowchart illustrating an example of a video processing method.

[0041] Figure 23 This is a flowchart illustrating an example of a method for storing video bitstreams. Detailed Implementation

[0042] In this document, chapter headings are used for ease of understanding and do not limit the applicability of the technologies and embodiments disclosed in each chapter to that chapter. Furthermore, the use of H.266 terminology in some descriptions is for ease of understanding only and not intended to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.

[0043] 1 Overview

[0044] This document relates to video codec technology. Specifically, it concerns inter-frame prediction and related techniques in video codec. This technology can be applied to existing video codec standards, such as HEVC, or upcoming (General Video Codec) standards. It can also be applied to future video codec standards or codecs.

[0045] 2 Background

[0046] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 video. The two organizations jointly developed the H.262 / MPEG-2 video standard, the H.264 / MPEG-4 Advanced Video Coding (AVC) standard, and the H.265 / HEVC standard. Since H.262, video codec standards have been based on a hybrid video codec architecture, employing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Team (JVET) in 2015. Since then, JVET has adopted many new methods and applied them to reference software called the Joint Exploration Model (JEM). JVET meetings are held quarterly, and the goal of the new codec standard is a 50% reduction in bitrate compared to HEVC. The new video codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. With ongoing efforts to standardize VVC, new codec technologies have been introduced for the VVC standard at each JVET meeting. The VVC working draft and test model VTM are updated after each meeting. The current goal of the VVC project is to finalize the Technical Documentation Standard (FDIS) at the July 2020 meeting.

[0047] 2.1 Extended Merge Forecast

[0048] In VTM, the Merge candidate list is constructed by sequentially including the following five candidates:

[0049] 1) Spatial MVP from adjacent CUs in the airspace;

[0050] 2) Temporal MVP from co-located CUs;

[0051] 3) History-based MVP from FIFO table;

[0052] 4) Paired average MVP;

[0053] 5) Zero MV.

[0054] The size of the merge list is signaled in the stripe header, and the maximum allowed size of the merge list in VTM is 6. For each CU code in Merge mode, the index of the best merge candidate is encoded using truncated univariate binarization (TU). The first bin of the merge index is encoded and decoded using context, while bypass encoding and decoding are used for the other bins.

[0055] This section provides the process for generating Merge candidates for each category.

[0056] 2.1.1 Derivation of Airspace Candidates

[0057] The derivation of spatial merge candidates in VVC is the same as that in HEVC. Figure 1 At most four merge candidates are selected from the candidates at the indicated positions. The derivation order is A0, B0, B1, A1, and B2. Position B2 is considered only if any CU at positions A0, B0, B1, or A1 is unavailable (e.g., because it belongs to another stripe or slice) or if it is intra-frame encoding / decoding. After adding the candidate at position A1, a redundancy check is performed on the addition of the remaining candidates to ensure that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only... Figure 2 The pairs connected by the middle arrow are added to the list only if the corresponding candidate used for redundancy checking does not have the same motion information.

[0058] Figure 1 An example of the location of the spatial merge candidate is shown.

[0059] Figure 2 An example of candidate pairs considered for redundancy checks in the spatial merge candidate is shown.

[0060] 2.1.2 Time-domain candidate derivation

[0061] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, the scaled motion vector is derived based on the juxtaposed CUs belonging to the juxtaposed reference image. The signaling in the strip header explicitly notifies the list of reference images used for deriving the juxtaposed CUs. The scaled motion vectors used for the temporal merge candidate are obtained, such as... Figure 3 As shown by the dashed lines, the scaled motion vector is obtained by scaling the motion vector of the juxtaposed CU using POC distances tb and td, where tb is defined as the POC difference between the reference image and the current image, and td is defined as the POC difference between the reference image and the juxtaposed image. The reference image index of the temporal merge candidate is set to zero.

[0062] Figure 3 This is a diagram illustrating motion vector scaling for temporal Merge candidates.

[0063] Select the position of the time-domain candidate between candidate C0 and C1, such as Figure 4 As shown. If the CU at position C0 is unavailable, is intra-frame encoded, or is outside the current CTU line, then position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0064] 2.1.3 Historical Merge Candidate Derivation

[0065] Historically based MVP (HMVP) merge candidates are added to the merge list after the spatial MVP and TMVP. In this method, motion information from previous codec blocks is stored in a table and used as the MVP for the current CU. A table with multiple HMVP candidates is maintained during encoding / decoding. This table is reset (cleared) when a new CTU row is encountered. Whenever a non-sub-block inter-frame codec CU exists, the associated motion information is added to the last entry of this table as a new HMVP candidate.

[0066] In VTM, the size S of the HMVP table is set to 6, meaning that a maximum of 6 history-based MVP (HMVP) candidates can be added to the table. When inserting a new motion candidate into the table, a restricted first-in-first-out (FIFO) rule is used, where a redundancy check is first applied to see if the same HMVP already exists in the table. If it does, the same HMVP is removed from the table, and then all subsequent HMVP candidates are moved forward.

[0067] HMVP candidates can be used during the Merge candidate list construction process. The latest few HMVP candidates in the table are checked sequentially, and these latest HMVP candidates are inserted into the candidate list after the TMVP candidates. Redundancy checks are applied from HMVP candidates to spatial or temporal Merge candidates.

[0068] To reduce the number of redundant check operations, the following simplifications are introduced:

[0069] 1. Set the number of HMPV candidates used to generate the Merge list to (N<=4)? M: (8-N), where N represents the number of existing candidates in the Merge list and M represents the number of available HMVP candidates in the table.

[0070] 2. Once the total number of available Merge candidates reaches the maximum allowed number of Merge candidates minus 1, the process of building the Merge candidate list from HMVP is terminated.

[0071] 2.1.4 Derivation of Pairwise Average Merge Candidates

[0072] Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing Merge candidate list. These predefined candidate pairs are defined as {(0,1),(0,2),(1,2),(0,3),(1,3),(2,3)}, where the numbers represent the Merge indices in the Merge candidate list. The average motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid.

[0073] If the Merge list is not full after adding pairwise average Merge candidates, insert zero MVP at the end until the maximum number of Merge candidates is reached.

[0074] 2.2 Triangulation for Inter-Frame Prediction

[0075] In VTM, inter-frame prediction supports Triangle Partition Mode (TPM). Triangle Partition Mode is only applicable to CUs with 64 samples or more that are encoded and decoded in Skip or Merge mode instead of in regular Merge mode, MMVD mode, CIIP mode, or Subblock Merge mode. CU level flags are used to indicate whether Triangle Partition Mode is applied.

[0076] When using this mode, use either diagonal division or anti-diagonal division. Figure 5The CU is divided into two triangular segments. Each triangular segment in the CU uses its own motion for inter-frame prediction; each segment allows only unidirectional prediction, meaning each segment has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, each CU requires only two motion-compensated predictions. The unidirectional prediction motion of each segment directly derives the Merge candidate list constructed in 2.1 for extended Merge prediction, and the unidirectional prediction motion is selected from the given Merge candidates in the list according to the process in 2.2.1.

[0077] If the triangulation pattern is used for the current CU, further signaling is provided indicating the direction of the triangulation (diagonal or anti-diagonal) and two Merge indices (one Merge index per segment). After predicting each triangulation, a blending process with adaptive weights is used to adjust the sample values ​​along the diagonal or anti-diagonal edges. This is the predicted signal for the entire CU, and like other prediction patterns, the transformation and quantization processes are applied to the entire CU. Finally, the motion field of the CU predicted using the triangulation pattern is stored in 4×4 cells, as shown in 2.2.3.

[0078] 2.2.1 Construction of One-Way Prediction Candidate List

[0079] Given the Merge candidate index, the unidirectional predicted motion vector is derived from the Merge candidate list constructed for the extended Merge prediction using the procedure in 2.1, as follows: Figure 6 As shown. For a candidate in the list, the LX motion vector with parity equal to the Merge candidate index value is used as the unidirectional predicted motion vector for the triangular segmentation pattern. These motion vectors in Figure 6 The L(1-X) motion vector is marked with "x". In the absence of a corresponding LX motion vector, the L(1-X) motion vector of the same candidate in the extended Merge prediction candidate list is used as the unidirectional prediction motion vector of the triangular segmentation mode.

[0080] 2.2.2 Blending along the edges of the triangle segment

[0081] After each triangulation is predicted using its own motion, a blend is applied to the two predicted signals to derive samples around the diagonal or anti-diagonal edges. The following weights are used in the blending process:

[0082] ●Brightness values ​​are {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8}, and chromaticity values ​​are {6 / 8, 4 / 8, 2 / 8}, such as... Figure 7 As shown.

[0083] Figure 7 An example of weights used in hybrid processing is shown.

[0084] 2.2.3 Sports Field Storage

[0085] The motion field of the CU encoded using the triangular segmentation pattern is stored in 4×4 cells. Based on the position of each 4×4 cell, a unidirectional or bidirectional predicted motion vector is stored. Mv1 and Mv2 are represented as the unidirectional predicted motion vectors of segment 1 and segment 2, respectively. If the 4×4 cell is located in... Figure 7 In the example, in the unweighted region, either Mv1 or Mv2 is stored for the 4×4 cell. Otherwise, if the 4×4 cell is in the weighted region, the bidirectional predicted motion vector is stored. The bidirectional predicted motion vector is derived from Mv1 and Mv2 according to the following procedure:

[0086] 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then simply combine Mv1 and Mv2 to form a bidirectional predicted motion vector.

[0087] 2) Otherwise, if Mv1 and Mv2 come from the same list, and without loss of generality, assume that Mv1 and Mv2 both come from L0. In this case,

[0088] 2.a) If a reference image for Mv2 (or Mv1) appears in L1, then use the reference image in L1 to convert Mv2 (or Mv1) into an L1 motion vector. Then combine the two motion vectors to form a bidirectional predicted motion vector;

[0089] Otherwise, instead of bidirectional predicted motion, only unidirectional predicted motion Mv1 is stored.

[0090] 2.3 Geometric Segmentation (GEO) for Inter-Frame Prediction

[0091] At the 15th JVET conference in Gothenburg, the Geometric Merge Model (GEO) was proposed as an extension of the existing Triangulation Prediction Model (TPM). At the 16th JVET conference in Geneva, a simpler-designed GEO model was selected as the CE anchor point for further study. Currently, the GEO model is being investigated as a replacement for the existing TPM in VVC.

[0092] Figure 8 This section explains the TPM in VTM-6.0 and the additional shapes proposed for non-rectangular inter-frame blocks.

[0093] The boundary of the geometric Merge pattern is determined by angles. and distance offset ρ i To describe, such as Figure 9 As shown. Angle Represents the quantized angle between 0 and 360 degrees, distance offset ρ i Represents the maximum distance ρmax The quantization offset. Furthermore, partitioning directions that overlap with binary tree partitioning and TPM partitioning are excluded.

[0094] GEO is applied to blocks of size 8×8 or larger. For each block size, there are 82 different ways to divide the blocks, distinguished by 24 angles and 4 edges relative to the center of the CU. Figure 10A The display starts at Edge0, which passes through the center of the CU, with four edges evenly distributed within the CU along the normal vector direction. Each segmentation pattern in the GEO (i.e., a pair of angle indices and edge indices) is assigned a sample adaptive weight table to blend samples from the two segmentation parts. The sampled weight values ​​range from 0 to 8 and are determined by the distance L2 from the center position of the sample to the edge. Essentially, a unity-gain constraint is followed when assigning weight values, meaning that when a small weight value is assigned to one GEO segment, a large complementary value is assigned to the other segment, summing to 8.

[0095] The calculation of the weight value for each sample point consists of two parts: (a) calculating the displacement from the sample point location to the given edge, and (c) mapping the calculated displacement to the weight value using a predefined lookup table. The method for calculating the displacement from the sample point location (x,y) to the given edge Edgei is actually the same as the method for calculating the displacement from (x,y) to Edge0 and subtracting the distance ρ between Edge0 and Edgei from that displacement. Figure 10B This illustrates the geometric relationship between (x, y) and the edge. Specifically, the displacement from (x, y) to the edge can be expressed as follows:

[0096]

[0097] Figure 10A The edges supported in GEO are shown. Figure 10B The geometric relationship between a given sample point location (x, y) and the two edges is shown.

[0098] The value of ρ is a function of the maximum length of the normal vector (denoted by ρmax) and the edge index i, that is:

[0099]

[0100] Where N is the number of edges supported by GEO, and "1" is to prevent the last edge, EdgeN-1, from being too close to the CU angle for some angular indices. Substituting Equation (8) into Equation (6), the displacement from each sample point (x, y) to a given Edgei can be calculated. In short, This is represented as wIdx(x,y). ρ needs to be calculated once for each CU, and wIdx(x,y) needs to be calculated once for each sample point, which involves multiplication.

[0101] 2.4 Merge (MMVD) with Motion Vector Difference

[0102] MMVD is also known as Ultimate Motion Vector Expression (UMVE).

[0103] UMVE is proposed. UMVE can be used for skip or merge modes using the proposed motion vector representation method.

[0104] UMVE reuses the same Merge candidates as those included in the regular Merge candidate list in VVC. Within the Merge candidates, a base candidate can be selected and further extended using the proposed motion vector representation method.

[0105] UMVE provides a novel representation of motion vector difference (MVD), which uses the origin, amplitude, and direction of motion to represent MVD.

[0106] Figure 11 An example of UMVE search processing is shown.

[0107] Figure 12 An example of a UMVE search point is shown.

[0108] The proposed technique uses the Merge candidate list as is. However, it only considers candidates of the default Merge type (MRG_TYPE_DEFAULT_N) for UMVE extensions.

[0109] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.

[0110] Table 2.4.1 Basic Candidate IDX

[0111] Basic candidate IDX 0 1 2 3 The Nth MVP First MVP Second MVP 3rd MVP 4th MVP

[0112] If the number of basic candidates is equal to 1, then no signaling is sent to the basic candidate IDX.

[0113] The distance index is motion amplitude information. The distance index represents a predefined distance from the starting point. The predefined distances are shown below:

[0114] Table 2.4.2a Distance to IDX

[0115] Distance from IDX 0 1 2 3 4 5 6 7 Sample distance 1 / 4 pixel 1 / 2 pixel 1 pixel 2 pixels 4 pixels 8 pixels 16 pixels 32 pixels

[0116] During entropy encoding and decoding, truncated unary codes are used to binarize the distance IDX in bin as follows:

[0117] Table 2.4.2b Distance IDX Binarization

[0118] Distance from IDX 0 1 2 3 4 5 6 7 Binarization 0 10 110 1110 11110 111110 1111110 1111111

[0119] In arithmetic encoding and decoding, the first bin uses probabilistic context encoding and decoding, while subsequent bins use equal probability model encoding and decoding, also known as bypass encoding and decoding.

[0120] The direction index indicates the direction of MVD relative to the starting point. The direction index can represent four directions, as shown below.

[0121] Table 2.4.3 Directional IDX

[0122] Directional IDX 00 01 10 11 X-axis + – N / A N / A Y-axis N / A N / A + –

[0123] Immediately after sending the skip or merge flag, signal the UMVE flag. If the skip or merge flag is true, the UMVE flag is resolved. If the UMVE flag is equal to 1, the UMVE syntax is resolved. However, if the UMVE flag is not 1, the AFFINE flag is resolved. If the AFFINE flag is equal to 1, it indicates an AFFINE mode. However, if the AFFINE flag is not equal to 1, the skip / merge index is resolved for the VTM's skip / merge mode.

[0124] No additional line buffer is needed due to UMVE candidates. Because the software's skip / merge candidates are used directly as the base candidates, the MV supplement can be determined before motion compensation using the input UMVE index. There is no need to reserve a long line buffer for this.

[0125] In VVC, the first or second Merge candidate in the Merge candidate list can be selected as the base candidate.

[0126] Figure 11 and Figure 12 The MVD search process and search point for UMVE are shown.

[0127] 3. Technical problems solved by the technical solutions disclosed in this paper

[0128] In the current TPM / GEO design, the segmented MV is directly inherited from the Merge candidate without any refinement, which may be inaccurate.

[0129] 4 Solution Examples

[0130] The detailed items below should be considered as examples to illustrate general concepts. These items should not be interpreted narrowly. Furthermore, these items can be combined in any way.

[0131] The term "GEO" can refer to an encoding / decoding method that divides a block into two or more sub-regions, which is not possible with segmentation types (e.g., QT / BT / TT). The term "GEO" can also be referred to as "GPM". The term "GPM" can represent Triangular Prediction Mode (TPM), and / or Geometric Merge Mode (GEO), and / or Wedge Prediction Mode.

[0132] The term "block" can refer to a codec block of CU and / or PU and / or TU.

[0133] A coding / decoding method called "GMVD" is proposed. In GMVD, more than one motion vector difference (MVD) can be derived for a block signaling notification. At least one MVD from the multiple MVDs is added to the MV derived from the merge candidate to obtain a refined MV. This refined MV is used for motion compensation of at least one sample within the block and / or stored for subsequent processing. It should be noted that the following bullet points can be applied to either GMVD or the traditional MMVD method.

[0134] Vector operations in this document are performed in a manner common to mathematics. For example, (a,b)*c is equal to (a*c,b*c).

[0135] 1. A GMVD codec block can be divided into more than one segment.

[0136] a. In one example, at least one segment is non-rectangular.

[0137] b. In one example, the block is split into two segments by TPM, and the merge candidate can be a TPM merge candidate.

[0138] c. In one example, the block is split into two segments via GEO, and the Merge candidate can be a GEO Merge candidate.

[0139] 2. When a GMVD codec block is divided into N segments, signal or derive at most N MVDs for these segments, where N>=2. In one example, N equals 2.

[0140] a. In one example, signaling informs or derives N MVDs, and each MVD corresponds to a specific segment.

[0141] i. Alternatively, further, the MVD corresponding to a segment is added to the MV of the segment to obtain a refined MV of the segment, which is used for motion compensation of the segment.

[0142] b. In one example, K (K < N) MVDs are signaled or derived, and one MVD can correspond to more than one partition.

[0143] c. In one example, how many MVDs are signaled may depend on decoding information such as block size, low-delay check flag, prediction direction, or reference picture list associated with N partitions.

[0144] d. In one example, the number of MVDs is signaled.

[0145] e. In one example, one partition can use multiple MVDs.

[0146] i. For example, multiple MVDs can be added to the MV of a partition to obtain a refined MV of the partition.

[0147] 3. It is proposed to signal a first message (e.g., a flag) to indicate whether at least one of the multiple MVDs is not equal to 0.

[0148] a. In one example, the first message is context-coded or bypass-coded in arithmetic coding. <00004​​​​​​​​​​​​​​​​​​​​d. Alternatively, if the first message indicates that at least one of the multiple MVDs is not equal to 0 and all MVDs before the last MVD are zero, then no signaling is sent to the last MVD indicating that it is non-zero and the last MVD is inferred to be non-zero.

[0157] 5. It is recommended to use syntax elements that characterize the level components of an MVD to signal at least one of multiple MVDs.

[0158] 6. It is recommended to use syntax elements that characterize the vertical components of an MVD to signal at least one of multiple MVDs.

[0159] 7. It is recommended to use syntax elements that represent the absolute values ​​of the horizontal and / or vertical components of an MVD to signal at least one of multiple MVDs.

[0160] 8. It is recommended to derive the second MVD based on the first MVD, and to notify or derive the first MVD before signaling notification or deriving the second MVD.

[0161] 9. Similar to MMVD design, the MVD used in GMVD can be represented by two variables: direction and magnitude (or distance) represented by D. Additionally, the following may also apply:

[0162] a. Signal at least one of the multiple MVDs using syntax elements that characterize the direction of the MVD (e.g., horizontal or vertical).

[0163] i. In one example, the direction may include, but is not limited to:

[0164] 1) The form of MVD is (1,0)*D, where D is not less than 0;

[0165] 2) The form of MVD is (-1,0)*D, where D is not less than 0;

[0166] 3) The form of MVD is (0,1)*D, where D is not less than 0;

[0167] 4) The form of MVD is (0, -1) * D, where D is not less than 0;

[0168] 5) The form of MVD is (1,1)*D, where D is not less than 0;

[0169] 6) The form of MVD is (-1,1)*D, where D is not less than 0;

[0170] 7) The form of MVD is (1, -1) * D, where D is not less than 0;

[0171] 8) The form of MVD is (-1, -1) * D, where D is not less than 0.

[0172] ii. In one example, only the four directions mentioned in 9.ai1) to 9.ai4) are used.

[0173] iii. In one example, the direction can be binarized using a fixed-length codec, a unary codec, or an exponential Columbus codec.

[0174] b. In one example, D is greater than 0.

[0175] c. For GMVD codec blocks, the indication of MVD amplitude (or distance) can be represented by signaling notification or derived D.

[0176] i. In one example, D is restricted to the candidate set.

[0177] 1) In one example, the candidate set may include zero.

[0178] 2) In one example, the candidate set may include 1 / 4-samples, 1 / 2-samples, or other fractional samples.

[0179] 3) In one example, the candidate set may include 1-sample, 2-sample, 4-sample, 8-sample, 16-sample, 32-sample, 64-sample, 128-sample, or other 2-sample sets. N - Sample points.

[0180] 4) In one example, the candidate set may include 3-samples, 6-samples, 12-samples, 24-samples or other integer samples.

[0181] 5) In one example, the candidate set can be {1 / 4-sample, 1 / 2-sample, 1-sample, 2-sample, 4-sample, 8-sample, 16-sample, 32-sample}.

[0182] 6) In one example, the candidate set can be {1 / 4-sample, 1 / 2-sample, 1-sample, 2-sample, 3-sample, 4-sample, 6-sample, 8-sample, 16-sample}.

[0183] 7) In one example, the candidate set can be set to be equal to the candidate set used for the MMVD codec block in the same video processing unit (e.g., strip / slice / sub-picture / picture / sequence).

[0184] a. Optionally, the candidate set may include additional candidates besides those for MMVD codec blocks in the same video processing unit.

[0185] b. Optionally, at least one candidate in the candidate set may be different from the candidate used for the MMVD codec block in the same video processing unit.

[0186] c. Optionally, at least one candidate in the candidate set may be equal to one of the candidates for the MMVD coding / decoding block in the same video processing unit.

[0187] 8) In one example, the index of the selected MVD magnitude in the candidate set is signaled instead of directly signaling D.

[0188] ii. In one example, multiple candidate sets of MVD may be predefined, and one of them may be selected to encode / decode the current GMVD coding / decoding block.

[0189] 1) In one example, the selection depends on the message signaled at the slice / picture / sequence level, such as the slice header / picture header / PPS / SPS.

[0190] 2) In one example, the selection depends on the same message for the MMVD coding / decoding block, such as sps_fpel_mmvd_enabled_flag.

[0191] iii. In one example, D is binarized into fixed-length coding, or unary coding, or exponential Golomb coding.

[0192] iv. In one example, at least one binarized bin of D is context-coded in arithmetic coding.

[0193] v. Alternatively, at least one binarized bin of the message is bypass-coded in arithmetic coding.

[0194] vi. In one example, the first bin of D is context-coded in arithmetic coding.

[0195] 1) Optionally, further, the other bins of D are bypass-coded.

[0196] d. For example, the MVD derived from D may be further modified before being used to derive the final MVD for the partition.

[0197] i. In one example, D may be modified to D = D << S, where S is an integer, such as 2.

[0198] ii. It may be implicitly inferred whether to apply the modification.

[0199] 1) For example, if the width and / or height of the current picture is greater than a threshold, the modification may be applied. <​iii. Whether to apply the modification may depend on the signaling notification message, which may be at the sequence level (e.g., in the SPS and / or sequence header), picture level (e.g., in the PPS and / or picture header), stripe level (e.g., in the stripe header), sub-picture level, slice level, CTU line level, or CTU level.

[0201] 1) For example, the same messages in VVC, sps_fpel_mmvd_enabled_flag and / or pic_fpel_mmvd_enabled_flag, can be used to control MMVD and GMVD.

[0202] 2) Alternatively, separate signaling messages can be used to control MMVD and GMVD separately.

[0203] 10. In the example above, the block can be split into two segments using TPM or GEO, and two MVDs can be signaled or derived for these two segments.

[0204] a. Optionally, further, the sum of the MV derived for the segmentation via TPM or GEO and the MVD of the segmentation is calculated as the MV of the segmentation.

[0205] b. Optionally, further, the two motion compensations and the weighted summation of the two MVs of the two segments are performed in the same manner as TPM or GEO.

[0206] c. Optionally, further, the MV stored procedures for the two split MVs are performed in the same manner as TPM or GEO.

[0207] 11. In the above example, whether signaling notifications or inferences derive multiple recommended MVDs can be conditional.

[0208] a. In one example, whether signaling notifications or inferences are used to suggest multiple MVDs may depend on whether GEO is enabled for the current block.

[0209] b. In one example, if there is no signaling notification or derivation of multiple suggested MVDs, then the multiple suggested MVDs are inferred to be zero.

[0210] c. In one example, signaling notifications can be made at the sequence level (e.g., in the SPS and / or sequence header), at the picture level (e.g., in the PPS and / or picture header), at the stripe level (e.g., in the stripe header), at the sub-picture level, at the slice level, at the CTU line level, or at the CTU level to indicate whether signaling notifications are made or to derive the recommended GMVD.

[0211] i. Optionally, further, it may be conditionally signaled whether to enable GMVD, for example, under the condition of enabling TPM / GEO.

[0212] d. In one example, whether to signal or derive the proposed multiple MVDs may depend on the block width (W) and / or the block height (H). For example, if at least one or any combination of the following conditions is met, the proposed multiple MVDs may not be signaled or derived.

[0213] i. W >= T1, for example, T1 = 64;

[0214] ii. H >= T2, for example, T2 = 64;

[0215] iii. W >= T3 * H, for example, T3 = 4;

[0216] iv. H >= T4 * H, for example, T4 = 8;

[0217] v. W <= T5, for example, T5 = 8;

[0218] vi. H <= T6, for example, T6 = 8;

[0219] vii. W * H >= T7, for example, T7 = 2048;

[0220] viii. W * H <= T8, for example, T8 = 64;

[0221] ix. W == T9 or H == T10 (for example, T9 = T10 = 4);

[0222] x. W / H > T11 or max(W, H) / min(W, H) >= T11;

[0223] xi. W / H < T11 or max(W, H) / min(W, H) < T11.

[0224] e. Whether to signal or derive the proposed multiple MVDs may depend on the Merge candidate index.

[0225] i. In one example, the Merge candidate index is greater than a value.

[0226] 12. For the GMVD coding / decoding block, the MVD may be signaled before the Merge candidate index. <​​​​b. Optionally, further, how the signaling informs the Merge candidate index of the GMVD codec block (e.g., binarization processing) may depend on the use of GMVD.

[0229] i. In one example, if multiple suggested MVDs are signaled or inferred, the maximum Merge candidate index that can be signaled can be reduced.

[0230] 5 Examples

[0231] 5.1 Draft Amendment Example

[0232] ■Merge Data Syntax

[0233]

[0234]

[0235] ■Merge Data Semantics

[0236]

[0237] `tmvd_flag[x0][y0]` indicates whether triangulation prediction with motion vector difference is applied to the current codec unit. Array indices `x0` and `y0` indicate the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image. When `tmvd_flag[x0][y0]` does not exist, it is inferred that `tmvd_flag[x0][y0]` is equal to 0.

[0238] `tmvd_part_flag[x0][y0][partIdx]` indicates whether triangulation prediction with motion vector difference is applied to the segment in the current codec unit whose index is equal to `partIdx`. The array indices `x0` and `y0` indicate the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image. When `tmvd_part_flag[x0][y0][partIdx]` does not exist, if `tmvd_flag[x0][y0]` equals 1 and `partIdx` equals 1, then `tmvd_part_flag[x0][y0][partIdx]` is inferred to be equal to 1. Otherwise, it is inferred to be equal to 0.

[0239] `tmvd_distance_idx[x0][y0][partIdx]` indicates the index used to derive `TmvdDistance[x0][y0][partIdx]`. The array indices `x0` and `y0` indicate the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0240] TmvdDistanceArray[] is set to equal to {4,8,16,32,48,64,96,128,256} and TmvdDistance[x0][y0][partIdx] is set to equal to TmvdDistanceArray[tmvd_distance_idx[x0][y0][partIdx]].

[0241] mmvd_direction_idx[x0][y0] indicates the index used to derive TmvdBaseMv[x0][y0][partIdx][compIdx] with compIdx = 0..1. The array indices x0 and y0 indicate the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0242] TmvdBaseArray[][] is set to equal to {{1,0},{-1,0},{0,1},{0,-1},{1,1},{1,-1},{-1,1},{-1,-1}.

[0243] For compIdx = 0..1, TmvdBase[x0][y0][partIdx][compIdx] is set to equal TmvdBaseArray[mmvd_direction_idx[x0][y0]][compIdx].

[0244] When tmvd_part_flag[x0][y0][partIdx] equals 0, for compIdx = 0..1, TmvdOffset[x0][y0][partIdx][compIdx] is set to zero. Otherwise, for compIdx = 0..1, TmvdOffset[x0][y0][partIdx][compIdx] is set to equal to TmvdBaseMv[x0][y0][partIdx][compIdx] * TmvdDistance[x0][y0][partIdx] * (pic_fpel_mmvd_enabled_flag == 1? 4:1).

[0245] ■Derivation and processing of brightness motion vectors in Merge triangle mode

[0246] This process is invoked only when MergeTriangleFlag[xCb][yCb] equals 1, where (xCb, yCb) indicates the position of the top-left sample of the current luminance codec block relative to the top-left luminance sample of the current image.

[0247] The input for this process is:

[0248] – The brightness position (xCb, yCb) of the top-left sample of the current luminance block relative to the top-left luminance sample of the current image.

[0249] – The variable cbWidth indicates the width of the current codec block in the luma sample.

[0250] – The variable cbHeight indicates the height of the current codec block in the luminance sample.

[0251] The output of this process is:

[0252] – Brightness motion vectors mvA and mvB with 1 / 16 fractional sample point accuracy,

[0253] –Refer to indices refIdxA and refIdxB,

[0254] – Prediction list flags predListFlagA and predListFlagB.

[0255] Motion vectors mvA and mvB, reference indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB are derived through the following ordered steps:

[0256] 1. Call the derivation process of the brightness motion vector of the Merge mode specified in Section 8.5.2.2, take the brightness position (xCb, yCb), variables cbWidth and cbHeight as input, and output: brightness motion vector mvL0[0][0], mvL1[0][0], reference index refIdxL0, refIdxL1, prediction list use flags predFlagL0[0][0] and predFlagL1[0][0], bidirectional prediction weight index bcwIdx and Merge candidate list mergeCandList.

[0257] 2. The variables m and n, which are the merge indices of 0 and 1 respectively for the triangular segmentation, are derived using merge_triangle_idx0[xCb][yCb] and merge_triangle_idx1[xCb][yCb] by the following formula:

[0258] m=merge_triangle_idx0[xCb][yCb] (657)

[0259] n=merge_triangle_idx1[xCb][yCb]+(merge_triangle_idx1[xCb][yCb]>=m)? 1:0 (658)

[0260] 3. Let refIdxL0M and refIdxL1M, predFlagL0M and predFlagL1M, mvL0M and mvL1M be the reference index, prediction list flag and motion vector of the Merge candidate M (M = mergeCandList[m]) at position m in the Merge candidate list mergeCandList.

[0261] 4. Set variable X to equal (m&0x01).

[0262] 5. When predFlagLXM equals 0, X is set to equal to (1-X).

[0263] 6. Apply the following formula:

[0264] mvA[0]=mvLXM[0]+TmvdOffset[x0][y0][0][0] (659)

[0265] mvA[1]=mvLXM[1]+TmvdOffset[x0][y0][0][1] (660)

[0266] refIdxA=refIdxLXM (661)

[0267] predListFlagA = X (662)

[0268] 7. Let refIdxL0N and refIdxL1N, predFlagL0N and predFlagL1N, mvL0N and mvL1N be the reference index, prediction list flag and motion vector of the Merge candidate N (N = mergeCandList[n]) at position m in the Merge candidate list mergeCandList.

[0269] 8. Set variable X to equal (n&0x01).

[0270] 9. When predFlagLXN equals 0, X is set to equal to (1-X).

[0271] 10. Apply the following formula:

[0272] mvB[0]=mvLXN[0]+TmvdOffset[x0][y0][1][0] (663)

[0273] mvB[1]=mvLXN[1]+TmvdOffset[x0][y0][1][1] (664)

[0274] refIdxB=refIdxLXN (665)

[0275] predListFlagB = X (666)

[0276] Table 123 Syntax Elements and Related Binarization

[0277]

[0278] Table 128 assigns ctxInc to syntax elements with context encoding / decoding bin.

[0279]

[0280]

[0281] Figure 13 A block diagram of an example video processing system 1900 is shown, which can implement various techniques disclosed herein. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format, such as 8-bit or 10-bit multi-component pixel values, or it may be received in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0282] System 1900 may include codec component 1904, which implements the various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via connected communication, as shown in component 1906. The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or to send displayable video to display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder, and the corresponding decoding tools or operations for the reverse encoded results will be performed by the decoder.

[0283] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The techniques described herein can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0284] Figure 14 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be implemented in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing circuitry 3606. Processor 3602 can be configured to implement one or more methods described herein. Memory(s) 3604 can be used to store data and code for implementing the methods and techniques described herein. Video processing circuitry 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0285] Figure 16 This is a block diagram illustrating an example of a video encoding / decoding system 100 to which the technology of this disclosure can be applied.

[0286] like Figure 16 As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and can be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and can be referred to as a video decoding device.

[0287] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0288] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations of these sources. Video data may include one or more images. Video encoder 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data may be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.

[0289] The target device 120 may include an I / O interface 126, a video decoder 124, and a display device 122.

[0290] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with target device 120, or it may be external to target device 120, which is configured to connect to an external display device.

[0291] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Codec (HEVC) standard, the Universal Video Codec (VVM) standard, and other current and / or further standards.

[0292] Figure 17 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 16 The video encoder 114 in the system 100 described herein.

[0293] The video encoder 200 can be configured to perform any or all of the techniques disclosed herein. Figure 18 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0294] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202 (which may include a mode selection unit 203), a motion estimation unit 204, a motion compensation unit 205, an intra-frame prediction unit 206, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214.

[0295] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0296] Furthermore, some components, such as the motion estimation unit 204 and the motion compensation unit 205, can be highly integrated, but for illustrative purposes, in Figure 17 The examples are shown separately.

[0297] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support video blocks of various sizes.

[0298] The mode selection unit 203 can select one of the coding modes (intra-frame or inter-frame) based, for example, on the error result, and provide the resulting intra-frame or inter-frame coded block to the residual generation unit 207 to generate residual block data and to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select a combination of intra-frame and inter-frame prediction (CIIP) modes, where prediction is based on inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select the resolution of the motion vector for the block (e.g., sub-pixel or integer pixel precision).

[0299] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing one or more reference frames from buffer 213 with the current video block. Motion compensation unit 205 can determine the predicted video block for the current video block based on motion information and decoded samples from images other than those associated with the current video block from buffer 213.

[0300] The motion estimation unit 204 and the motion compensation unit 205 can perform different operations for the current video block, for example, depending on whether the current video block is in an I-band, P-band, or B-band.

[0301] In some examples, motion estimation unit 204 can perform unidirectional prediction for the current video block, and can search reference images in list 0 or list 1 to find a reference video block for the current video block. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0302] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0 and can also search for another reference video block for the current video block in the reference images of list 1. Motion estimation unit 204 can then generate a reference index indicating the reference image containing the reference video block in list 0 or list 1, and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as the motion information of the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video block indicated by the motion information of the current video block.

[0303] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder to use in the decoding process.

[0304] In some examples, motion estimation unit 204 may not output the complete set of motion information for the current video. Instead, motion estimation unit 204 may signal the motion information of the current video block to another video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0305] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0306] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) within the syntax structure associated with the current video block. This motion vector difference represents the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0307] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0308] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block may include the predicted video block and various syntax elements.

[0309] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a negative sign) the predicted video block of the current video block from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0310] In other examples, for the current video block, such as in skip mode, there may not be residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0311] The transform processing unit 208 can generate one or more transform coefficient video blocks of the current video block by applying one or more transforms to the residual video block associated with the current video block.

[0312] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video.

[0313] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding sample from one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block, which is stored in buffer 213.

[0314] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0315] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0316] Figure 18 This is a block diagram illustrating an example of a video decoder 300. The video decoder 300 can be... Figure 16 The video decoder 114 in the system 100 described herein.

[0317] The video decoder 300 can be configured to perform any or all of the techniques disclosed herein. Figure 17 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0318] exist Figure 18 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform functions typically associated with video encoder 200. Figure 17 The decoding path is the opposite of the encoding path described.

[0319] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and based on the entropy-decoded video data, the motion compensation unit 302 can determine motion information including motion vectors, motion vector precision, reference image list index, and other motion information. For example, the motion compensation unit 302 can determine this information by executing AMVP and Merge modes.

[0320] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The identifier of the interpolation filter used with sub-pixel precision can be included in the syntax elements.

[0321] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during the encoding of video blocks to calculate the interpolated values ​​of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0322] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode the frames and / or stripes of the encoded video sequence, segmentation information describing how each macroblock of the picture of the encoded video sequence is divided, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame coded block, and other information for decoding the encoded video sequence.

[0323] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 303 performs inverse quantization, i.e., dequantization, on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 303 applies an inverse transform.

[0324] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove blocky artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on the display device.

[0325] The following is a list of preferred solutions for some embodiments.

[0326] The following scheme shows example embodiments of the techniques discussed in the preceding sections (e.g., item 1).

[0327] 1. A video processing method (e.g., Figure 15 The method 1500 shown includes: for the conversion between video units of a video and the codec representation of the video, segmenting the video into multiple segments, and performing the conversion on at least some of the multiple segments using a Merge pattern with motion vector differences and based on the multiple segments.

[0328] 2. The method described in Scheme 1, wherein at least one of the plurality of segments is a non-rectangular segment.

[0329] 3. The method as described in any one of Schemes 1-2, wherein the segmentation uses a triangular segmentation pattern, wherein the Merge pattern with motion vector difference uses the triangular segmentation pattern Merge candidate as the Merge candidate.

[0330] The following scheme shows example embodiments of the techniques discussed in the preceding sections (e.g., item 2).

[0331] 4. The method as described in any one of Schemes 1-3, wherein the plurality of segments comprises N segments, where N is an integer greater than or equal to 2, wherein the encoding / decoding representation comprises at most N motion vector differences, or the conversion uses at most N motion vector differences.

[0332] 5. The method as described in Scheme 4, wherein the encoding / decoding representation includes a field indicating K segments, where K < N, and one motion vector difference is shared by more than one of the multiple segments.

[0333] The following scheme shows example embodiments of the techniques discussed in the preceding sections (e.g., item 3).

[0334] 6. The method as described in schemes 4-5, wherein the field in the encoding / decoding representation indicates whether at least one of the N motion vector differences is a non-zero value.

[0335] 7. The method described in Scheme 6, wherein the field is encoded and decoded in the encoding / decoding representation via context encoding / decoding.

[0336] The following schemes show example embodiments of the techniques discussed in the preceding sections (e.g., items 5, 6, and 7).

[0337] 8. The method of any one of schemes 4-7, wherein at least one of the N motion vector differences is signaled as a syntax element indicating the horizontal component of the motion vector difference.

[0338] 9. The method of any one of schemes 4-7, wherein at least one of the N motion vector differences is signaled as a syntax element indicating the vertical component of the motion vector difference.

[0339] 10. The method of any one of schemes 4-7, wherein at least one of the N motion vector differences is signaled as a syntax element indicating the absolute value of the horizontal or vertical component of the motion vector difference.

[0340] The following scheme shows example embodiments of the techniques discussed in the preceding sections (e.g., item 8).

[0341] 11. The method of any one of schemes 4-7, wherein at least one of the N motion vector differences is derived from another motion vector difference.

[0342] The following scheme shows example embodiments of the techniques discussed in the preceding sections (e.g., item 9).

[0343] 12. A video processing method comprising: for a conversion between video units of a video and a codec representation of the video, segmenting the video into multiple segments, and performing the conversion based on the multiple segments, wherein the codec representation includes a field indicating motion vector differences, the field indicating motion vector differences using one or more of two variables corresponding to the direction and magnitude of the motion vector differences.

[0344] 13. The method described in Scheme 12, wherein one of the two variables is a direction variable.

[0345] 14. The method described in schemes 12-13, wherein the amplitude is allowed to have values ​​belonging to a set.

[0346] The following scheme shows example embodiments of the techniques discussed in the preceding sections (e.g., item 12).

[0347] 15. A video processing method, comprising: for a conversion between video units of a video and a codec representation of the video, dividing the video into multiple segments, and performing the conversion based on the multiple segments, wherein the codec representation contains fields arranged in sequence.

[0348] 16. The method of Scheme 15, wherein the order includes motion vector differences that occur before the Merge candidate index in the codec representation.

[0349] 17. The method of embodiment 15, wherein the order depends on the characteristics of the video unit.

[0350] 18. The method described in any of the above schemes, wherein the video unit includes an encoding / decoding unit.

[0351] 19. The method described in any of the above schemes, wherein the video unit comprises a video block.

[0352] 20. The method of any one of schemes 1 to 19, wherein the conversion includes encoding the video into the codec representation.

[0353] 21. The method of any one of claims 1 to 19, wherein the conversion includes decoding the codec representation to generate pixel values ​​of the video.

[0354] 22. A video decoding apparatus, including a processor configured to implement one or more of the methods described in schemes 1 to 21.

[0355] 23. A video encoding apparatus, including a processor configured to implement one or more of the methods described in schemes 1 to 21.

[0356] 24. A computer program product having computer code stored thereon, which, when executed by a processor, causes the processor to implement the method described in any one of schemes 1 to 21.

[0357] 25. A method, apparatus or system described herein.

[0358] Figure 19A flowchart illustrating an example of a video processing method is shown. The method includes: determining, for a conversion between a current video block and the bitstream of the current video, that the current video block is encoded and decoded using a geometric segmentation mode (1902); deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block (1904); and performing the conversion based on the refined MV (1906).

[0359] In some examples, the current video block is a geometric motion vector difference (GMVD) codec block or a Merge (MMVD) codec block with motion vector difference.

[0360] In some examples, the geometric segmentation pattern includes multiple segmentation schemes, at least one of which divides the current video block into two or more segments, at least one of which is neither square nor rectangular.

[0361] In some examples, the geometric segmentation pattern includes a triangular segmentation pattern.

[0362] In some examples, the geometric segmentation pattern includes a geometric merge pattern.

[0363] In some examples, the current video block is divided into N segments, and for each segment of the current video block, at most N MVDs are signaled or derived, where N>=2.

[0364] In some examples, N = 2.

[0365] In some examples, for the segmentation of the current video block, signaling notifies or derives N MVDs, and each MVD corresponds to a specific segmentation.

[0366] In some examples, a segment's corresponding MVD is added to the segment's MV to obtain a refined MV for that segment, which is then used for motion compensation of that segment.

[0367] In some examples, signaling informs or derives K MVDs, and at least one MVD corresponds to more than one partition, where K <N。

[0368] In some examples, the number of MVDs in the signaling notification depends on the decoding information, which includes at least one of the following: block size associated with N segments, low-latency check flag, prediction direction, or a list of reference images.

[0369] In some examples, the number of MVDs is included in the bitstream.

[0370] In some examples, a single split uses multiple MVDs.

[0371] In some examples, multiple MVDs are added to the segmented MV to obtain a refined MV for that segment.

[0372] In some examples, the first message is included in the bitstream to indicate whether at least one of the multiple MVDs is not equal to 0.

[0373] In some examples, the first message is a flag.

[0374] In some examples, the first message is encoded in arithmetic codecs via context codecs or bypass codecs.

[0375] In some examples, the first message is encoded and decoded using at least one context identical to the syntax element to indicate whether MMVD is applied.

[0376] In some examples, information related to a plurality of MVDs is included in the bitstream only if the first message indicates that at least one of the plurality of MVDs is not equal to 0.

[0377] In some examples, a second message for at least one of the plurality of MVDs is included in the bitstream to indicate whether the MVD is equal to zero, wherein the second message is a non-zero message.

[0378] In some examples, at least one binarized bin of the second message is encoded and decoded via context in arithmetic encoding and decoding.

[0379] In some examples, at least one binarized bin of the second message is bypassed during arithmetic encoding / decoding.

[0380] In some examples, the second message is encoded and decoded using at least one context identical to the syntax element to indicate whether MMVD is applied.

[0381] In some examples, the information associated with the MVD is included in the bitstream only if the second message indicates that the MVD is not equal to 0.

[0382] In some examples, at least one of the plurality of MVDs is contained in a bitstream having a syntax element that indicates the horizontal component of the MVD, or at least one of the plurality of MVDs is contained in a bitstream having a syntax element that indicates the vertical component of the MVD.

[0383] In some examples, at least one of the multiple MVDs is included in a bitstream having syntax elements that indicate the absolute values ​​of the horizontal and / or vertical components of the MVD.

[0384] In some examples, a second MVD is derived from a first MVD, and the first MVD is notified or derived before the signaling notification or derivation of the second MVD.

[0385] In some examples, the MVD used in GMVD can be represented by two variables: the first variable is the direction, and the second variable is the magnitude represented by D.

[0386] In some examples, at least one of the multiple MVDs is contained in a bitstream having syntax elements that indicate the direction of the MVD.

[0387] In some examples, the direction may include at least one of the following:

[0388] 1) The form of MVD is (1,0)*D, where D is not less than 0;

[0389] 2) The form of MVD is (-1,0)*D, where D is not less than 0;

[0390] 3) The form of MVD is (0,1)*D, where D is not less than 0;

[0391] 4) The form of MVD is (0, -1) * D, where D is not less than 0;

[0392] 5) The form of MVD is (1,1)*D, where D is not less than 0;

[0393] 6) The form of MVD is (-1,1)*D, where D is not less than 0;

[0394] 7) The form of MVD is (1, -1)*D, where D is not less than 0; and

[0395] 8) The form of MVD is (-1, -1) * D, where D is not less than 0.

[0396] In some examples, the direction may include at least one of the following:

[0397] 1) The form of MVD is (1,0)*D, where D is not less than 0;

[0398] 2) The form of MVD is (-1,0)*D, where D is not less than 0;

[0399] 3) The form of MVD is (0,1)*D, where D is not less than 0; and

[0400] 4) The form of MVD is (0, -1) * D, where D is not less than 0.

[0401] In some examples, direction can be binarized using fixed-length encoding / decoding, unary encoding / decoding, or exponential Columbus encoding / decoding.

[0402] In some examples, D is greater than 0.

[0403] In some examples, for GMVD codec blocks, the indication of MVD magnitude can be represented by signaling notification or derived D.

[0404] In some examples, D is restricted to the candidate set.

[0405] In some examples, the candidate set may include zero.

[0406] In some examples, the candidate set may include 1 / 4-samples, 1 / 2-samples, or other fractional samples.

[0407] In some examples, the candidate set may include 1-sample, 2-sample, 4-sample, 8-sample, 16-sample, 32-sample, 64-sample, 128-sample, or other 2N samples.

[0408] In some examples, the candidate set may include 3-samples, 6-samples, 12-samples, 24-samples, or other integer samples.

[0409] In some examples, the candidate set can be {1 / 4-sample, 1 / 2-sample, 1-sample, 2-sample, 4-sample, 8-sample, 16-sample, 32-sample}.

[0410] In some examples, the candidate set can be {1 / 4-sample, 1 / 2-sample, 1-sample, 2-sample, 3-sample, 4-sample, 6-sample, 8-sample, 16-sample}.

[0411] In some examples, the candidate set can be set to be equal to the candidate set used for MMVD codec blocks in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

[0412] In some examples, the candidate set may include more candidates than those for MMVD codec blocks in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

[0413] In some examples, at least one candidate in the candidate set is different from the candidate for an MMVD codec block in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

[0414] In some examples, at least one candidate in the candidate set is equal to one of the candidates for the MMVD codec block in the same video processing unit, where the video processing unit includes at least one of a stripe, a slice, a sub-picture, a picture, or a sequence.

[0415] In some examples, the index of the selected MVD magnitude in the candidate set is signaled instead of directly signaling D.

[0416] In some examples, multiple candidate MVD sets are predefined, and one of them is selected to encode / decode the current GMVD codec block.

[0417] In some examples, the selection depends on a message signaled at least at one of the stripe level, the picture level, or the sequence level.

[0418] In some examples, the message is signaled in at least one of a stripe header, a picture header, a picture parameter set (PPS), or a sequence parameter set (SPS).

[0419] In some examples, the selection depends on the same message for the MMVD codec block, and the same message is sps_fpel_mmvd_enabled_flag.

[0420] In some examples, D is binary coded into a fixed-length code, or a unary code, or a Golomb-Rice code.

[0421] In some examples, at least one binary bin of D is context-coded in arithmetic coding.

[0422] In some examples, at least one binary bin of D is bypass-coded in arithmetic coding.

[0423] In some examples, the first bin of D is context-coded in arithmetic coding.

[0424] In some examples, the other bins of D are bypass-coded in arithmetic coding.

[0425] In some examples, the MVD derived from D can be further modified before being used to derive the final MVD for the partition.

[0426] In some examples, D can be modified to D = D << S, where S is an integer.

[0427] In some examples, S = 2.

[0428] In some examples, it is implicitly inferred whether to apply the modification.

[0429] In some examples, the modification is applied if the width and / or height of the current image is greater than a threshold.

[0430] In some examples, whether the modification is applied depends on the signaling notification message, which is at least one of the sequence level, picture level, stripe level, sub-picture level, slice level, CTU line level, and CTU level.

[0431] In some examples, the message is signaled in the SPS and / or sequence header, the PPS and / or picture header, or the stripe header.

[0432] In some examples, the same message, including sps_fpel_mmvd_enabled_flag and / or pic_fpel_mmvd_enabled_flag, can be used to control MMVD and GMVD.

[0433] In some examples, separate signaling messages can be used to control MMVD and GMVD separately.

[0434] In some examples, the current video block can be divided into two segments using a triangular segmentation mode or a geometric merge mode, and two MVDs can be signaled or derived for these two segments.

[0435] In some examples, the sum of the MV derived for the segmentation through the triangular segmentation pattern or the geometric merge pattern and the MVD of the segmentation is calculated as the MV of the segmentation.

[0436] In some examples, the two motion compensations of the two segments and the weighted summation of the two MVs are performed in the same way as in the triangular segmentation mode or the geometric merge mode.

[0437] In some examples, the MV stored procedures for the two MVs of the two partitions are executed in the same way as in the triangular partitioning mode or the geometric merge mode.

[0438] In some examples, whether signaling notifies or derives the multiple MVDs depends on one or more conditions.

[0439] In some examples, whether signaling notifications or derivations of the suggested multiple MVDs depend on whether the geometry merge mode is enabled for the current video block.

[0440] In some examples, if there is no signaling notification or derivation of the multiple MVDs, the multiple MVDs are inferred to be zero.

[0441] In some examples, signaling at least one of the following levels—sequence level, picture level, stripe level, sub-picture level, slice level, CTU line level, or CTU level—can indicate whether GMVD is enabled.

[0442] In some examples, it is signaled whether GMVD is enabled in the SPS and / or sequence header, PPS and / or picture header, or slice header.

[0443] In some examples, it is signaled whether GMVD is enabled under the condition that the triangular partitioning mode or geometric Merge mode is enabled.

[0444] In some examples, whether the proposed multiple MVDs are signaled or derived may depend on the block width (W) and / or block height (H) of the current video block.

[0445] In some examples, if at least one or any combination of the following conditions is satisfied, the multiple MVDs are not signaled or derived:

[0446] i. W >= T1, where T1 = 64;

[0447] ii. H >= T2, where T2 = 64;

[0448] iii. W >= T3 * H, where T3 = 4;

[0449] iv. H >= T4 * H, where T4 = 8;

[0450] v. W <= T5, where T5 = 8;

[0451] vi. H <= T6, where T6 = 8; <00​​​​​​​​​​​​​​​​​​​​​​​​

[0460] In some examples, the index of the Merge candidate is signaled based on whether the multiple MVDs are signaled or inferred.

[0461] In some examples, how the signaling informs the GMVD codec block of the index of the Merge candidate depends on the use of GMVD.

[0462] In some examples, if the multiple MVDs are signaled or inferred, the maximum index of the Merge candidate that can be signaled can be reduced.

[0463] In some examples, the conversion involves encoding the current video block into the bitstream.

[0464] In some examples, the conversion involves decoding the current video block from the bitstream.

[0465] In some examples, the conversion includes generating the bitstream from the current video block; the method further includes storing the bitstream in a non-transitory computer-readable recording medium.

[0466] In some examples, the means for processing video data includes a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to: determine that the current video block is encoded and decoded using a geometric segmentation mode for a conversion between a current video block and the bitstream of the current video; derive at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) signaled or derived for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block; and perform the conversion based on the refined MV.

[0467] In some examples, a non-transitory computer-readable medium stores instructions that cause a processor to: determine that the current video block is encoded and decoded using a geometric segmentation mode for a conversion between the current video block and the bitstream of the current video; derive at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) signaled or derived for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block; and perform the conversion based on the refined MV.

[0468] In some examples, a non-transitory computer-readable medium stores a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining that the current video block is encoded and decoded using a geometric segmentation mode for a conversion between a current video block and the bitstream of the current video; deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block; and generating a bitstream from the current video block based on the refined MV.

[0469] Figure 20 A flowchart illustrating an example of a method for storing a bitstream of video is shown. The method includes: determining, for a conversion between a current video block and the bitstream of the current video, that the current video block is encoded and decoded using a geometric segmentation mode (2002); deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to motion vectors (MVs) derived from Merge candidates associated with the current video block (2004); generating the bitstream from the current video block based on the refined MV (2006); and storing the bitstream in a non-transitory computer-readable recording medium (2008).

[0470] Figure 21 A flowchart illustrating an example of a video processing method is shown. The method includes: determining, for a conversion between a current video block and the bitstream of the video, that the current video block is encoded and decoded using a geometric segmentation mode (2102); deriving at least one refined MV of the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block, the MV being related to an offset distance and / or offset direction (21904); and performing the conversion based on the refined MV (2206).

[0471] In some examples, the current video block is a geometric motion vector differential codec block or a Merge codec block with motion vector differences.

[0472] In some examples, the geometric segmentation pattern includes multiple segmentation schemes, in which the current video block comprises two or more segments.

[0473] In some examples, at least one of the two or more divisions is neither square nor rectangular.

[0474] In some examples, the geometric segmentation pattern includes a triangular segmentation pattern.

[0475] In some examples, the geometric segmentation pattern includes a geometric merge pattern.

[0476] In some examples, the same lookup table is used to derive the offset distance from the distance index used for the geometric motion vector difference and the Merge with the motion vector difference.

[0477] In some examples, the same lookup table is used to derive the offset direction from the direction index used for the geometric motion vector difference and the Merge with the motion vector difference.

[0478] Figure 22 A flowchart illustrating an example of a video processing method is shown. The method includes: for a conversion between a current video block and the bitstream of the video, determining that the current video block is encoded and decoded using a geometric segmentation mode (2202); determining whether a geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD from a plurality of motion vector differences (MVDs) signaled or derived for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block (2204); and performing the conversion based on the refined MV (1906).

[0479] In some examples, the geometric segmentation pattern includes multiple segmentation schemes, in which the current video block comprises two or more segments.

[0480] In some examples, at least one of the two or more divisions is neither square nor rectangular.

[0481] In some examples, whether the geometric motion vector differential encoding / decoding method is enabled or disabled for different segments within the current video block is determined explicitly or implicitly.

[0482] In some examples, conditional signaling informs the first syntax element to indicate whether the geometric motion vector differential encoding / decoding method is applied to the current video block or whether at least one of the plurality of MVDs of the segmentation of the current video block is a non-zero MVD.

[0483] In some examples, when the first syntax element indicates the application of the geometric motion vector differential encoding / decoding method, how the signaling notification of information associated with the MVD and / or the on / off control information of the geometric motion vector differential encoding / decoding method depends on the segmentation index or on the segmentation position relative to the current video block.

[0484] In some examples, for the first segment in the encoding / decoding sequence, the signaling informs the on / off control information of the geometric motion vector differential encoding / decoding method or whether the MVD is zero for both the horizontal and vertical components.

[0485] In some examples, if the geometric motion vector differential codec is turned off for the first segment, then for the second segment in the codec sequence, no signaling is sent to the geometric motion vector differential codec to provide on / off control information or to indicate whether the MVD is zero for both the horizontal and vertical components.

[0486] In some examples, it is inferred that the geometric motion vector differential encoding / decoding method is applied to the second partition.

[0487] In some examples, signaling notifies indicators that enable / disable the geometric motion vector differential encoding / decoding method at high levels including sequence level and / or low levels including picture level and stripe level.

[0488] In some examples, the signaling notifies an indicator that enables / disables the geometric motion vector differential encoding / decoding method in at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture header, or strip header.

[0489] In some examples, the indicator is conditionally signaled based on whether the geometry merge mode is enabled, and / or whether the stripe type or picture type is type B, and / or whether the sequence allows at least one B stripe / picture.

[0490] In some examples, the indicator for enabling / disabling the geometric motion vector differential encoding / decoding method is implicitly derived without signaling notification.

[0491] In some examples, the indicator for enabling / disabling the geometric motion vector differential encoding / decoding method is derived based on whether the Merge encoding / decoding method with motion vector difference and / or the geometric Merge mode is enabled.

[0492] In some examples, the geometric motion vector differential encoding / decoding method is enabled if the Merge encoding / decoding method with motion vector difference and / or the geometric Merge mode are enabled.

[0493] In some examples, if the Merge encoding / decoding method with motion vector difference is disabled, the geometric motion vector difference encoding / decoding method is also disabled.

[0494] In some examples, disabling the geometry merge mode disables the geometry motion vector differential encoding / decoding method.

[0495] In some examples, the conversion involves encoding the current video block into the bitstream.

[0496] In some examples, the conversion involves decoding the current video block from the bitstream.

[0497] In some examples, the conversion includes generating the bitstream from the current video block; the method also includes storing the bitstream in a non-transitory computer-readable recording medium.

[0498] In some examples, a non-transitory computer-readable medium is provided that stores instructions that cause a processor to implement the methods described above.

[0499] In some examples, a non-transitory computer-readable medium is provided that stores a bitstream of video generated by the methods described above.

[0500] Figure 23 A flowchart illustrating an example of a method for storing a bitstream of video is shown. The method includes: determining, for a conversion between a current video block and the bitstream of the video, that the current video block is encoded and decoded using a geometric segmentation mode (2302); determining whether a geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD from a plurality of motion vector differences (MVDs) signaled or derived for the current video block to motion vectors (MVs) derived from Merge candidates associated with the current video block (2304); generating the bitstream from the current video block based on the refined MV (2306); and storing the bitstream in a non-transitory computer-readable recording medium (2308).

[0501] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, video compression algorithms can be applied during the conversion from the pixel representation of a video to its corresponding bitstream representation, and vice versa. As defined by the syntax elements, for example, the bitstream representation of the current video block can correspond to bits juxtaposed within the bitstream or distributed at different positions within the bitstream. For example, macroblocks can be encoded based on the transform and encoding / decoding error residuals and also using headers and other fields in the bitstream.

[0502] The solutions, examples, embodiments, modules, and functional operations disclosed and otherwise described herein can be implemented in digital electronic circuits, or computer software, firmware, or hardware, including the structures disclosed herein and their equivalents, or combinations thereof. The disclosed embodiments and other embodiments can be implemented as one or more computer program products, i.e., one or more modules of computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a material composition that influences machine-readable propagation signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, programmable processors, computers, or multiprocessors or computer groups. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. Propagation signals are man-made signals, such as machine-generated electrical, optical, or electromagnetic signals, generated to encode information for transmission to a suitable receiving device.

[0503] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored as part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to that program, or as multiple coordinating files (e.g., files that store one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.

[0504] The processing and logic flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0505] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or receive data from or transfer data to one or more mass storage devices via operative coupling, or both. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The processor and memory may be supplemented by or integrated into special-purpose logic circuitry.

[0506] While this patent document contains numerous details, it should not be construed as limiting any implementation or scope of the claims, but rather as a description of features specific to particular embodiments of a particular technology. Some features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in some combinations, or even initially claimed to be so, in some cases, one or more features from a claimed combination may be removed from the combination, and a claimed combination may refer to a sub-combination or a variation of a sub-combination.

[0507] Similarly, although operations are described in a specific order in the accompanying drawings, this should not be construed as requiring that these operations be performed in the specific order shown or sequentially, or that all of the shown operations be performed to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0508] Only some implementation methods and examples have been described. Other implementation methods, enhancements and variations can be made based on the content described and illustrated in this patent document.

Claims

1. A video processing method, comprising: For the conversion between the current video block and the bitstream of the video, it is determined that the current video block is encoded and decoded using a geometric segmentation mode; Determine whether the geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD of a plurality of motion vector differences (MVDs) that are signaled or derived for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block. as well as The transformation is performed based on the refined MV.

2. The method according to claim 1, wherein, The geometric segmentation mode includes multiple segmentation schemes, and in at least one segmentation scheme, the current video block includes two or more segments.

3. The method according to claim 2, wherein, At least one of the two or more divisions is neither square nor rectangular.

4. The method according to claim 3, wherein, Whether the geometric motion vector differential encoding / decoding method is enabled or disabled for different segments in the current video block is determined explicitly or implicitly.

5. The method according to any one of claims 1 to 4, wherein, A conditional signaling notification is sent to the first syntax element to indicate whether the geometric motion vector differential encoding / decoding method is applied to the current video block or whether at least one of the plurality of motion vector differences (MVDs) of the segmentation of the current video block is a non-zero MVD.

6. The method according to claim 5, wherein, When the first syntax element indicates the application of the geometric motion vector differential encoding / decoding method, how signaling notifications are sent to information associated with the MVD and / or the on / off control information of the geometric motion vector differential encoding / decoding method depend on the segmentation index or on the segmentation position relative to the current video block.

7. The method according to claim 6, wherein, For the first segment in the encoding / decoding sequence, the signaling notifies the on / off control information of the geometric motion vector differential encoding / decoding method, or the signaling notifies whether the MVD for the horizontal and vertical components is zero.

8. The method according to claim 7, wherein, If the geometric motion vector differential encoding / decoding method is turned off for the first segment, then for the second segment in the encoding / decoding sequence, no signaling is sent to the geometric motion vector differential encoding / decoding method to provide on / off control information, or a signaling is sent to indicate whether the MVD for both the horizontal and vertical components is zero.

9. The method according to claim 8, wherein, It is inferred that the geometric motion vector differential encoding and decoding method is applied to the second segmentation.

10. The method according to any one of claims 1 to 4, wherein, In high-level signaling, including sequence level and / or low-level signaling, including image level and stripe level, the signaling notifies an indicator to enable / disable the geometric motion vector differential encoding / decoding method.

11. The method according to claim 10, wherein, In at least one of the Sequence Parameter Set (SPS), Picture Parameter Set (PPS), Picture header, or strip header, the signaling notifies an indicator that enables / disables the geometric motion vector differential encoding / decoding method.

12. The method according to claim 11, wherein, The indicator is conditionally signaled based on whether the geometry merge mode is enabled, and / or whether the stripe type or picture type is type B, and / or whether the sequence allows at least one B stripe / picture.

13. The method according to any one of claims 1 to 4, wherein, In the absence of signaling notification, the indicator for enabling / disabling the geometric motion vector differential encoding / decoding method is implicitly derived.

14. The method according to claim 13, wherein, The indicator for enabling / disabling the geometric motion vector differential encoding / decoding method is derived based on whether the Merge encoding / decoding method with motion vector difference and / or the geometric Merge mode is enabled.

15. The method according to claim 14, wherein, If the Merge encoding / decoding method with motion vector difference and / or the geometric Merge mode are enabled, then the geometric motion vector difference encoding / decoding method is enabled.

16. The method of claim 14, wherein, If the Merge encoding / decoding method with motion vector difference is disabled, then the geometric motion vector difference encoding / decoding method is disabled.

17. The method of claim 14, wherein, If the geometry merge mode is disabled, the geometry motion vector differential encoding / decoding method is also disabled.

18. The method according to any one of claims 1 to 4, wherein, The conversion includes encoding the current video block into the bitstream.

19. The method according to any one of claims 1 to 4, wherein, The conversion includes decoding the current video block from the bitstream.

20. The method according to any one of claims 1 to 4, wherein, The conversion includes generating the bitstream from the current video block; The method also includes: The bit stream is stored in a non-transitory computer-readable recording medium.

21. A video processing method, comprising: For the conversion between the current video block and the bitstream of the video, it is determined that the current video block is encoded and decoded using a geometric segmentation mode; At least one refined MV of the current video block is derived by adding at least one of a plurality of motion vector differences (MVDs) derived from a Merge candidate associated with the current video block to a motion vector (MV) derived from a signaling notification or derivation for the current video block. The MVs are related to offset distance and / or offset direction. The transformation is performed based on the refined MV.

22. The method according to claim 21, wherein, The current video block is a geometric motion vector differential codec block or a Merge codec block with motion vector differences.

23. The method according to claim 21, wherein, The geometric segmentation mode includes multiple segmentation schemes, and in at least one segmentation scheme, the current video block includes two or more segments.

24. The method according to claim 23, wherein, At least one of the two or more divisions is neither square nor rectangular.

25. The method according to any one of claims 21 to 24, wherein, The geometric segmentation pattern includes the triangular segmentation pattern.

26. The method according to any one of claims 21 to 24, wherein, The geometric segmentation mode includes the geometric merge mode.

27. The method according to any one of claims 22 to 24, wherein, The offset distance is derived from the distance index used for the geometric motion vector difference and the Merge with the motion vector difference using the same lookup table.

28. The method according to any one of claims 22 to 24, wherein, The offset direction is derived using the same lookup table from the direction index used for the geometric motion vector difference and the Merge with the motion vector difference.

29. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein, When the instruction is executed by the processor, the processor is caused to implement the method as described in any one of claims 1 to 28.

30. A non-transitory computer-readable medium storing instructions that cause a processor to perform the method as described in any one of claims 1 to 28.

31. A non-transitory computer-readable medium storing a bitstream of video generated by the method of any one of claims 1 to 28.

32. A non-transitory computer-readable recording medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein, The method includes: The current video block of the video is determined to be encoded and decoded using a geometric segmentation mode; Determine whether the geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD from a plurality of motion vector differences (MVDs) derived for signaling notification or derivation of the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block; and The bitstream of the video is generated based on the refined MV.

33. A method for storing a video bitstream, comprising: The current video block of the video is determined to be encoded and decoded using a geometric segmentation mode; Determine whether the geometric motion vector differential encoding / decoding method is enabled or disabled, wherein the geometric motion vector differential encoding / decoding method derives at least one refined MV of the current video block by adding at least one MVD of a plurality of motion vector differences (MVDs) that are signaled or derived for the current video block to a motion vector (MV) derived from a Merge candidate associated with the current video block. Based on the refined MV, the bitstream is generated from the current video block; and The bit stream is stored in a non-transitory computer-readable recording medium.

Citation Information

Patent Citations

  • Update of lookup table: FIFO, constrained FIFO

    CN110662039A

  • Signaling of motion vector precision indication with adaptive motion vector resolution

    CN110944191A