Motion vector difference for blocks with geometric partitioning

By adding multiple motion vector differences (MVDs) to the merge candidate exported motion vectors (MVs) for signaling notification or export of video blocks to achieve refinement, the problem of inaccurate direct inheritance of motion vectors in the prior art is solved, and the accuracy and efficiency of video encoding and decoding are improved.

CN115176463BActive Publication Date: 2026-07-24DOUYIN VISION CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2020-12-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In existing video codec standards, the segmented motion vectors in inter-frame prediction are directly inherited from the merge candidate without refinement, resulting in insufficient accuracy.

Method used

A coding and decoding method called GMVD is proposed, which adds more than one motion vector difference (MVD) to the merge candidate derived motion vector (MV) to obtain a refined MV for motion compensation and storage within the block by notifying or deriving the block signaling.

Benefits of technology

It improves the accuracy and efficiency of video encoding and decoding, especially when using geometric segmentation mode, by improving the accuracy of motion compensation through refined motion vectors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115176463B_ABST
    Figure CN115176463B_ABST
Patent Text Reader

Abstract

Motion vector differences for blocks with geometric partitioning are described. An example method of video processing includes, for a conversion between a current video block of a video and a bitstream of the current video, determining that the current video block is coded with a geometric partitioning mode, deriving at least one refined motion vector (MV) for the current video block by adding at least one motion vector difference (MVD) signaled or derived for the current video block to a MV derived from a merge candidate associated with the current video, and performing the conversion based on the refined MV.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This application is the national phase of International Patent Application No. PCT / CN2020 / 141332, filed on December 30, 2020, which claims priority and interest in International Patent Application No. PCT / CN2019 / 129812, filed on December 30, 2019. The entire disclosure of International Patent Application No. PCT / CN2019 / 129812 is incorporated herein by reference as part of the disclosure of this application. Technical Field

[0003] This patent document relates to image and video encoding and decoding. Background Technology

[0004] Digital video accounts for the largest share of bandwidth usage on the Internet and other digital communication networks. As the number of connected user devices capable of receiving and displaying video increases, the bandwidth demand for digital video is expected to continue to grow. Summary of the Invention

[0005] This document discloses techniques that can be used by video encoders and decoders to perform cross-component adaptive loop filtering during video encoding or decoding.

[0006] In one example aspect, a video processing method is disclosed. The method includes: segmenting the video into multiple segments for a conversion between video units and a codec representation of the video; and performing the conversion on at least some of the multiple segments using a merge pattern with motion vector differences.

[0007] In another example, a different video processing method is disclosed. The method includes: a conversion between video units of a video and a codec representation of the video; segmenting the video into multiple segments; and performing the conversion based on the multiple segments, wherein the codec representation includes a field indicating the motion vector difference using one or more of two variables corresponding to the direction and magnitude of the motion vector difference.

[0008] In another example, a different video processing method is disclosed. The method includes: a conversion between video units of a video and a codec representation of the video; segmenting the video into multiple segments; and performing the conversion based on the multiple segments, wherein the codec representation includes sequentially ordered fields.

[0009] In another example aspect, a different video processing method is disclosed. The method includes: for a conversion between a current video block and a bitstream of the current video, determining the current video block to be encoded and decoded using a geometric segmentation mode; deriving at least one refined MV for the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to motion vectors (MVs) derived from merge candidates associated with the current video; and performing the conversion based on the refined MV.

[0010] In another example aspect, a method for storing a bitstream of video is disclosed. The method includes a conversion between a current video block and a bitstream of the current video; determining the current video block to be encoded and decoded using a geometric segmentation pattern; deriving at least one refined MV for the current video block by adding at least one MVD from a plurality of motion vector differences (MVDs) signaled or derived for the current video block to motion vectors (MVs) derived from merge candidates associated with the current video; and storing the bitstream in a non-transitory computer-readable recording medium.

[0011] In another example, a video encoder apparatus is disclosed. The video encoder includes a processor configured to implement the methods described above.

[0012] In yet another example, a video decoder apparatus is disclosed. The video decoder includes a processor configured to implement the methods described above.

[0013] In yet another example, a computer-readable medium is disclosed having code stored thereon. The code embodies one of the methods described herein in the form of processor-executable code.

[0014] This document describes these and other features throughout. Attached Figure Description

[0015] Figure 1 Example locations for spatial merge candidates are shown.

[0016] Figure 2 An example of candidate pairs considering redundancy checks for spatial merge candidates is shown.

[0017] Figure 3 This is a schematic diagram of the scaling of motion vectors for temporal merge candidates.

[0018] Figure 4 Examples of candidate positions C0 and C1 for temporal merge candidates are shown.

[0019] Figure 5An example of triangle segmentation based on intra-frame prediction is shown.

[0020] Figure 6 An example of unidirectional prediction MV selection for a triangular segmentation pattern is shown.

[0021] Figure 7 An example of the weights used in the mixing process is shown.

[0022] Figure 8 Examples of proposals and TPM designs in VTM-6.0 are shown.

[0023] Figure 9 An example of a GEO split boundary description is shown.

[0024] Figure 10A The edges supported in GEO are shown.

[0025] Figure 10B The diagram illustrates the geometric relationship between a given sample point location (x, y) and two edges.

[0026] Figure 11 An example of the UMVE search process is shown.

[0027] Figure 12 An example of a UMVE search point is shown.

[0028] Figure 13 This is a block diagram illustrating an example video processing system.

[0029] Figure 14 This is a block diagram of an example hardware platform used for video processing.

[0030] Figure 15 This is a flowchart of an example method for video processing.

[0031] Figure 16 This is a block diagram illustrating a video encoding / decoding system according to some embodiments of the present disclosure.

[0032] Figure 17 This is a block diagram illustrating an encoder according to some embodiments of the present disclosure.

[0033] Figure 18 This is a block diagram illustrating a decoder according to some embodiments of the present disclosure.

[0034] Figure 19 A flowchart of an example method for video processing is shown.

[0035] Figure 20 A flowchart of an example method for storing video bitstreams is shown. Detailed Implementation

[0036] The use of section headings in this document is for ease of understanding and does not limit the applicability of the technologies and embodiments disclosed in each section to that section. Furthermore, the use of H.266 terminology in some descriptions is merely for ease of understanding and not intended to limit the scope of the disclosed technologies. Therefore, the technologies described herein are also applicable to other video codec protocols and designs.

[0037] 1. Summary

[0038] This document relates to video codec technology. Specifically, it concerns inter-frame prediction and related techniques in video codecs. It can be applied to existing video codec standards (such as HEVC) or standards yet to be finalized (Multi-Functional Video Codec). It can also be applied to future video codec standards or video codecs.

[0039] 2. Background

[0040] Video codec standards have primarily evolved through the development of well-known ITU-T and ISO / IEC standards. ITU-T developed H.261 and H.263, while ISO / IEC developed MPEG-1 and MPEG-4 Visual. Together, these two organizations developed the H.262 / MPEG-2 video codec and the H.264 / MPEG-4 Advanced Video Codec (AVC) and H.265 / HEVC standards. Starting with H.262, video codec standards are based on a hybrid video codec architecture, utilizing temporal prediction plus transform coding. To explore future video codec technologies beyond HEVC, VCEG and MPEG jointly established the Joint Video Exploration Group (JVET) in 2015. Since then, JVET has adopted many new methods and incorporated them into reference software called the "Joint Exploration Model" (JEM). JVET meetings are held quarterly, and the goal of the new codec standard is to reduce the bitrate by 50% compared to HEVC. The new video codec standard was officially named Versatile Video Coding (VVC) at the JVET meeting in April 2018, and the first version of the VVC Test Model (VTM) was released at that time. With ongoing efforts to contribute to VVC standardization, new codec technologies have been adopted into the VVC standard at each JVET meeting. The VVC working draft and test model VTM are then updated after each meeting. The VVC project now aims for Technical Completion (FDIS) at the July 2020 meeting.

[0041] 2.1. Extended merge prediction

[0042] In VTM, the merge candidate list is constructed by sequentially including the following five types of candidates:

[0043] 1) MVP from airspace adjacent to CU

[0044] 2) Temporal MVP from collocated CU

[0045] 3) Historical MVPs from FIFO tables

[0046] 4) Paired average MVP

[0047] 5) Zero MV.

[0048] The size of the merge list is signaled in the stripe header, and the maximum allowed size of the merge list is 6 in the VTM. For each CU code in merge mode, the index of the best merge candidate is encoded using truncated univariate binarization (TU). The first bin (binary bits) of the merge index is decoded using context encoding, and bypass encoding is used for the other bins.

[0049] This session provides the process for generating merge candidates for each category.

[0050] 2.1.1. Derivation of Airspace Candidates

[0051] The derivation of spatial merge candidates in VVC is the same as that in HEVC. (The last sentence appears to be incomplete and possibly refers to a separate topic.) Figure 1 At most four merge candidates are selected from the candidates for the depicted positions. The derivation order is A0, B0, B1, A1, and B2. Position B2 is considered only if any CU at positions A0, B0, B1, or A1 is unavailable (e.g., because it belongs to another stripe or slice) or is intra-frame encoded / decoded. After adding the candidate at position A1, the addition of the remaining candidates is subject to a redundancy check, which ensures that candidates with the same motion information are excluded from the list, thereby improving encoding / decoding efficiency. To reduce computational complexity, not all possible candidate pairs are considered in the aforementioned redundancy check. Instead, only... Figure 2 Pairs linked by arrows are selected, and if the corresponding candidates used for redundancy checks do not have the same motion information, only one candidate is added to the list.

[0052] Figure 1 Example locations for spatial merge candidates are shown.

[0053] Figure 2 An example of candidate pairs is shown that take into account redundancy checks for spatial merge candidates.

[0054] 2.1.2. Derivation of Time-Domain Candidates

[0055] In this step, only one candidate is added to the list. Specifically, in the derivation of this temporal merge candidate, a scaled motion vector is derived based on the co-located CU belonging to the co-located reference image. The list of reference images used for deriving the co-located CU is explicitly signaled in the strip header. The scaled motion vector of the temporal merge candidate is obtained, as follows: Figure 3 The dashed lines in the diagram illustrate scaling from the motion vector of the co-located CU using POC distances tb and td, where tb is defined as the POC difference between the current image and the reference image, and td is defined as the POC difference between the co-located image and the reference image. The reference image index for the temporal merge candidate is set to zero.

[0056] Figure 3 This is a schematic diagram of the scaling of motion vectors for temporal merge candidates.

[0057] like Figure 4 As described, a temporary candidate position is selected between candidate C0 and C1. If the CU at position C0 is unavailable, intra-frame encoded or decoded, or outside the current line of the CTU, position C1 is used. Otherwise, position C0 is used in the derivation of the temporal merge candidate.

[0058] 2.1.3. Derivation of historical merge candidates

[0059] Following the spatial MVP and TMVP, historical MVP (HMVP) merge candidates are added to the merge list. In this method, motion information from previously encoded / decoded blocks is stored in a table and used as the MVP of the current CU. A table with multiple HMVP candidates is maintained during the encoding / decoding process. The table is reset (cleared) when a new CTU row is encountered. Whenever a CU that is not intra-frame encoded / decoded in a sub-block exists, the associated motion information is added as the last entry in the table as a new HMVP candidate.

[0060] In VTM, the HMVP table size S is set to 6, meaning that up to 6 history-based MVP (HMVP) candidates can be added to the table. When a new motion candidate is inserted into the table, a constrained First-In-First-Out (FIFO) rule is used, where a redundancy check is first applied to find if a duplicate HMVP exists in the table. If found, the duplicate HMVP is removed from the table, and all subsequent HMVP candidates are moved forward.

[0061] HMVP candidates can be used in the merge candidate list construction process. The most recent HMVP candidates in the table are checked sequentially and inserted into the candidate list after the TMVP candidates. Redundancy checks are applied to HMVP candidates for spatial or temporal merge candidates.

[0062] To reduce the number of redundant check operations, the following simplification is introduced:

[0063] 1. Is the number of HMPV candidates used for merging the list set to (N<=4)? M: (8–

[0064] M), where N represents the number of existing candidates in the merge list and M represents the number of available HMVP candidates in the table.

[0065] 2. Once the total number of available merge candidates reaches the maximum allowed merge candidates minus 1, the merge candidate list building process from HMVP is terminated.

[0066] 2.1.4. Derivation of Pairwise Average Merge Candidates

[0067] Pairwise averaging candidates are generated by averaging predefined candidate pairs from an existing merge candidate list, where the predefined pairs are defined as {(0,1), (0,2), (1,2), (0,3), (1,3), (2,3)}, where the numbers represent the merge index of the merge candidate list. The averaged motion vector is calculated separately for each reference list. If two motion vectors are available in a list, they are averaged even if they point to different reference images; if only one motion vector is available, that vector is used directly; if no motion vector is available, the list remains invalid.

[0068] If the merge list is not full after adding pairwise average merge candidates, insert zero MVP at the end until the maximum number of merge candidates is encountered.

[0069] 2.2. Triangle segmentation for inter-frame prediction

[0070] In VTM, Triangle Partitioning Mode (TPM) is supported for intra-frame prediction. Triangle Partitioning Mode is only applicable to CUs with 64 samples or larger and is used for encoding / decoding in skip or merge modes, but not in regular merge mode, MMVD mode, CIIP mode, or sub-block merge mode. CU-level flags are used to indicate whether Triangle Partitioning Mode is applied.

[0071] When using this pattern, use diagonal splitting or anti-diagonal splitting. Figure 5A CU is uniformly divided into two triangular segments. Each triangular segment in the CU performs inter-frame prediction using its own motion; only unidirectional prediction is allowed for each segment, meaning each segment has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with traditional bidirectional prediction, each CU requires only two motion-compensated predictions. The unidirectional prediction motion for each segment is directly derived from the merge candidate list constructed in 2.1 for extended merge prediction, and the selection of the unidirectional prediction motion from a given merge candidate in the list follows the procedure in 2.2.1.

[0072] If the triangular segmentation pattern is used for the current CU, further signaling is provided indicating the direction of the triangular segmentation (diagonal or anti-diagonal) and two merge indices (one for each segment). After predicting each triangular segment, a blending process with adaptive weights is used to adjust the sample values ​​along the diagonal or anti-diagonal sides. This is the predicted signal for the entire CU, and the transform and quantization processes are applied to the entire CU as in other prediction patterns. Finally, the motion field of the CU predicted using the triangular segmentation pattern is stored in a 4x4 cell, as described in 2.2.3.

[0073] 2.2.1. Construction of One-Way Prediction Candidate List

[0074] Given merge candidate indices, derive unidirectional predicted motion vectors from the merge candidate list constructed using the process in 2.1 for extended merge predictions, such as... Figure 6 As shown. For candidates in the list, their LX motion vectors are used as unidirectional predicted motion vectors for the triangle segmentation pattern, where X equals the parity of the merge candidate index values. These motion vectors are in Figure 6 The L(1-X) motion vector is marked with "x". In the absence of a corresponding LX motion vector, the L(1-X) motion vector of the same candidate in the expanded merge prediction candidate list is used as the unidirectional prediction motion vector of the triangle segmentation mode.

[0075] 2.2.2. Blending along the triangle partition edge

[0076] After each triangle segmentation using its own motion prediction, a blend is applied to the two predicted signals to derive samples near the diagonal or anti-diagonal sides. The following weights are used during the blending process:

[0077] • {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} are used for luminance, and {6 / 8, 4 / 8, 2 / 8} are used for chrominance, such as Figure 7 As shown.

[0078] Figure 7 An example of the weights used in the mixing process is shown.

[0079] 2.2.3. Sports Field Storage

[0080] The motion vectors of the CU encoded and decoded in a triangular segmentation pattern are stored in 4x4 cells. Depending on the position of each 4x4 cell, either a unidirectional or bidirectional predicted motion vector is stored. Mv1 and Mv2 are represented as the unidirectional predicted motion vectors for segmentation 1 and segmentation 2, respectively. If the 4x4 cell is located in... Figure 7 In the example shown, for an unweighted region, either Mv1 or Mv2 is stored for that 4x4 cell. Otherwise, if the 4x4 cell is located in a weighted region, a bidirectional predicted motion vector is stored. The bidirectional predicted motion vector is derived from Mv1 and Mv2 according to the following process:

[0081] 1) If Mv1 and Mv2 come from different lists of reference images (one from L0 and the other from L1), then simply combine Mv1 and Mv2 to form a bidirectional predicted motion vector.

[0082] 2) Otherwise, if Mv1 and Mv2 come from the same list, and without loss of generality, assume they both come from L0. In this case,

[0083] 2.a) If a reference image for Mv2 (or Mv1) appears in L1, then use that reference image in L1 to convert Mv2 (or Mv1) into an L1 motion vector. The two motion vectors are then combined to form a bidirectional predicted motion vector;

[0084] Otherwise, only store the unidirectional predicted motion Mv1 instead of the bidirectional predicted motion.

[0085] 2.3. Geometric Segmentation (GEO) for Inter-Frame Prediction

[0086] The Geometric Merge Model (GEO) was proposed at the 15th JVET conference in Gothenburg as an extension of the existing Triangle Prediction Model (TPM). At the 16th JVET conference in Geneva, a simpler-designed GEO model was selected as the CE anchor for further research. Currently, the GEO model is being investigated as an alternative to the existing TPM in VVC.

[0087] Figure 8 Examples of TPM in VTM-6.0 and other shapes proposed for non-rectangular inter-frame blocks are shown.

[0088] The split boundary of the geometric merge pattern is as follows Figure 9 Angle shown in and distance offset ρ i To describe. Angle. Represents the quantized angle between 0 and 360 degrees, distance offset ρ i Represents the maximum distance ρ max The quantization offset is not included. Furthermore, splitting directions that overlap with binary tree splitting and TPM splitting are not included.

[0089] GEO is applicable to block sizes of 8×8 or larger, and for each block size, there are 82 different division methods, distinguished by 24 angles and 4 edges relative to the center of the CU. Figure 10A Four edges are shown, evenly distributed within the CU along the normal vector direction, starting from Edge0, which passes through the center of the CU. Each segmentation pattern in GEO (i.e., a pair of angle indices and edge indices) is assigned a sample adaptive weight table to blend samples from parts of the two segments. The weight values ​​of the samples range from 0 to 8 and are determined by the L2 distance from the center of the sample to the edge. Essentially, a unity-gain constraint is followed when assigning weight values; that is, when a smaller weight value is assigned to one GEO segment, a larger complementary value is assigned to the other segment, for a total of 8.

[0090] The calculation of the weight value for each sample point is twofold: (a) calculating the displacement from the sample point location to the given edge and (c) mapping the calculated displacement to the weight value using a predefined lookup table. The method for calculating the displacement from the sample point location (x, y) to the given edge Edgei is essentially the same as calculating the displacement from (x, y) to Edge0 and subtracting the distance ρ between Edge0 and Edgei from that displacement. Figure 10B The geometric relationship between (x, y) and the edge is illustrated. Specifically, the displacement from (x, y) to Edge can be expressed by the following formula:

[0091]

[0092] Figure 10A The edges supported in GEO are shown. Figure 10B The geometric relationship between a given sample point location (x, y) and two edges is shown.

[0093] The value of ρ is a function of the maximum length of the normal vector (denoted by ρmax) and the edge index i, that is:

[0094]

[0095] Where N is the number of edges supported by GEO, and "1" is to prevent the last edge EdgeN-1 from being too close to the CU angle for some angular indices. Substituting equation (8) into equation (6), we can calculate the displacement from each sample point (x, y) to a given Edgei. In short, we will It is represented as wIdx(x, y). ρ needs to be calculated once for each CU, and wIdx(x, y) needs to be calculated once for each sample point, which involves multiplication.

[0096] 2.4. Merge with Motion Vector Difference (MMVD)

[0097] MMVD is also known as Ultimate Motion Vector Expression (UMVE).

[0098] UMVE is demonstrated. UMVE is used in skip mode or merge mode using the proposed motion vector representation method.

[0099] UMVE reuses the same merge candidates as those included in the regular merge candidate list in VVC. From these merge candidates, a base candidate can be selected and further extended using the proposed motion vector representation method.

[0100] UMVE provides a novel representation of motion vector difference (MVD) using a starting point, motion amplitude, and motion direction.

[0101] Figure 11 An example of the UMVE search process is shown.

[0102] Figure 12 An example of a UMVE search point is shown.

[0103] The proposed technique uses the merge candidate list as is. However, only candidates of the default merge type (MRG_TYPE_DEFAULT_N) are considered for UMVE extensions.

[0104] The base candidate index defines the starting point. The base candidate index indicates the best candidate among the candidates in the list, as shown below.

[0105] Table 2.4.1. Basic Candidate IDX

[0106]

[0107] If the number of basic candidates is equal to 1, then no signaling is sent to the basic candidate IDX.

[0108] The distance index is motion amplitude information. The distance index represents a predefined distance from the starting point. The predefined distances are as follows:

[0109] Table 2.4.2a. Distance to IDX

[0110]

[0111] During entropy encoding and decoding, the distance IDX is binarized in the bin containing truncated unary code as follows:

[0112] Table 2.4.2b. Distance IDX Binarization

[0113] Bin 0 10 110 1110 11110 111110 1111110 1111111

[0114] In arithmetic encoding and decoding, the first bin is encoded and decoded using a probabilistic context, and subsequent bins are encoded and decoded using an equal probability model (also known as bypass encoding and decoding).

[0115] The direction index represents the direction of MVD relative to the starting point. The direction index can represent four directions, as shown below.

[0116] Table 2.4.3. Directional IDX

[0117] x-axis + – N / A N / A y-axis N / A N / A + –

[0118] Immediately after sending the skip or merge flag, signal the UMVE flag. If the skip or merge flag is true, the UMVE flag is resolved. If the UMVE flag is equal to 1, the UMVE syntax is resolved. However, if it is not 1, the AFFINE flag is resolved. If the AFFINE flag is equal to 1, it is in AFFINE mode; otherwise, it is in skip / merge mode for the VTM, resolving the skip / merge index.

[0119] No additional line buffer is needed due to UMVE candidates. The software's skip / merge candidates are used directly as the base candidates. The input UMVE index is used to determine the MV supplement just before motion compensation. No long line buffer is needed for this.

[0120] In VVC, you can select the first or second merge candidate in the merge candidate list as the base candidate.

[0121] Figure 11 and Figure 12 The MVD search process and search points for UMVE are shown.

[0122] 3. The technical problems solved by the technical solutions disclosed in this document

[0123] In the current design of TPM / GEO, the segmented MV is directly inherited from the merge candidate without any refinement, which may be inaccurate.

[0124] 4. Solution Example

[0125] The detailed items below should be considered as examples to illustrate general concepts. These items should not be interpreted in a narrow way. Furthermore, these items can be combined in any way.

[0126] The term "GEO" can refer to an encoding / decoding method that splits a block into two or more sub-regions that cannot be achieved by a segmentation type (e.g., QT / BT / TT). The term "GEO" can also be referred to as "GPM". The term "GPM" can indicate Triangle Prediction Mode (TPM) and / or Geometric Merge Mode (GEO) and / or Wedge Prediction Mode.

[0127] The term "block" can refer to a codec block of CU and / or PU and / or TU.

[0128] A coding / decoding method called "GMVD" is proposed. In GMVD, more than one motion vector difference (MVD) can be derived for block signaling notification, wherein at least one MVD from multiple MVDs is added to the MV derived from the merge candidate to obtain a refined MV, which is used for motion compensation of at least one sample within the block and / or stored for subsequent processing. It should be noted that the following bullet points can be applied to either GMVD or the conventional MMVD method.

[0129] Vector operations in this document are performed in a manner common to mathematics. For example, (a, b)*c is equal to (a*c, b*c).

[0130] 1. A block of GMVD encoding / decoding can be divided into more than one segment.

[0131] a. In one example, at least one segment is non-rectangular.

[0132] b. In one example, the block is split into two segments by TPM, and the merge candidate can be a TPM merge candidate.

[0133] c. In one example, the block is split into two segments by GEO, and the merge candidate can be a GEO merge candidate.

[0134] 2. When a block of GMVD encoding / decoding is divided into N segments, signal or derive up to N MVDs for these segments, where N>=2. In one example, N equals 2.

[0135] a. In one example, signaling notifies or derives N MVDs, and each MVD corresponds to a specific segment.

[0136] i. Alternatively, the MVD corresponding to the segment is added to the MV of the segment to obtain a refined MV of the segment for motion compensation of the segment.

[0137] b. In one example, signaling notifies or derives K (K < N) MVDs, and one MVD can correspond to more than one partition.

[0138] c. In one example, how many MVDs to signal can depend on decoded information, such as block size, low-delay check flag, prediction direction, or reference picture list associated with N partitions.

[0139] d. In one example, signal the number of MVDs.

[0140] e. In one example, a partition can use multiple MVDs.

[0141] i. For example, multiple MVDs can be added to the MV of a partition to obtain a refined MV for that partition.

[0142] 3. Propose signaling a first message (such as a flag) to indicate whether at least one MVD among multiple MVDs is not equal to zero.

[0143] a. In one example, context encode or bypass encode the first message in arithmetic coding.

[0144] i. Alternatively, in addition, encode the first message with at least one context same as the syntax element indicating whether MMVD is applied.

[0145] b. Alternatively, in addition, signal the information related to multiple MVDs only when the first message indicates that at least one MVD among multiple MVDs is not equal to 0.

[0146] 4. Propose signaling a non-zero message for at least one MVD among multiple MVDs to indicate whether the MVD is equal to zero.

[0147] a. In one example, context encode at least one binarized bin of the message in arithmetic coding.

[0148] i. Alternatively, bypass encode at least one binarized bin of the message in arithmetic coding.

[0149] b. In one example, encode the message with at least one context same as the syntax element indicating whether MMVD is applied.

[0150] c. Alternatively, in addition, signal the information (such as direction and / or absolute value) associated with the MVD only when the non-zero message indicates that the MVD is not equal to zero.

[0151] d. Alternatively, if the first message indicates that at least one of the multiple MVDs is not equal to zero and all MVDs before the last MVD are zero, then no signaling is sent to the last MVD indicating that it is nonzero and the last MVD is inferred to be nonzero.

[0152] 5. Propose to notify at least one of the plurality of MVDs using syntax element signaling that indicates the level component of the MVD.

[0153] 6. Propose using syntax element signaling that indicates the vertical component of the MVD to notify at least one of the multiple MVDs.

[0154] 7. Propose a syntax element signaling to at least one of a plurality of MVDs using syntax elements representing the absolute values ​​of the horizontal and / or vertical components of the MVD.

[0155] 8. It can be proposed that the second MVD can be derived based on the first MVD that is signaled or derived prior to the signaling notification or the deriving of the second MVD.

[0156] 9. Similar to MMVD design, the MVD used in GMVD can be represented by two variables: direction and magnitude (or distance) represented by D. Furthermore, the following also apply:

[0157] a. Notify at least one of the multiple MVDs using a syntax element signaling that indicates the direction of the MVD (e.g., horizontal or vertical).

[0158] i. In one example, the direction may include, but is not limited to:

[0159] 1) MVD takes the form (1, 0) * D, where D is not less than 0.

[0160] 2) MVD takes the form (-1, 0) * D, where D is not less than 0.

[0161] 3) MVD adopts the form (0, 1) * D, where D is not less than 0.

[0162] 4) MVD takes the form (0, -1)*D, where D is not less than 0.

[0163] 5) MVD adopts the form (1, 1) * D, where D is not less than 0.

[0164] 6) MVD takes the form (-1, 1) * D, where D is not less than 0.

[0165] 7) MVD takes the form (1, -1)*D, where D is not less than 0.

[0166] 8) MVD takes the form (-1, -1)*D, where D is not less than 0.

[0167] ii. In one example, only the four directions mentioned in 9.ai1) to 9.ai4) are used.

[0168] iii. In one example, the direction can be binarized using a fixed-length codec, a unary codec, or an exponential Columbus codec.

[0169] b. In one example, D is greater than 0.

[0170] c. For blocks encoded and decoded using GMVD, the indication of MVD magnitude (or distance) can be represented by a D that can be signaled or derived.

[0171] i. In one example, D is restricted to the candidate set.

[0172] 1) In one example, the candidate set may contain zero.

[0173] 2) In one example, the candidate set may contain 1 / 4 sample points, 1 / 2 sample points, or other fractional sample points.

[0174] 3) In one example, the candidate set may include 1 sample point, 2 sample points, 4 sample points, 8 sample points, 16 sample points, 32 sample points, 64 sample points, 128 sample points, or other 2 samples. N Sample points.

[0175] 4) In one example, the candidate set may include 3 samples, 6 samples, 12 samples, 24 samples or other integer samples.

[0176] 5) In one example, the candidate set can be {1 / 4 sample points, 1 / 2 sample points, 1 sample point, 2 sample points, 4 sample points, 8 sample points, 16 sample points, 32 sample points}.

[0177] 6) In one example, the candidate set can be {1 / 4 sample, 1 / 2 sample, 1 sample, 2 sample, 3 sample, 4 sample, 6 sample, 8 sample, 16 sample}.

[0178] 7) In one example, the candidate set can be set to be equal to the blocks used for MMVD encoding and decoding in the same video processing unit (e.g., strip / slice / sub-picture / picture / sequence).

[0179] a. Alternatively, the candidate set may include more candidates in addition to those candidates used for MMVD encoding and decoding in the same video processing unit.

[0180] b. Alternatively, at least one candidate in the candidate set may be different from those candidates in the same video processing unit used for MMVD encoding and decoding.

[0181] c. Alternatively, at least one candidate in the candidate set can be equal to one of the candidates for the blocks used for MMVD encoding / decoding in the same video processing unit.

[0182] 8) In one example, instead of directly signaling D, the index of the selected MVD magnitude in the candidate set is signaled.

[0183] ii. In one example, multiple candidate sets of MVD can be predefined, and one of the multiple candidate sets can be selected for encoding / decoding the current GMVD-encoded / decoded block.

[0184] 1) In one example, the selection depends on a message signaled at the slice / picture / sequence level such as slice header / picture header / PPS / SPS.

[0185] 2) In one example, the selection depends on the same message for the blocks used for MMVD encoding / decoding, such as SPS_fpel_mmvd_enabled_flag.

[0186] iii. In one example, D is binarized into fixed-length coding, unary coding, or exponential Golomb coding.

[0187] iv. In one example, at least one binarized bin of D is context-coded in arithmetic coding.

[0188] v. Alternatively, at least one binarized bin of the message is bypass-coded in arithmetic coding.

[0189] vi. In one example, the first bin of D is context-coded in arithmetic coding.

[0190] 1) Alternatively, in addition, the other bins of D are bypass-coded.

[0191] d. For example, the MVD derived from D can be further modified before deriving the final MVD for the partition.

[0192] i. In one example, D can be modified to D = D << S, where S is an integer, such as 2.

[0193] ii. Whether the modification is applied can be implicitly inferred.

[0194] 1) For example, if the width and / or height of the current picture is greater than a threshold, the modification can be applied.

[0195] iii. Whether to apply modifications may depend on the message that can be signaled at the sequence level (e.g., in the SPS and / or sequence header), at the picture level (e.g., in the PPS and / or picture header), at the stripe level (e.g., in the stripe header), at the sub-picture level, at the slice level, at the CTU line level, or at the CTU level.

[0196] 1) For example, the same messages sps_fpel_mmvd_enabled_flag and / or pic_fpel_mmvd_enabled_flag in VVC can be used to control both MMVD and GMVD.

[0197] 2) Alternatively, separate signaling messages can be used to control MMVD and GMVD separately.

[0198] 10. In the example above, the block can be split into two segments using TPM or GEO, and two MVDs can be signaled or derived for these two segments.

[0199] a. Alternatively, the segmented MV is calculated as the sum of the MV derived from the segmented TPM or GEO and the segmented MVD.

[0200] b. Alternatively, the two motion compensations and the weighted summation of the two MVs of the two segments are performed in the same manner as TPM or GEO.

[0201] c. Alternatively, in addition, perform the two-part MV stored procedure in the same manner as TPM or GEO.

[0202] 11. In the above example, whether signaling notification or deriving of the proposed multiple MVDs can be conditional.

[0203] a. In one example, whether signaling notification or export of multiple proposed MVDs may depend on whether GEO is enabled for the current block.

[0204] b. In one example, if there is no signaling notification or derivation of the proposed multiple MVDs, they are inferred to be zero.

[0205] c. In one example, signaling notification may be provided at the sequence level (e.g., in the SPS and / or sequence header), at the picture level (e.g., in the PPS and / or picture header), at the stripe level (e.g., in the stripe header), at the sub-picture level, at the slice level, at the CTU line level, or at the CTU level, to indicate whether the proposed GMVD is being signaled or derived.

[0206] ii. Alternatively, GMVD can be conditionally signaled as to whether it is enabled, such as when TPM / GEO is enabled.

[0207] d. In one example, whether to signal or derive the proposed multiple MVDs may depend on the block width (W) and / or the block height (H). For example, if at least one or any combination of the following conditions is satisfied, the proposed multiple MVDs may not be signaled or derived.

[0208] i. W >= T1. For example, T1 = 64.

[0209] ii. H >= T2. For example, T2 = 64.

[0210] iii. W >= T3 * H. For example, T3 = 4.

[0211] iv. H >= T4 * H. For example, T4 = 8.

[0212] v. W <= T5. For example, T5 = 8.

[0213] vi. H <= T6. For example, T6 = 8.

[0214] vii. W * H >= T7, for example, T7 = 2048

[0215] viii. W * H <= T8, for example, T8 = 64

[0216] ix. W == T9 or H == T10. (For example, T9 = T10 = 4)

[0217] x. W / H > T11 or max(W, H) / min(W, H) >= T11

[0218] xi. W / H < T11 or max(W, H) / min(W, H) < T11

[0219] e. Whether to signal or derive the proposed multiple MVDs may depend on the merge candidate index.

[0220] i. In one example, the merge candidate index is greater than a predetermined value.

[0221] 12. For blocks with GMVD encoding / decoding, the MVD may be signaled before the merge candidate index.

[0222] a. Alternatively, in addition, the merge candidate index may be signaled according to whether the proposed multiple MVDs are signaled or derived.

[0223] b. Alternatively, in addition, how to signal the merge candidate index of a block with GMVD encoding / decoding (e.g., the binarization process) may depend on the use of GMVD.

[0224] i. In one example, if signaling notification or deriving of multiple proposed MVDs, the maximum merge candidate index that can be signaled can be reduced.

[0225] 5. Examples

[0226] 5.1. An example of draft changes

[0227] merge data syntax

[0228]

[0229]

[0230] merge data semantics

[0231] `tmvd_flag[x0][y0]` specifies whether to apply a triangle prediction with motion vector difference to the current codec unit. Array indices `x0` and `y0` specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image. If `tmvd_flag[x0][y0]` does not exist, it is inferred to be equal to 0.

[0232] `tmvd_part_flag[x0][y0][partIdx]` specifies whether to apply a triangular prediction with motion vector difference to the segment in the current codec unit whose index is equal to `partIdx`. The array indices `x0` and `y0` specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image. When `tmvd_part_flag[x0][y0][partIdx]` does not exist, if `tmvd_flag[x0][y0]` equals 1 and `partIdx` equals 1, it is inferred to be equal to 1. Otherwise, it is inferred to be equal to 0.

[0233] tmvd_distance_idx[x0][y0][partIdx] specifies the index used to export TmvdDistance[x0][y0][partIdx]. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0234] TmvdDistanceArray[] is set to equal to {4, 8, 16, 32, 48, 64, 96, 128, 256} and TmvdDistance[x0][y0][partIdx] is set to equal to TmvdDistanceArray[tmvd_distance_idx[x0]][y0][partIdx]].

[0235] mmvd_direction_idx[x0][y0] specifies the index used to export TmvdBaseMV[x0][y0][partIdx][compIdx] for compIdx = 0..1. The array indices x0 and y0 specify the position (x0, y0) of the top-left luminance sample of the codec block under consideration relative to the top-left luminance sample of the image.

[0236] TmvdBaseArray[][] is set to equal to {{1,0},{-1,0},{0,1},{0,-1},{1,1},{1,-1},{-1,1},{-1,-1}.

[0237] For compIdx = 0..1, TmvdBase[x0][y0][partIdx][compIdx] is set to be equal to TmvdBaseArray[mmvd_direction_idx[x0][y0]][compIdx].

[0238] When tmvd_part_flag[x0][y0][partIdx] equals zero, for compIdx = 0..1, TmvdOffset[x0][y0][partIdx][compIdx] is set to zero. Otherwise, for compIdx = 0..1, TmvdOffset[x0][y0][partIdx][compIdx] is set to equal to TmvdBaseMV[x0][y0][partIdx][compIdx] * TmvdDistance[x0][y0][partIdx] * (pic_fpel_mmvd_enabled_flag == 1? 4:1).

[0239] Derivation of the brightness motion vector for merge triangle mode

[0240] This procedure is called only when MergeTriangleFlag[xCb][yCb] equals 1, where (xCb, yCb) specifies the top-left sample of the current luminance codec block relative to the top-left luminance sample of the current image.

[0241] The input for this process is:

[0242] – The brightness position (xCb, yCb) of the top-left corner sample of the current luminance block relative to the top-left corner luminance sample of the current image.

[0243] – The variable cbWidth specifies the width of the current codec block in the luminance sample.

[0244] – The variable cbHeight specifies the height of the current codec block in the luminance sample.

[0245] The output of this process is:

[0246] Luminance motion vectors mvA and mvB with a precision of -1 / 16 fractional sample point.

[0247] –Refer to indices refIdxA and refIdxB,

[0248] – Prediction list flags predListFlagA and predListFlagB.

[0249] Motion vectors mvA and mvB, reference indices refIdxA and refIdxB, and prediction list flags predListFlagA and predListFlagB are derived by the following sequential steps:

[0250] 1. Using the brightness position (xCb, yCb), variables cbWidth and cbHeight as inputs, invoke the luminance motion vector push process as specified in Clause 8.5.2.2 and output the brightness motion vectors MVL0[0][0], MVL1[0][0], reference indices refIdxL0, refIdxL1, prediction list using flags predFlagL0[0][0] and predFlagL1[0][0], bidirectional prediction weight index bcwIdx and merge candidate list mergeCandList.

[0251] 2. Using merge_triangle_idx0[xCb][yCb] and merge_triangle_idx1[xCb][yCb], the variables m and n, which are the merge indices for the triangle segmentation of 0 and 1 respectively, are derived as follows:

[0252] m=merge_triangle_idx0[xCb][yCb] (657)

[0253] n=merge_triangle_idx1[xCb][yCb]+(merge_triangle_idx1[xCb][yCb]>=m)? 1:0(658)

[0254] 3. Let refIdxL0M and refIdxL1M, predFlagL0M and predFlagL1M, and MVL0M and MVL1M be the reference index, prediction list utilization flag, and motion vector (M = mergeCandList[m]) of the merge candidate M at position m in the merge candidate list mergeCandList.

[0255] 4. Variable X is set to equal to (m&0x01).

[0256] 5. When predFlagLXM equals 0, X is set to equal to (1-X).

[0257] 6. The following applies:

[0258] mvA[0]=MVLXM[0]+TmvdOffset[x0][y0][0][0] (659)

[0259] mvA[1]=MVLXM[1]+TmvdOffset[x0][y0][0][1] (660)

[0260] refIdxA=refIdxLXM (661)

[0261] predListFlagA = X (662)

[0262] 7. Let refIdxL0N and refIdxL1N, predFlagL0N and predFlagL1N, and MVL0N and MVL1N be the reference index, prediction list flag, and motion vector (N = mergeCandList[n]) of the merge candidate N at position m in the merge candidate list mergeCandList.

[0263] 8. Variable X is set to equal to (n&0x01).

[0264] 9. When predFlagLXN equals 0, X is set to (1-X).

[0265] 10. The following applies:

[0266] mvB[0]=MVLXN[0]+TmvdOffset[x0][y0][1][0] (663)

[0267] mvB[1]=MVLXN[1]+TmvdOffset[x0][y0][1][1] (664)

[0268] refIdxB=refIdxLXN (665)

[0269] predListFlagB = X (666)

[0270] Table 123 Syntax Elements and Associated Binarization

[0271]

[0272] Table 123 Syntax Elements and Associated Binarization

[0273]

[0274] Table 128 - Assigning ctxInc to syntax elements with context encoding / decoding bin

[0275]

[0276] Table 128 - Assigning ctxInc to syntax elements with context encoding / decoding bin

[0277]

[0278] Figure 13 This is a block diagram illustrating an example video processing system 1900 in which various techniques disclosed herein can be implemented. Various implementations may include some or all of the components of system 1900. System 1900 may include an input 1902 for receiving video content. The video content may be received in a raw or uncompressed format (e.g., 8 or 10-bit multi-component pixel values), or in a compressed or encoded format. Input 1902 may represent a network interface, a peripheral bus interface, or a storage interface. Examples of network interfaces include wired interfaces such as Ethernet, Passive Optical Network (PON), and wireless interfaces such as Wi-Fi or cellular interfaces.

[0279] System 1900 may include codec component 1904, which can implement various codec or encoding methods described in this document. Codec component 1904 can reduce the average bit rate of the video from input 1902 to the output of codec component 1904 to produce a codec representation of the video. Therefore, codec techniques are sometimes referred to as video compression or video transcoding techniques. The output of codec component 1904 can be stored or transmitted via a connected communication (as indicated by component 1906). The stored or communicated bitstream (or codec) representation of the video received at input 1902 can be used by component 1908 to generate pixel values ​​or sent to displayable video at display interface 1910. The process of generating user-viewable video from the bitstream representation is sometimes referred to as video decompression. Furthermore, although some video processing operations are referred to as “codec” operations or tools, it should be understood that codec tools or operations are used at the encoder and will be performed by the decoder with the corresponding decoding tools or operations preserving the results of the codec.

[0280] Examples of peripheral bus interfaces or display interfaces may include Universal Serial Bus (USB), High Definition Multimedia Interface (HDMI), or DisplayPort. Examples of storage interfaces include SATA (Serial Advanced Technology Accessory), PCI, IDE, etc. The technologies described in this document can be implemented in a variety of electronic devices, such as mobile phones, laptops, smartphones, or other devices capable of performing digital data processing and / or video display.

[0281] Figure 14 This is a block diagram of a video processing apparatus 3600. Apparatus 3600 can be used to implement one or more methods described herein. Apparatus 3600 can be embodied in smartphones, tablets, computers, Internet of Things (IoT) receivers, etc. Apparatus 3600 may include one or more processors 3602, one or more memories 3604, and video processing hardware 3606. The processors 3602 can be configured to implement one or more methods described in this document. The memories 3604 can be used to store data and code for implementing the methods and techniques described herein. The video processing hardware 3606 can be used to implement some of the techniques described herein in hardware circuitry.

[0282] Figure 16 This is a block diagram illustrating an exemplary video codec system 100 that can utilize the techniques disclosed herein.

[0283] like Figure 16As shown, the video encoding / decoding system 100 may include a source device 110 and a destination device 120. The source device 110 generates encoded video data and may be referred to as a video encoding device. The destination device 120 can decode the encoded video data generated by the source device 110 and may be referred to as a video decoding device.

[0284] The source device 110 may include a video source 112, a video encoder 114, and an input / output (I / O) interface 116.

[0285] Video source 112 may include sources such as video capture devices, interfaces for receiving video data from video content providers, and / or computer graphics systems for generating video data, or combinations thereof. Video data may include one or more images. Video codec 114 encodes the video data from video source 112 to generate a bitstream. The bitstream may include a sequence of bits forming a codec representation of the video data. The bitstream may include codec images and associated data. A codec image is a codec representation of an image. Associated data may include sequence parameter sets, image parameter sets, and other syntax structures. I / O interface 116 may include a modulator / demodulator (modem) and / or a transmitter. Encoded video data can be transmitted directly to destination device 120 via network 130a through I / O interface 116. Encoded video data may also be stored on storage medium / server 130b for access by destination device 120.

[0286] Destination device 120 may include I / O interface 126, video decoder 124 and display device 122.

[0287] I / O interface 126 may include a receiver and / or a modem. I / O interface 126 may acquire encoded video data from source device 110 or storage medium / server 130b. Video decoder 124 may decode the encoded video data. Display device 122 may display the decoded video data to a user. Display device 122 may be integrated with destination device 120, or may be external to destination device 120 configured to interface with an external display device.

[0288] The video encoder 114 and the video decoder 124 can operate according to video compression standards, such as the High Efficiency Video Coding (HEVC) standard, the Multi-Function Video Coding (VVC) standard, and other current and / or further standards.

[0289] Figure 17 This is a block diagram illustrating an example of a video encoder 200, which can be... Figure 16 The video encoder 114 in the system 100 shown.

[0290] The video encoder 200 can be configured to perform any or all of the technologies disclosed herein. Figure 17 In the example, the video encoder 200 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video encoder 200. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0291] The functional components of the video encoder 200 may include a segmentation unit 201, a prediction unit 202, a residual generation unit 207, a transform unit 208, a quantization unit 209, an inverse quantization unit 210, an inverse transform unit 211, a reconstruction unit 212, a buffer 213, and an entropy coding unit 214. The prediction unit 202 may include a mode selection unit 203, a motion estimation unit 204, a motion compensation unit 205, and an intra-frame prediction unit 206.

[0292] In other examples, the video encoder 200 may include more, fewer, or different functional components. In one example, the prediction unit 202 may include an intra-block copy (IBC) unit. The IBC unit can perform prediction in IBC mode, where at least one reference picture is the picture containing the current video block.

[0293] Furthermore, some components (such as motion estimation unit 204 and motion compensation unit 205) can be highly integrated, but for illustrative purposes, in Figure 17 The examples are shown separately.

[0294] The segmentation unit 201 can segment an image into one or more video blocks. The video encoder 200 and the video decoder 300 can support various video block sizes.

[0295] The mode selection unit 203 can, for example, select one of the encoding / decoding modes (intra-frame or inter-frame) based on the error result, and provide the resulting intra-frame or inter-frame encoded / decoded block to the residual generation unit 207 to generate residual block data, and provide it to the reconstruction unit 212 to reconstruct the coded block for use as a reference picture. In some examples, the mode selection unit 203 can select an intra-inter-frame (CIIP) mode in which prediction is based on a combination of inter-frame prediction signals and intra-frame prediction signals. In the case of inter-frame prediction, the mode selection unit 203 can also select a motion vector resolution (e.g., sub-pixel or integer pixel precision) for the block.

[0296] To perform inter-frame prediction on the current video block, motion estimation unit 204 can generate motion information for the current video block by comparing it with one or more reference frames from buffer 213. Motion compensation unit 205 can determine the predicted video block for the current video block based on the motion information and the decoded samples from buffer 213 that are different from the images associated with the current video block (e.g., reference images).

[0297] For example, depending on whether the current video block is in an I-band, P-band, or B-band, the motion estimation unit 204 and the motion compensation unit 205 can perform different operations on the current video block.

[0298] In some examples, motion estimation unit 204 can perform unidirectional prediction on the current video block, and can search for a reference video block for the current video block in the reference images of list 0 or list 1. Then, motion estimation unit 204 can generate a reference index indicating the reference image in list 0 or list 1 containing the reference video block and a motion vector indicating the spatial displacement between the current video block and the reference video block. Motion estimation unit 204 can output the reference index, prediction direction indicator, and motion vector as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current block based on the reference video block indicated by the motion information of the current video block.

[0299] In other examples, motion estimation unit 204 can perform bidirectional prediction on the current video block. Motion estimation unit 204 can search for a reference video block for the current video block in the reference images of list 0, and can also search for another reference video block for the current video in the reference images of list 1. Then, motion estimation unit 204 can generate reference indices indicating the reference images in lists 0 and 1 containing the reference video blocks, and motion vectors indicating the spatial displacement between the reference video blocks and the current video block. Motion estimation unit 204 can output the reference index and motion vector of the current video block as motion information for the current video block. Motion compensation unit 205 can generate a predicted video block for the current video block based on the reference video blocks indicated by the motion information of the current video block.

[0300] In some examples, the motion estimation unit 204 can output a complete set of motion information for the decoder's decoding processing.

[0301] In some examples, motion estimation unit 204 may not output the full set of motion information for the current video. Instead, motion estimation unit 204 may refer to the motion information of another video block to signal the motion information of the current video block. For example, motion estimation unit 204 may determine that the motion information of the current video block is sufficiently similar to the motion information of neighboring video blocks.

[0302] In one example, the motion estimation unit 204 may indicate a value in the syntax structure associated with the current video block that indicates to the video decoder 300 that the current video block has the same motion information as another video block.

[0303] In another example, motion estimation unit 204 can identify another video block and motion vector difference (MVD) in the syntax structure associated with the current video block. The motion vector difference indicates the difference between the motion vector of the current video block and the motion vector of the indicated video block. Video decoder 300 can use the motion vector of the indicated video block and the motion vector difference to determine the motion vector of the current video block.

[0304] As described above, the video encoder 200 can predictively signal motion vectors. Two examples of predictive signaling techniques that can be implemented by the video encoder 200 include Advanced Motion Vector Prediction (AMVP) and Merge Pattern Signaling.

[0305] Intra-prediction unit 206 can perform intra-prediction on the current video block. When intra-prediction unit 206 performs intra-prediction on the current video block, it can generate prediction data for the current video block based on decoded samples from other video blocks in the same frame. The prediction data for the current video block can include the predicted video block and various syntax elements.

[0306] The residual generation unit 207 can generate residual data for the current video block by subtracting (e.g., indicated by a minus sign) multiple predicted video blocks from the current video block. The residual data for the current video block can include residual video blocks corresponding to different sample components of the samples in the current video block.

[0307] In other examples, such as in skip mode, there may be no residual data for the current video block, and the residual generation unit 207 may not perform the subtraction operation.

[0308] The transform processing unit 208 can generate one or more transform coefficient video blocks for the current video block by applying one or more transforms to the residual video blocks associated with the current video block.

[0309] After the transform processing unit 208 generates a transform coefficient video block associated with the current video block, the quantization unit 209 can quantize the transform coefficient video block associated with the current video block based on one or more quantization parameter (QP) values ​​associated with the current video block.

[0310] Inverse quantization unit 210 and inverse transform unit 211 can apply inverse quantization and inverse transform to the transform coefficient video block, respectively, to reconstruct the residual video block from the transform coefficient video block. Reconstruction unit 212 can add the reconstructed residual video block to the corresponding samples of one or more predicted video blocks generated by prediction unit 202 to produce a reconstructed video block associated with the current block for storage in buffer 213.

[0311] After the video block is reconstructed by the reconstruction unit 212, a loop filtering operation can be performed to reduce video block artifacts in the video block.

[0312] Entropy encoding unit 214 can receive data from other functional components of video encoder 200. When entropy encoding unit 214 receives data, it can perform one or more entropy encoding operations to generate entropy encoded data and output a bit stream including the entropy encoded data.

[0313] Figure 18 It shows that it can be Figure 16 A block diagram of an example of video decoder 300 in system 100, showing video decoder 114.

[0314] The video decoder 300 can be configured to perform any or all of the technologies disclosed herein. Figure 17 In the example, the video decoder 300 includes multiple functional components. The techniques described in this disclosure can be shared among the various components of the video decoder 300. In some examples, the processor can be configured to perform any or all of the techniques described in this disclosure.

[0315] exist Figure 18 In the example, video decoder 300 includes an entropy decoding unit 301, a motion compensation unit 302, an intra-frame prediction unit 303, an inverse quantization unit 304, an inverse transform unit 305, a reconstruction unit 306, and a buffer 307. In some examples, video decoder 300 can perform encoding passes typically described with respect to video encoder 200. Figure 17 The opposite decoding pass.

[0316] The entropy decoding unit 301 can retrieve the encoded bitstream. The encoded bitstream may include entropy-coded video data (e.g., encoded blocks of video data). The entropy decoding unit 301 can decode the entropy-coded video data, and the motion compensation unit 302 can determine motion information from the entropy-coded video data, including motion vectors, motion vector precision, reference image list index, and other motion information. For example, the motion compensation unit 302 can determine this information by executing AMVP and Merge modes.

[0317] The motion compensation unit 302 can generate motion compensation blocks, possibly performing interpolation based on an interpolation filter. The syntax elements may include identifiers for the interpolation filters used at sub-pixel precision.

[0318] The motion compensation unit 302 can use the interpolation filter used by the video encoder 200 during video block encoding to calculate the interpolated values ​​of sub-integer pixels of the reference block. The motion compensation unit 302 can determine the interpolation filter used by the video encoder 200 based on the received syntax information, and use the interpolation filter to generate the prediction block.

[0319] The motion compensation unit 302 may use some syntax information to determine the size of the blocks used to encode (one or more) frames and / or (one or more) stripes of the encoded video sequence, segmentation information describing how each macroblock of the image of the encoded video sequence is segmented, a mode indicating how each segment is encoded, one or more reference frames (and a list of reference frames) for each inter-frame encoded block, and other information for decoding the encoded video sequence.

[0320] Intra-prediction unit 303 can use, for example, an intra-prediction mode received in the bitstream to form prediction blocks from spatially adjacent blocks. Inverse quantization unit 304 performs inverse quantization (i.e., dequantization) on the quantized video block coefficients provided in the bitstream and decoded by entropy decoding unit 301. Inverse transform unit 305 applies an inverse transform.

[0321] The reconstruction unit 306 can add the residual block to the corresponding prediction block generated by the motion compensation unit 202 or the intra-frame prediction unit 303 to form a decoded block. If necessary, a deblocking filter can also be applied to filter the decoded block to remove block artifacts. The decoded video block is then stored in a buffer 307, which provides a reference block for subsequent motion compensation / intra-frame prediction and also generates the decoded video for presentation on a display device.

[0322] The following provides a list of preferred solutions for some of the embodiments.

[0323] The following solutions illustrate example embodiments of the techniques discussed in the previous section (e.g., item 1).

[0324] 1. A video processing method (e.g., Figure 15 The method 1500 shown includes, for the conversion between video units of a video and the codec representation of the video, segmenting the video into multiple segments, and performing the conversion based on the multiple segments using a merge mode with motion vector differences for at least some of the multiple segments.

[0325] 2. The method according to claim 1, wherein at least one of the plurality of partitions is a non-rectangular partition.

[0326] 3. The method according to any one of claims 1-2, wherein the partitioning uses a triangular partitioning pattern, and wherein the merge mode with a motion vector difference uses a triangular partitioning pattern merge candidate as a merge candidate.

[0327] The following scenarios illustrate example embodiments of the techniques discussed in the previous section (e.g., item 2).

[0328] 4. The method according to any one of claims 1-3, wherein the plurality of partitions includes N partitions, where N is an integer greater than or equal to 2, and wherein the codec representation includes at most N motion vector differences, or wherein the transformation uses at most N motion vector differences.

[0329] 5. The method according to claim 4, wherein the codec representation includes a field indicating K partitions, where K < N, and wherein one or more of the plurality of partitions share a motion vector difference.

[0330] The following scenarios illustrate example embodiments of the techniques discussed in the previous section (e.g., item 3).

[0331] 6. The method according to claims 4-5, wherein the field in the codec representation indicates whether at least one of the N motion vector differences is a non-zero value.

[0332] 7. The method according to claim 6, wherein the field is contextually coded in the codec representation.

[0333] The following scenarios illustrate example embodiments of the techniques discussed in the previous section (e.g., items 5, 6, 7).

[0334] 8. The method according to any one of claims 4-7, wherein at least one of the N motion vector differences is signaled as a syntax element indicating the horizontal component of the motion vector difference.

[0335] 9. The method according to any one of claims 4-7, wherein at least one of the N motion vector differences is signaled as a syntax element indicating the vertical component of the motion vector difference.

[0336] 10. The method according to any one of claims 4-7, wherein at least one of the N motion vector differences is signaled as a syntax element indicating the absolute value of the horizontal or vertical component of the motion vector difference.

[0337] The following scheme illustrates example embodiments of the techniques discussed in the previous section (e.g., item 8).

[0338] 11. The method according to any one of Schemes 4-7, wherein at least one of the N motion vector differences is derived from another motion vector difference.

[0339] The following scheme illustrates example embodiments of the techniques discussed in the previous section (e.g., item 9).

[0340] 12. A video processing method comprising: converting between video units of a video and a codec representation of the video, segmenting the video into a plurality of segments, and performing the conversion based on the plurality of segments, wherein the codec representation includes a field indicating a motion vector difference using one or more of two variables corresponding to the direction and magnitude of the motion vector difference.

[0341] 13. The method according to scheme 12, wherein one of the two variables is a direction variable.

[0342] 14. The method according to schemes 12-13, wherein the amplitude is allowed to have values ​​belonging to a set.

[0343] The following scheme illustrates example embodiments of the techniques discussed in the previous section (e.g., item 12).

[0344] 15. A video processing method comprising: converting between video units of a video and a codec representation of the video; segmenting the video into multiple segments; and performing the conversion based on the multiple segments, wherein the codec representation includes fields in sequence.

[0345] 16. The method according to scheme 15, wherein the order includes motion vector differences occurring before the merge candidate index in the codec representation.

[0346] 17. The method according to claim 15, wherein the order depends on the characteristics of the video unit.

[0347] 18. The method according to any one of the above schemes, wherein the video unit includes an encoding / decoding unit.

[0348] 19. The method according to any one of the above schemes, wherein the video unit comprises a video block.

[0349] 20. The method according to any one of claims 1 to 19, wherein the conversion includes encoding / decoding the video into the codec representation.

[0350] 21. The method according to any one of claims 1 to 19, wherein the conversion includes decoding the codec representation to generate pixel values ​​of the video.

[0351] 22. A video decoding apparatus, comprising a processor configured to implement one or more of the methods described in claims 1-21.

[0352] 23. A video encoding apparatus, comprising a processor configured to implement one or more of the methods described in schemes 1-21.

[0353] 24. A computer program product having stored computer code thereon, which, when executed by a processor, causes the processor to implement the method described in any one of schemes 1-21.

[0354] 25. The methods, apparatus or systems described in this document.

[0355] Figure 19 A flowchart of an example method for video processing is shown. The method includes a conversion between a current video block and a bitstream of the current video, determining (1902) the current video block to be encoded and decoded using a geometric segmentation mode; deriving (1904) at least one refined MV for the current video block by adding at least one of a plurality of motion vector differences (MVDs) derived or signaled for the current video block to a motion vector (MV) derived from a merge candidate associated with the current video block; and performing a conversion based on the refined MV (1906).

[0356] In some examples, the current video block is a block encoded with Geometric Motion Vector Difference (GMVD) or a block encoded with Merge (MMVD) with Motion Vector Difference.

[0357] In some examples, the geometric segmentation pattern includes multiple segmentation schemes and at least one segmentation scheme splits the current video block into two or more segments, at least one of which is neither square nor rectangular.

[0358] In some examples, the geometric segmentation pattern includes the triangle segmentation pattern.

[0359] In some examples, the geometric segmentation pattern includes the geometric merge pattern.

[0360] In some examples, the current video block is divided into N segments, and the segmentation signaling for the current video block is notified or at most N MVDs are derived, where N>=2.

[0361] In some examples, N = 2.

[0362] In some examples, the segmentation signaling for the current video block is notified or N MVDs are derived, and each MVD corresponds to a specific segmentation.

[0363] In some examples, the MVD corresponding to the segment is added to the MV of the segment to obtain a refined MV of the segment for use in motion compensation of the segment.

[0364] In some examples, signaling notifies or derives K MVDs, and at least one MVD corresponds to more than one segment, where K <N。

[0365] In some examples, the number of MVDs to be signaled depends on the decoded information, which includes at least one of the following: block size associated with N segments, low-latency check flag, predicted direction, or a list of reference images.

[0366] In some examples, the number of MVDs is included in the bitstream.

[0367] In some examples, multiple MVDs are used by a single split.

[0368] In some examples, multiple MVDs are added to the segmented MV to obtain a more refined MV.

[0369] In some examples, the first message is included in the bitstream to indicate whether at least one of the multiple MVDs is not equal to zero.

[0370] In some examples, the first message is a flag.

[0371] In some examples, the first message is either context-coded or bypass-coded in arithmetic encoding / decoding.

[0372] In some examples, the first message is encoded and decoded using the same context as the syntax element that indicates whether MMVD has been applied.

[0373] In some examples, information related to multiple MVDs is included in the bitstream only if the first message indicates that at least one of the multiple MVDs is not equal to 0.

[0374] In some examples, a second message for at least one of the multiple MVDs is included in the bitstream to indicate whether the MVD is equal to zero, where the second message is a non-zero message.

[0375] In some examples, in arithmetic encoding and decoding, context encoding and decoding of at least one binarized bin of the second message.

[0376] In some examples, at least one binarized bin of the second message is bypassed in arithmetic coding.

[0377] In some examples, the second message is encoded and decoded using at least one context that is the same as the syntax element indicating whether MMVD has been applied.

[0378] In some examples, the information associated with MVD is included in the bitstream only when the second message indicates that MVD is not equal to 0.

[0379] In some examples, at least one of the multiple MVDs is included in a bitstream with a syntax element that indicates the horizontal component of the MVD, or at least one of the multiple MVDs is included in a bitstream with a syntax element that indicates the vertical component of the MVD.

[0380] In some examples, at least one of the multiple MVDs is included in a bitstream with syntax elements that indicate the absolute values ​​of the horizontal and / or vertical components of the MVD.

[0381] In some examples, the second MVD is derived based on the first MVD that is signaled or exported before the second MVD is signaled or exported.

[0382] In some examples, the MVD used in GMVD is represented by two variables: the first variable is the direction, and the second variable is the magnitude represented by D.

[0383] In some examples, at least one of the multiple MVDs is included in a bitstream with syntax elements that indicate the direction of the MVD.

[0384] In some examples, the direction includes at least one of the following:

[0385] 1) MVD adopts the form (1, 0) * D, where D is not less than 0;

[0386] 2) MVD takes the form (-1, 0) * D, where D is not less than 0;

[0387] 3) MVD adopts the form (0, 1) * D, where D is not less than 0;

[0388] 4) MVD takes the form (0, -1)*D, where D is not less than 0;

[0389] 5) MVD adopts the form (1, 1) * D, where D is not less than 0;

[0390] 6) MVD takes the form (-1, 1) * D, where D is not less than 0;

[0391] 7) MVD takes the form (1, -1)*D, where D is not less than 0; and

[0392] 8) MVD takes the form (-1, -1)*D, where D is not less than 0.

[0393] In some examples, the direction includes at least one of the following:

[0394] 1) MVD adopts the form (1, 0) * D, where D is not less than 0;

[0395] 2) MVD takes the form (-1, 0) * D, where D is not less than 0;

[0396] 3) MVD adopts the form (0, 1) * D, where D is not less than 0; and

[0397] 4) MVD takes the form (0, -1)*D, where D is not less than 0.

[0398] In some examples, fixed-length encoding / decoding, unary encoding / decoding, or exponential Columbus encoding / decoding are used to binarize the direction.

[0399] In some examples, D is greater than 0.

[0400] In some examples, for blocks encoded and decoded by GMVD, the indication of the MVD magnitude is provided by signaling notification or derived D representation.

[0401] In some examples, D is restricted to the candidate set.

[0402] In some examples, the candidate set contains zero.

[0403] In some examples, the candidate set includes 1 / 4 sample points, 1 / 2 sample points, or other fractional sample points.

[0404] In some examples, the candidate set includes 1 sample, 2 samples, 4 samples, 8 samples, 16 samples, 32 samples, 64 samples, 128 samples, or other 2 samples. N Sample points.

[0405] In some examples, the candidate set includes 3 samples, 6 samples, 12 samples, 24 samples, or other integer samples.

[0406] In some examples, the candidate set is {1 / 4 sample, 1 / 2 sample, 1 sample, 2 sample, 4 sample, 8 sample, 16 sample, 32 sample}.

[0407] In some examples, the candidate set is {1 / 4 sample, 1 / 2 sample, 1 sample, 2 sample, 3 sample, 4 sample, 6 sample, 8 sample, 16 sample}.

[0408] In some examples, the candidate set is set to be equal to the candidate set of blocks used for MMVD encoding and decoding in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

[0409] In some examples, the candidates include more candidates than those for blocks used for MMVD encoding and decoding in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

[0410] In some examples, at least one candidate in the candidate set is different from those candidates used for MMVD encoding and decoding in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

[0411] In some examples, at least one candidate in the candidate set is equal to one of the candidates for a block used for MMVD encoding / decoding in the same video processing unit, wherein the video processing unit includes at least one or a sequence of strips, slices, sub-pictures, pictures.

[0412] In some examples, instead of directly notifying D, it is the index of the MVD magnitude selected from the signaling notification candidate set.

[0413] In some examples, multiple candidate sets of MVD are predefined, and one candidate set from these multiple candidate sets is selected for encoding / decoding the current GMVD encoding / decoding block.

[0414] In some examples, the selection depends on at least one signaling notification at the strip level, picture level, or sequence level.

[0415] In some examples, at least one signaling notification message is included in the strip header, picture header, picture parameter set (PPS), or sequence parameter set (SPS).

[0416] In some examples, the selection depends on the same message used for the MMVD codec, which is SPS_fpel_mmvd_enabled_flag.

[0417] In some examples, D is binarized into a fixed-length codec, a unary codec, or an exponential Columbus codec.

[0418] In some examples, at least one binarized bin of the context codec D in arithmetic codec.

[0419] In some examples, at least one binarized bin of the codec D is bypassed in arithmetic codecs.

[0420] In some examples, the first bin of context codec D is used in arithmetic codec.

[0421] In some examples, other bins of the codec D are bypassed in arithmetic codecs.

[0422] In some examples, the MVD derived from D is further modified before the final MVD for export segmentation.

[0423] In some examples, D is modified to D = D << S, where S is an integer.

[0424] In certain examples, S = 2.

[0425] In some examples, it is implicitly inferred whether to apply the modification.

[0426] In some examples, if the width and / or height of the current picture is greater than a threshold, the modification is applied.

[0427] In some examples, whether to apply the modification depends on a signaling message at at least one of the sequence level, picture level, slice level, sub - picture level, tile level, CTU row level, or CTU level.

[0428] In some examples, the signaling message is in the SPS and / or sequence header, in the PPS and / or picture header, or in the slice header.

[0429] In some examples, the same message including sps_fpel_mmvd_enabled_flag and / or pic_fpel_mmvd_enabled_flag is used to control both MMVD and GMVD.

[0430] In some examples, a separate message is signaled to control MMVD and GMVD separately.

[0431] In some examples, the current video block is split into two segments by a triangular partitioning mode or a geometric merge mode, and two MVDs are signaled or exported for the two segments.

[0432] In some examples, the split MV is calculated as the sum of the MV derived from the split triangular partitioning mode or geometric merge mode and the split MVD.

[0433] In some examples, the weighted summation process of the two MVs of the two splits and two motion compensations is performed in the same way as the triangular partitioning mode or geometric merge mode.

[0434] In some examples, the MV storage process of the two MVs of the two splits is performed in the same way as the triangular partitioning mode or geometric merge mode.

[0435] In some examples, whether to signal or export multiple MVDs depends on one or more conditions.

[0436] In some examples, whether the proposed multiple MVDs are signaled or derived depends on whether the geometric merge mode is enabled for the current video block.

[0437] In some examples, if the multiple MVDs are not signaled or derived, they are inferred to be zero.

[0438] In some examples, whether GMVD is enabled is signaled at least at one of sequence level, picture level, slice level, sub-picture level, tile level, CTU row level or CTU level.

[0439] In some examples, whether GMVD is enabled is signaled in the SPS and / or sequence header, in the PPS and / or picture header, or in the slice header.

[0440] In some examples, whether GMVD is enabled is signaled on condition that the triangle partitioning mode or the geometric merge mode is enabled.

[0441] In some examples, whether the multiple MVDs are signaled or derived depends on the block width (W) and / or block height (H) of the current video block.

[0442] In certain examples, the multiple MVDs are not signaled or derived if at least one or any combination of the following conditions is satisfied:

[0443] i. W >= T1, where T1 = 64;

[0444] ii. H >= T2, where T2 = 64;

[0445] iii. W >= T3 * H, where T3 = 4;

[0446] iv. H >= T4 * H, where T4 = 8;

[0447] v. W <= T5, where T5 = 8;

[0448] vi. H <= T6, where T6 = 8;

[0449] vii. W * H >= T7, where T7 = 2048;

[0450] viii. W * H <= T8, where T8 = 64;

[0451] ix. W == T9 or H == T10, where T9 = T10 = 4;

[0452] x. W / H > T11 or max(W, H) / min(W, H) >= T11;

[0453] xi. W / H < T11 or max(W, H) / min(W, H) < T11.

[0454] In some examples, whether signaling notifications or exporting multiple MVDs depends on the index of the merge candidate.

[0455] In some examples, the index of the merge candidate is greater than the predetermined value.

[0456] In some examples, for blocks encoded and decoded by GMVD, signaling is used to notify MVD before merging candidate indices.

[0457] In some examples, signaling notifications are used to merge candidate indexes based on whether signaling notifications are sent or whether multiple MVDs are exported.

[0458] In some examples, how the block signaling notification for GMVD encoding and decoding informs the index of merge candidates depends on the use of GMVD.

[0459] In some examples, if signaling notification or export of multiple MVDs is performed, the maximum index of merge candidates that can be signaled is reduced.

[0460] In some examples, the conversion involves encoding the current video block into a bitstream.

[0461] In some examples, the conversion involves decoding the current video chunk from the bitstream.

[0462] In some examples, the transformation includes generating a bitstream from the current block; the method also includes storing the bitstream in a non-transitory computer-readable recording medium.

[0463] In some examples, an apparatus for processing video data includes a processor and a non-transitory memory thereon with instructions, wherein the instructions, when executed by the processor, cause the processor to: determine, for a conversion between a current video block and a bitstream of the current video, encode and decode the current video block using a geometric segmentation mode; derive at least one refined MV for the current video block by adding at least one of a plurality of motion vector differences (MVDs) signaled or derived for the current video block to a motion vector (MV) derived from a merge candidate associated with the current video; and perform a conversion based on the refined MV.

[0464] In some examples, a non-transitory computer-readable medium stores instructions that cause a processor to: determine, for a conversion between a current video block and a bitstream of the current video, encode and decode the current video block using a geometric segmentation mode; derive at least one refined MV for the current video block by adding at least one of a plurality of motion vector differences (MVDs) signaled or derived for the current video block to a motion vector (MV) derived from a merge candidate associated with the current video; and perform a conversion based on the refined MV.

[0465] In some examples, a non-transitory computer-readable medium stores a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: determining, for a current video block of video and the bitstream of the current video, to encode and decode the current video block using a geometric segmentation mode; deriving at least one refined MV for the current video block by adding at least one of a plurality of motion vector differences (MVDs) signaled or derived for the current video block to motion vectors (MVs) derived from merge candidates associated with the current video; and generating a bitstream from the current video block based on the refined MV.

[0466] Figure 20 A flowchart of an example method for storing a bitstream of video is shown. The method includes a conversion between a current video block and a bitstream of the current video, determining (2002) encoding and decoding the current video block using a geometric segmentation mode; deriving at least one refined MV for the current video block by adding at least one of a plurality of motion vector differences (MVDs) signaled or derived for the current video block to a motion vector (MV) derived from a merge candidate associated with the current video block; generating (2006) a bitstream from the current video block based on the refined MV; and storing (2008) the bitstream in a non-transitory computer-readable recording medium.

[0467] In this document, the term "video processing" can refer to video encoding, video decoding, video compression, or video decompression. For example, a video compression algorithm can be applied during the conversion from a pixel representation of a video to a corresponding bitstream representation, and vice versa. For example, the bitstream representation of the current video block can correspond to bits that are in the same position or distributed at different positions within the bitstream, as defined by the syntax. For example, macroblocks can be encoded based on the error residuals from the transformation and encoding / decoding, and also using bits in the header and other fields in the bitstream.

[0468] The disclosed and other solutions, examples, embodiments, modules, and functional operations described herein can be implemented in digital electronic circuits or in computer software, firmware, or hardware, including the structures disclosed herein and their structural equivalents, or combinations thereof. The disclosed and other embodiments can be implemented as one or more computer program products, i.e., modules of one or more computer program instructions encoded on a computer-readable medium for execution by a data processing apparatus or for controlling the operation of the data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a memory device, a material composition affecting machine-readable propagated signals, or a combination thereof. The term "data processing apparatus" encompasses all means, devices, and machines for processing data, including, for example, a programmable processor, a computer, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for a computer program, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination thereof. The propagated signals are artificially generated signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device.

[0469] Computer programs (also known as programs, software, software applications, scripts, or code) can be written in any programming language (including compiled or interpreted languages) and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinating files (e.g., a file storing one or more modules, subroutines, or portions of code). Computer programs can be deployed to execute on one or more computers located at a single site or distributed across multiple sites interconnected by a communication network.

[0470] The processing and logic flows described herein can be executed by one or more programmable processors that execute one or more computer programs to perform functions by manipulating input data and generating outputs. The processing and logic flows can also be executed by special-purpose logic circuitry, and the devices can be implemented as special-purpose logic circuitry, such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits).

[0471] For example, processors suitable for executing computer programs include general-purpose and special-purpose microprocessors, as well as any one or more processors in any type of digital computer. Typically, the processor receives instructions and data from read-only memory or random access memory, or both. The basic components of a computer are a processor that executes instructions and one or more storage devices that store the instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or for receiving data from or transferring data to one or more mass storage devices, or both, through operative coupling. However, a computer does not necessarily have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by or incorporated into special-purpose logic circuitry.

[0472] While this patent document contains numerous details, it should not be construed as limiting any subject matter or scope of the claims, but rather as a description of features of specific embodiments of a particular technology. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various functions described in the context of a single embodiment may also be implemented individually in multiple embodiments, or in any suitable sub-combination. Furthermore, although the foregoing features may be described as functioning in certain combinations, or even initially claimed to be so, in some cases one or more features from the claimed combination may be removed from the combination, and the claimed combination may refer to a sub-combination or a variation of the sub-combination.

[0473] Similarly, although the operations are described in a specific order in the accompanying drawings, this should not be construed as requiring the specific order or sequence of operations shown, or all the described operations, to be performed in order to obtain the desired result. Furthermore, the separation of various system components in the embodiments described in this patent document should not be construed as requiring such separation in all embodiments.

[0474] Only some implementations and examples are described; other implementations, enhancements, and variations can be made based on the content described and illustrated in this patent document.

Claims

1. A method for video processing, comprising: For the conversion between the current video block and the bitstream of the video, determine to encode and decode the current video block using the first prediction mode; as well as The conversion is performed based on the determination. In the first prediction mode, first motion information for the first geometric segmentation of the current video block and second motion information for the second geometric segmentation of the current video block are determined, and a weighting process for generating a final prediction for the samples in the weighted region of the current video block is applied based on the weighted sum of the prediction samples, wherein the prediction samples are derived based on the first motion information and the second motion information. Wherein, at least one motion vector in the first motion information and the second motion information is refined by adding at least one motion vector difference, and The motion vector difference is derived based on step size information indicating the absolute value and direction information indicating the direction of the motion vector difference.

2. The method of claim 1, wherein at least one first flag is signaled in the bit stream to indicate whether the at least one motion vector is refined by adding at least one motion vector difference.

3. The method of claim 2, wherein each of the at least one first mark corresponds to a motion vector difference.

4. The method of claim 2, wherein whether the motion vector difference is derived is based on the value of the first flag.

5. The method of claim 1, wherein whether the step size information and the direction information are included in the bitstream is based on the value of a first flag indicating whether the motion vector is refined by adding the motion vector difference.

6. The method of claim 5, wherein when the value of the first flag is not equal to 0, the step size information and the direction information are included in the bit stream.

7. The method according to claim 1, wherein the absolute value is determined based on a table representing the correspondence between step size information and absolute value and the step size information of the current video block.

8. The method of claim 1, wherein the first prediction mode comprises a plurality of segmentation schemes, and at least one segmentation scheme divides the current video block into two or more segments, at least one of the two or more segments being non-square and non-rectangular.

9. The method of claim 1, wherein geometric segmentation includes a triangular segmentation pattern.

10. The method of claim 1, wherein geometric segmentation includes a geometric merge pattern.

11. The method according to any one of claims 1-10, wherein the conversion comprises encoding the current video block into the bitstream.

12. The method according to any one of claims 1-10, wherein the conversion comprises decoding the current video block from the bitstream.

13. An apparatus for encoding and decoding video data, comprising a processor and a non-transitory memory, the non-transitory memory having instructions, wherein the instructions, when executed by the processor, cause the processor to: For the conversion between the current video block and the bitstream of the video, determine to encode and decode the current video block using a first prediction mode; and The conversion is performed based on the determination. in, In the first prediction mode, first motion information for a first geometric segmentation of the current video block and second motion information for a second geometric segmentation of the current video block are determined, and a weighting process for generating a final prediction for samples within a weighted region of the current video block is applied based on a weighted sum of the predicted samples, the predicted samples being derived based on the first motion information and the second motion information. Wherein, at least one motion vector in the first motion information and the second motion information is refined by adding at least one motion vector difference, and The motion vector difference is derived based on step size information indicating the absolute value and direction information indicating the direction of the motion vector difference.

14. The apparatus of claim 13, wherein at least one first flag is signaled in the bit stream to indicate whether the at least one motion vector is refined by adding at least one motion vector difference.

15. The apparatus of claim 14, wherein each of the at least one first mark corresponds to a motion vector difference.

16. The apparatus of claim 14, wherein whether the motion vector difference is derived is based on the value of the first flag.

17. A non-transitory computer-readable medium storing instructions that cause a processor to: For the conversion between the current video block and the bitstream of the video, determine to encode and decode the current video block using a first prediction mode; and The conversion is performed based on the determination. in, In the first prediction mode, first motion information for a first geometric segmentation of the current video block and second motion information for a second geometric segmentation of the current video block are determined, and a weighting process for generating a final prediction for samples within a weighted region of the current video block is applied based on a weighted sum of the predicted samples, the predicted samples being derived based on the first motion information and the second motion information. Wherein, at least one motion vector in the first motion information and the second motion information is refined by adding at least one motion vector difference, and The motion vector difference is derived based on step size information indicating the absolute value and direction information indicating the direction of the motion vector difference.

18. The non-transitory computer-readable medium of claim 17, wherein signaling in the bit stream notifies at least one first flag to indicate whether the at least one motion vector is refined by adding at least one motion vector difference.

19. The non-transitory computer-readable medium of claim 18, wherein each of the at least one first mark corresponds to a motion vector difference.

20. A non-transitory computer-readable medium storing a bitstream of video generated by a method performed by a video processing apparatus, wherein the method includes: For the current video block of the video, determine to encode and decode the current video block using the first prediction mode; as well as The bit stream is generated based on the determination. In the first prediction mode, first motion information for the first geometric segmentation of the current video block and second motion information for the second geometric segmentation of the current video block are determined, and a weighting process for generating a final prediction for the samples in the weighted region of the current video block is applied based on the weighted sum of the prediction samples, wherein the prediction samples are derived based on the first motion information and the second motion information. Wherein, at least one motion vector in the first motion information and the second motion information is refined by adding at least one motion vector difference, and The motion vector difference is derived based on step size information indicating the absolute value and direction information indicating the direction of the motion vector difference.

21. A method for storing a video bitstream, comprising: For the current video block of the video, determine to encode and decode the current video block using the first prediction mode; The bit stream is generated based on the determination; as well as The bitstream is stored in a non-transitory computer-readable recording medium. In the first prediction mode, first motion information for the first geometric segmentation of the current video block and second motion information for the second geometric segmentation of the current video block are determined, and a weighting process for generating a final prediction for the samples in the weighted region of the current video block is applied based on the weighted sum of the prediction samples, wherein the prediction samples are derived based on the first motion information and the second motion information. Wherein, at least one motion vector in the first motion information and the second motion information is refined by adding at least one motion vector difference, and The motion vector difference is derived based on step size information indicating the absolute value and direction information indicating the direction of the motion vector difference.

22. The method according to claim 1, further comprising: For the conversion between the current video block and the current video bitstream, determine to encode and decode the current video block using a geometric segmentation mode; At least one refined MV is derived for the current video block by adding at least one of the multiple motion vector differences (MVDs) that are signaled or derived for the current video block to the motion vector (MV) derived from the merge candidate associated with the current video. as well as The transformation is performed based on the refined MV.

23. The method of claim 22, wherein the current video block is a block encoded with Geometric Motion Vector Difference (GMVD) or a block encoded with Merge (MMVD) using Motion Vector Difference.

24. The method of claim 22, wherein the geometric segmentation mode comprises a plurality of segmentation schemes, and at least one segmentation scheme divides the current video block into two or more segments, at least one of the two or more segments being non-square and non-rectangular.

25. The method of claim 22, wherein the geometric segmentation mode includes a triangle segmentation mode.

26. The method of claim 22, wherein the geometric segmentation mode includes a geometric merge mode.

27. The method of claim 22, wherein the current video block is divided into N segments, and the segmentation signaling of the current video block is notified or derives up to N MVDs, where N>=2.

28. The method of claim 27, wherein N=2.

29. The method of claim 27, wherein the segmentation signaling for the current video block is notified or N MVDs are derived, and each MVD corresponds to a specific segmentation.

30. The method of claim 29, wherein a segmented MVD is added to the segmented MV to obtain a refined MV of the segment for use in motion compensation of the segment.

31. The method of claim 27, wherein signaling notifies or derives K MVDs, and at least one MVD corresponds to more than one segmentation, wherein K <N。 32. The method of claim 27, wherein the number of MVDs to be signaled depends on the decoded information, the decoded information including at least one of the block size associated with the N segments, a low-latency check flag, a predicted direction, or a list of reference images.

33. The method of claim 32, wherein the bitstream includes the number of MVDs.

34. The method of claim 27, wherein a plurality of MVDs are used in a single partition.

35. The method of claim 34, wherein a plurality of MVDs are added to the segmented MV to obtain a refined MV of the segment.

36. The method of claim 22, wherein the bit stream includes a first message indicating whether at least one of the plurality of MVDs is not equal to zero.

37. The method of claim 36, wherein the first message is a flag.

38. The method of claim 36, wherein the first message is either context-coded or bypass-coded in arithmetic encoding / decoding.

39. The method of claim 38, wherein the first message is encoded and decoded using at least one context that is the same as the syntax element indicating whether MMVD is applied.

40. The method of claim 36, wherein information relating to the plurality of MVDs is included in the bitstream only when the first message indicates that at least one of the plurality of MVDs is not equal to zero.

41. The method of claim 22, wherein the bitstream includes a second message for at least one of the plurality of MVDs to indicate whether the MVD is equal to zero, wherein the second message is a non-zero message.

42. The method of claim 41, wherein at least one binarized bin of the second message is context-coded in arithmetic encoding / decoding.

43. The method of claim 41, wherein at least one binarized bin of the second message is bypassed during arithmetic encoding and decoding.

44. The method of claim 41, wherein the second message is encoded and decoded using at least one context that is the same as the syntax element indicating whether MMVD is applied.

45. The method of claim 41, wherein information related to the MVD is included in the bitstream only when the second message indicates that the MVD is not equal to 0.

46. ​​The method of claim 22, wherein at least one of the plurality of MVDs is included in a bitstream having syntax elements indicating the horizontal component of the MVD, or at least one of the plurality of MVDs is included in a bitstream having syntax elements indicating the vertical component of the MVD.

47. The method of claim 22, wherein at least one of the plurality of MVDs is included in a bitstream having syntax elements indicating the absolute values ​​of the horizontal and / or vertical components of the MVD.

48. The method of claim 22, wherein the second MVD is derived based on the first MVD that is signaled or derived prior to signaling notification or deriving the second MVD.

49. The method of claim 23, wherein the MVD used in GMVD is represented by two variables, the first variable being the direction and the second variable being the magnitude represented by D.

50. The method of claim 49, wherein at least one of the plurality of MVDs is included in a bitstream having syntax elements indicating the direction of the MVD.

51. The method of claim 50, wherein the direction comprises at least one of the following: 1) MVD adopts the form (1, 0) * D, where D is not less than 0; 2) MVD takes the form (-1, 0) * D, where D is not less than 0; 3) MVD adopts the form (0, 1) * D, where D is not less than 0; 4) MVD takes the form (0, -1)*D, where D is not less than 0; 5) MVD adopts the form (1, 1) * D, where D is not less than 0; 6) MVD takes the form (-1, 1) * D, where D is not less than 0; 7) MVD takes the form (1, -1)*D, where D is not less than 0; and 8) MVD takes the form (-1, -1)*D, where D is not less than 0.

52. The method of claim 50, wherein the direction comprises at least one of the following: 1) MVD adopts the form (1, 0) * D, where D is not less than 0; 2) MVD takes the form (-1, 0) * D, where D is not less than 0; 3) MVD adopts the form (0, 1) * D, where D is not less than 0; and 4) MVD takes the form (0, -1)*D, where D is not less than 0.

53. The method of claim 50, wherein the direction is binarized using a fixed-length codec, a unary codec, or an exponential Columbus codec.

54. The method of claim 49, wherein D is greater than 0.

55. The method of claim 49, wherein for a block encoded and decoded by GMVD, the indication of the MVD magnitude is represented by a D notified by signaling or derived.

56. The method of claim 55, wherein D is restricted to the candidate set.

57. The method of claim 56, wherein the candidate set includes zero.

58. The method of claim 56, wherein the candidate set comprises 1 / 4 sample points, 1 / 2 sample points, or other fractional sample points.

59. The method of claim 56, wherein the candidate set comprises 1 sample point, 2 sample points, 4 sample points, 8 sample points, 16 sample points, 32 sample points, 64 sample points, 128 sample points, or other 2N sample points.

60. The method of claim 56, wherein the candidate set comprises 3 samples, 6 samples, 12 samples, 24 samples, or other integer samples.

61. The method of claim 56, wherein the candidate set is {1 / 4 sample point, 1 / 2 sample point, 1 sample point, 2 sample points, 4 sample points, 8 sample points, 16 sample points, 32 sample points}.

62. The method of claim 56, wherein the candidate set is {1 / 4 sample point, 1 / 2 sample point, 1 sample point, 2 sample points, 3 sample points, 4 sample points, 6 sample points, 8 sample points, 16 sample points}.

63. The method of claim 56, wherein the candidate set is set to be equal to the candidate set of blocks used for MMVD encoding and decoding in the same video processing unit, wherein the video processing unit includes at least one of strip, slice, sub-picture, picture, or sequence.

64. The method of claim 56, wherein the candidates include additional candidates besides those for blocks used for MMVD encoding / decoding in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

65. The method of claim 56, wherein at least one candidate in the candidate set is different from those candidates for blocks used for MMVD encoding / decoding in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

66. The method of claim 56, wherein at least one candidate in the candidate set is equal to one of the candidates of blocks used for MMVD encoding / decoding in the same video processing unit, wherein the video processing unit includes at least one of strips, slices, sub-pictures, pictures, or sequences.

67. The method of claim 56, wherein instead of directly signaling D, the index of the selected MVD magnitude in the candidate set is signaled.

68. The method of claim 55, wherein a plurality of candidate sets of MVDs are predefined, and one of the plurality of candidate sets is selected for encoding / decoding a block of the current GMVD encoding / decoding.

69. The method of claim 68, wherein the selection depends on at least one signaling notification message at the strip level, picture level, or sequence level.

70. The method of claim 69, wherein the message is signaled in at least one of a strip header, a picture header, a picture parameter set (PPS), or a sequence parameter set (SPS).

71. The method according to claim 68, wherein the selection depends on the same message for the MMVD coding and decoding block, and the same message is sps_fpel_mmvd_enabled_flag.

72. The method according to claim 55, wherein D is binarized into fixed-length coding and decoding, or unary coding or exponential Golomb coding.

73. The method according to claim 55, wherein at least one binarized bin of D is context-coded in arithmetic coding and decoding.

74. The method according to claim 55, wherein at least one binarized bin of D is bypass-coded in arithmetic coding and decoding.

75. The method according to claim 55, wherein the first bin of D is context-coded in arithmetic coding, and the other bins of D are bypass-coded in arithmetic coding and decoding.

76. The method according to claim 55, wherein the MVD derived from D is further modified before deriving the final MVD for the partition.

77. The method according to claim 76, wherein D is modified to D = D << S, where S is an integer.

78. The method according to claim 77, wherein S = 2.

79. The method according to claim 77, wherein it is implicitly inferred whether to apply the modification.

80. The method according to claim 79, wherein if the width and / or height of the current picture is greater than a threshold, the modification is applied.

81. The method according to claim 78, wherein whether to apply the modification depends on at least one signaling message at the sequence level, picture level, slice level, sub-picture level, slice level, CTU row level or CTU level.

82. The method according to claim 81, wherein the message is signaled in the SPS and / or sequence header, in the PPS and / or picture header or in a slice header.

83. The method according to claim 82, wherein the same message including sps_fpel_mmvd_enabled_flag and / or pic_fpel_mmvd_enabled_flag is used to control both MMVD and GMVD.

84. The method according to claim 82, wherein a separate message is signaled to control MMVD and GMVD separately.

85. The method according to claim 22, wherein the current video block is divided into two partitions by a triangular partitioning mode or a geometric merge mode, and two MVDs are signaled or derived for the two partitions.

86. The method according to claim 85, wherein the MV of one partition is calculated as the sum of the MV of the partition derived from the triangular partitioning mode or the geometric merge mode and the MVD of the partition.

87. The method according to claim 85, wherein the weighted summation process of the two MVs and two motion compensations for the two partitions is performed in the same manner as the triangular partitioning mode or the geometric merge mode.

88. The method according to claim 85, wherein the MV storage process for the two MVs of the two partitions is performed in the same manner as the triangular partitioning mode or the geometric merge mode.

89. The method according to claim 22, wherein whether to signal or derive the plurality of MVDs depends on one or more conditions.

90. The method according to claim 89, wherein whether to signal or derive the proposed plurality of MVDs depends on whether the geometric merge mode is enabled for the current video block.

91. The method according to claim 89, wherein if the plurality of MVDs are not signaled or derived, they are inferred to be zero.

92. The method according to claim 89, wherein whether GMVD is enabled is signaled at least at one of sequence level, picture level, slice level, sub-picture level, tile level, CTU row level or CTU level.

93. The method according to claim 92, wherein whether GMVD is enabled is signaled in the SPS and / or sequence header, in the PPS and / or picture header or in the slice header.

94. The method according to claim 93, wherein whether GMVD is enabled is signaled under the condition that the triangular partitioning mode or the geometric merge mode is enabled.

95. The method according to claim 89, wherein whether to signal or derive the plurality of MVDs depends on the block width (W) and / or block height (H) of the current video block.

96. The method according to claim 95, wherein if at least one or any combination of the following conditions is satisfied, the plurality of MVDs are not signaled or derived: i. W >= T1, T1 = 64; ii. H >= T2, T2 = 64; iii. W >= T3 * H, T3 = 4; iv. H >= T4 * H, T4 = 8; v. W <= T5, T5 = 8; vi. H <= T6, T6 = 8; vii. W * H >= T7, T7 = 2048; viii. W * H <= T8, T8 = 64; ix. W == T9 or H == T10, T9 = T10 = 4; x. W / H > T11 or max(W, H) / min(W, H) >= T11; xi. W / H < T11 or max(W, H) / min(W, H) < T11.

97. The method according to claim 89, wherein whether to signal or derive the plurality of MVDs depends on the index of the merge candidate.

98. The method according to claim 97, wherein the index of the merge candidate is greater than a predetermined value.

99. The method according to claim 23, wherein for the block encoded and decoded by GMVD, the MVD is signaled before the index of the merge candidate.

100. The method according to claim 99, wherein the index of the merge candidate is signaled according to whether the plurality of MVDs are signaled or derived.

101. The method of claim 99, wherein how the block signaling for GMVD encoding / decoding notifies the merge candidate of the index depends on the use of GMVD.

102. The method of claim 101, wherein if the plurality of MVDs are signaled or derived, the maximum index of the merge candidates that can be signaled is reduced.

103. The method according to any one of claims 22-102, wherein the conversion comprises encoding the current video block into the bitstream.

104. The method according to any one of claims 22-102, wherein the conversion includes decoding the current video block from the bitstream.

105. The method according to any one of claims 22-102, wherein the conversion comprises generating the bitstream from the current block; The method further includes: The bit stream is stored in a non-transitory computer-readable recording medium.

106. An apparatus for processing video data, comprising a processor and a non-transitory memory having instructions thereon, wherein the instructions, when executed by the processor, cause the processor to perform the method according to any one of claims 1-12 and 21-105.

107. A non-transitory computer-readable medium storing instructions that cause a processor to perform the method according to any one of claims 1-12 and 21-105.

108. A non-transitory computer-readable medium storing a bitstream of video generated by a video processing apparatus performing the method according to any one of claims 1-12 and 21-105.

109. The method according to claim 1, further comprising: For the conversion between the current video block and the current video bitstream, determine to encode and decode the current video block using a geometric segmentation mode; At least one refined MV is derived for the current video block by adding at least one of the multiple motion vector differences (MVDs) that are signaled or derived for the current video block to the motion vectors (MVs) derived from the merge candidate associated with the current video. The bitstream is generated from the current video block based on the refined MV; as well as The bit stream is stored in a non-transitory computer-readable recording medium.