Geometry Partition Merge Mode with Merge Mode Having Motion Vector Difference

By providing additional motion vector refinement and new syntax elements for GEO segmentation in VVC, GEO's limitations in MV accuracy are solved, encoding efficiency is improved, and video encoding is achieved more efficiently.

CN115486069BActive Publication Date: 2025-07-04NOKIA TECHNOLOGIES OY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202180026750.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-04-03
Filing Date
2021-03-25
Publication Date
2025-07-04
Estimated Expiration
2041-03-25

AI Technical Summary

Technical Problem

In multifunction video encoding (VVC), geometric segmentation merge mode (GEO) has limitations in motion vector (MV) accuracy, hindering higher coding efficiency, especially since GEO segmentation supports 64 angles and distances, but only supports 30 MV combinations selected from the 6 merge candidates.

Method used

By providing additional motion vector refinement for GEO segmentation, using new syntax elements and MMVD modes, the motion vector difference value is calculated to improve prediction accuracy, allowing more flexible MV refinement, and combining the existing VVC GEO syntax mechanism, indicating whether additional refinement is performed and motion vector difference value is applied according to the enabled value.

Benefits of technology

It improves the prediction accuracy of the geometric segmentation merge mode, improves coding efficiency, reduces the cost of coding complexity and bit overhead, and achieves more efficient video encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115486069B_ABST
    Figure CN115486069B_ABST
Patent Text Reader

Abstract

A method, apparatus, and computer program product for improving the prediction accuracy of geometric segmentation merge modes. In the context of a method, the method accesses a coding unit that is segmented into a first partition and a second partition. The method also determines whether an enable value associated with the coding unit indicates additional motion vector refinement, and determines at least one motion vector difference for at least one of the first partition and the second partition based on determining that the enable value associated with the coding unit indicates additional motion vector refinement. The method also applies prediction of the coding unit based on the at least one motion vector difference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments generally relate to techniques in video coding, and more particularly to techniques for improving the prediction accuracy of geometric segmentation merge modes. Background Art

[0002] Hybrid video codecs can encode video information in several stages. In one stage, pixel values in a predicted picture (e.g., a block) are predicted. This prediction can be performed by means of motion compensation (e.g., finding and indicating in a previously encoded video frame a region closely corresponding to the block being encoded) or by spatial means (e.g., using pixel values around the block to be encoded in a specified manner). Prediction using motion compensation means can be referred to as inter-frame prediction, while prediction using spatial means can be referred to as intra-frame prediction.

[0003] Versatile Video Coding (VVC) is an international video coding standard being developed by the Joint Video Exploration Team (JVET). VVC is expected to be the successor to the High Efficiency Video Coding (HEVC) standard. In VVC, many new coding tools are available for encoding video, including many new and improved inter-frame prediction coding tools, such as but not limited to the merge mode with motion vector difference (MMVD).

[0004] In VVC, the Geometric Segmentation Merge mode (GEO) is an inter-frame coding mode that allows for flexible segmentation of coding units. However, under the current VVC design, MMVD cannot be used for GEO coding units. In the VVC design, the GEO mode is more flexible in terms of segmentation and more restrictive in terms of motion vector (MV) accuracy. Specifically, GEO segmentation supports 64 angles and distances, while GEO MV only supports 30 MV combinations selected from 6 merge candidates. The restriction on MV accuracy hinders GEO from achieving higher coding efficiency. Summary of the Invention

[0005] In one embodiment, there is provided an apparatus that includes at least one processor and at least one memory including computer program code, the at least one memory and the computer program code being configured to, with the at least one processor, cause the apparatus to: access a coding unit that is divided into a first partition and a second partition. The at least one memory and the computer program code are further configured to, with the at least one processor, cause the apparatus to: determine whether an enable value associated with the coding unit indicates additional motion vector refinement, and based on determining that the enable value associated with the coding unit indicates additional motion vector refinement, determine at least one motion vector difference for at least one of the first partition and the second partition. The at least one memory and the computer program code are configured to, with the at least one processor, cause the apparatus to: apply prediction of the coding unit based on the at least one motion vector difference.

[0006] In some embodiments of the apparatus, the coding unit is geometrically partitioned according to a geometric split merge pattern. In some embodiments of the apparatus, at least one motion vector difference is at least based on a distance value and a direction value associated with a corresponding partition. In some embodiments of the apparatus, applying prediction of the coding unit according to at least one motion vector difference includes applying the same motion vector difference to both a first partition and a second partition. In some embodiments of the apparatus, applying prediction of the coding unit according to at least one motion vector difference includes applying the motion vector difference only to the larger one of the first partition or the second partition. In some embodiments of the apparatus, determining at least one motion vector difference for at least one of the first partition and the second partition is based on corresponding enable values associated with the first partition and the second partition. In some embodiments of the apparatus, when the corresponding enable value of a particular partition indicates that the motion vector difference will not be used, no at least one motion vector difference is determined for the particular partition.

[0007] In another embodiment, an apparatus is provided that includes means for accessing a coding unit partitioned into a first partition and a second partition. The apparatus further includes means for determining whether an enable value associated with the coding unit indicates additional motion vector refinement, and means for determining at least one motion vector difference for at least one of the first partition and the second partition according to the determination that the enable value associated with the coding unit indicates additional motion vector refinement. The apparatus further includes means for applying prediction of the coding unit according to at least one motion vector difference.

[0008] In some embodiments of the apparatus, the coding unit is geometrically partitioned according to a geometric split merge pattern. In some embodiments of the apparatus, at least one motion vector difference is at least based on a distance value and a direction value associated with a corresponding partition. In some embodiments of the apparatus, the means for applying prediction of the coding unit according to at least one motion vector difference includes means for applying the same motion vector difference to both a first partition and a second partition. In some embodiments of the apparatus, the means for applying prediction of the coding unit according to at least one motion vector difference includes means for applying the motion vector difference only to the larger one of the first partition or the second partition. In some embodiments of the apparatus, determining at least one motion vector difference for at least one of the first partition and the second partition is based on corresponding enable values associated with the first partition and the second partition. In some embodiments of the apparatus, when the corresponding enable value of a particular partition indicates that the motion vector difference will not be used, no at least one motion vector difference is determined for the particular partition.

[0009] In another embodiment, a method is provided that includes accessing a coding unit that is partitioned into a first partition and a second partition. The method further includes: determining whether an enable value associated with the coding unit indicates additional motion vector refinement, and based on determining that the enable value associated with the coding unit indicates additional motion vector refinement, determining at least one motion vector difference for at least one of the first partition and the second partition. The method further includes applying a prediction of the coding unit based on the at least one motion vector difference.

[0010] In some embodiments of the method, the coding unit is geometrically partitioned according to a geometric partition merge mode. In some embodiments of the method, the at least one motion vector difference is at least based on a distance value and a direction value associated with the corresponding partition. In some embodiments of the method, applying the prediction of the coding unit based on the at least one motion vector difference includes applying the same motion vector difference to both the first partition and the second partition. In some embodiments of the method, applying the prediction of the coding unit based on the at least one motion vector difference includes applying the motion vector difference only to the larger one of the first partition or the second partition. In some embodiments of the method, determining at least one motion vector difference for at least one of the first partition and the second partition is based on the respective enable values associated with the first partition and the second partition. In some embodiments of the method, when the respective enable value of a particular partition indicates that the motion vector difference will not be used, at least one motion vector difference is not determined for the particular partition.

[0011] In another embodiment, a computer program product is provided that includes a non-transitory computer-readable storage medium having program code portions stored thereon, the program code portions being configured, when executed, to: access a coding unit that is partitioned into a first partition and a second partition. The program code portions are further configured, when executed, to: determine whether an enable value associated with the coding unit indicates additional motion vector refinement, and based on determining that the enable value associated with the coding unit indicates additional motion vector refinement, the program code portions are further configured, when executed, to: determine at least one motion vector difference for at least one of the first partition and the second partition. The program code portions are further configured, when executed, to: apply a prediction of the coding unit based on the at least one motion vector difference.

[0012] In some embodiments of a computer program product, encoding units are geometrically partitioned according to a geometric partitioning merge mode. In some embodiments of a computer program product, at least one motion vector difference is at least based on a distance value and a direction value associated with a corresponding partition. In some embodiments of a computer program product, applying prediction of an encoding unit according to at least one motion vector difference includes applying the same motion vector difference to both a first partition and a second partition. In some embodiments of a computer program product, applying prediction of an encoding unit according to at least one motion vector difference includes applying the motion vector difference only to the larger one of the first partition or the second partition. In some embodiments of a computer program product, determining at least one motion vector difference for at least one of the first partition and the second partition is based on corresponding enable values associated with the first partition and the second partition. In some embodiments of a computer program product, when the corresponding enable value of a specific partition indicates that a motion vector difference will not be used, at least one motion vector difference is not determined for the specific partition. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Certain exemplary embodiments of the present disclosure have been described generally above, and hereinafter reference will be made to the accompanying drawings, which are not necessarily drawn to scale, and in which:

[0014] Figure 1 is a block diagram of an apparatus that can be specifically configured according to an exemplary embodiment of the present disclosure;

[0015] Figure 2 is a flowchart showing operations performed according to an exemplary embodiment;

[0016] Figure 3A is a representation of a specific semantic change of the VVC specification according to an exemplary embodiment; and

[0017] Figure 3B is a representation of a specific semantic change of the VVC specification according to an exemplary embodiment. DETAILED DESCRIPTION

[0018] Some embodiments of the present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, in which some, but not all, embodiments of the present disclosure are shown. In fact, some embodiments may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy applicable legal requirements. Like reference numerals always refer to like elements. As used herein, the terms "data", "content", "information" and similar terms may be used interchangeably to refer to data that can be transmitted, received, and / or stored according to embodiments of the present disclosure. Thus, the use of any such term should not be regarded as limiting the spirit and scope of the embodiments of the present disclosure.

[0019] In addition, as used herein, the term "circuitry" refers to (a) a pure hardware circuit implementation (e.g., an implementation using analog circuitry and / or digital circuitry); (b) a combination of circuitry and computer program product(s), including software and / or firmware instructions stored on one or more computer-readable memories that work together to cause an apparatus to perform one or more functions described herein; (c) circuitry, such as a microprocessor or a portion of a microprocessor, which also requires software or firmware to operate even if the software or firmware is not physically present. This definition of "circuitry" applies to all uses of the term herein, including in any claims. As another example, as used herein, the term "circuitry" also includes an implementation that includes one or more processors and / or portions thereof, as well as accompanying software and / or firmware. As another example, the term "circuitry" as used herein also includes, for example, a baseband integrated circuit or an application processor integrated circuit for a mobile phone, or a similar integrated circuit in a server, a cellular network device, other network devices (such as core network apparatuses), a field programmable gate array, and / or other computing devices.

[0020] As described above, hybrid video codecs (e.g., ITU-T H.263, H.264 / Advanced Video Coding (AVC), and High Efficiency Video Coding (HEVC)) can encode video information in stages. During a first stage, pixel values in a block are predicted, for example, by means of motion compensation or by spatial means. During the first stage, predictive coding can be applied, for example, as sample or syntax prediction. In sample prediction, pixels or sample values in a predicted block are predicted. For example, these pixel or sample values can be predicted using one or more motion compensation or intra prediction mechanisms.

[0021] A motion compensation mechanism (also referred to as inter prediction, temporal prediction, motion compensated temporal prediction, or motion compensation prediction (MCP)) involves locating and indicating in a previously encoded video frame a region that closely corresponds to the block being encoded. In this regard, inter prediction can reduce temporal redundancy.

[0022] Intra prediction of pixel or sample values by spatial means involves finding and indicating spatial region relationships. Intra prediction takes advantage of the fact that neighboring pixels within the same picture are likely to be correlated. Intra prediction can be performed in the spatial domain or the transform domain (e.g., sample values or transform coefficients can be predicted). Intra prediction is typically used in intra coding, where inter prediction is not applied.

[0023] In syntax prediction, which can also be referred to as parameter prediction, a syntax element and / or a syntax element value and / or a variable derived from a syntax element are predicted from an earlier (decoded) encoded syntax element and / or an earlier derived variable.

[0024] For example, in motion vector prediction, a motion vector (e.g., for inter-frame and / or inter-view prediction) can be differentially encoded relative to a block-specific predicted motion vector. In many video codecs, the predicted motion vector is created in a predefined manner, e.g., by computing the median of the encoded or decoded motion vectors of neighboring blocks. Another method of creating a motion vector prediction (sometimes referred to as Advanced Motion Vector Prediction (AMVP)) is to generate a candidate prediction list from neighboring blocks and / or co-located blocks in a temporal reference picture and signal the candidate that is selected as the motion vector predictor. In addition to predicting the motion vector value, the reference index of a previously encoded / decoded picture can also be predicted. The reference index can be predicted from neighboring blocks and / or co-located blocks in a temporal reference picture. Differential encoding of motion vectors is typically disabled across slice boundaries. For example, block partitioning can be predicted from coding tree units to coding units and prediction units (PUs).

[0025] In filter parameter prediction, filtering parameters such as those for sample adaptive offset can be predicted. As described above, prediction methods that use image information from a previous encoded image are also referred to as inter-frame prediction, temporal prediction, and / or motion compensation. Prediction methods that use image information within the same image are also referred to as intra-frame prediction.

[0026] In a second stage, the prediction error (e.g., the difference between the predicted pixel block and the original pixel block) is encoded. This can be done by transforming the difference in pixel values using a specified transform (e.g., Discrete Cosine Transform (DCT) or a variant of DCT), quantizing the coefficients, and entropy encoding the quantized coefficients. By varying the fidelity of the quantization process, the encoder can control the balance between the accuracy of the pixel representation (e.g., picture quality) and the size of the resulting encoded video representation (e.g., file size of the transmission bitrate).

[0027] In video codecs including H.264 / AVC and HEVC, motion information is indicated by motion vectors associated with each motion-compensated image block. Each of these motion vectors represents the displacement of an image block in a picture to be encoded (at the encoder) or decoded (at the decoder) and a predicted source block in a previously encoded or decoded image (or picture). In H.264 / AVC and HEVC and many other video compression standards, a picture is partitioned into a rectangular grid, and for each in the rectangular grid, a similar block in one of the reference pictures is indicated for inter-frame prediction. The position of the predicted block is encoded as a motion vector indicating the position of the predicted block relative to the block being encoded.

[0028] Versatile Video Coding (VVC) provides many new coding tools. For example, VVC includes new coding tools for intra prediction, including but not limited to 67 intra modes with wide-angle mode extension, block-size and mode-dependent 4-tap interpolation filters, position-dependent intra prediction combination (PDPC), cross-component linear model intra prediction, multi-reference line intra prediction, intra sub-partitioning, and weighted intra prediction with matrix multiplication. In addition, VVC includes new coding tools for inter prediction, such as block motion copy with spatial, temporal, history-based, and pairwise average merge candidates, affine motion inter prediction, sub-block-based temporal MV prediction, adaptive MV resolution, 8×8 block-based motion compression for temporal motion prediction, high-precision motion vector storage and motion compensation with 8-tap interpolation filter for the luminance component and 4-tap interpolation filter for the chrominance component, triangular partitioning, combined intra and inter prediction (CIIP), merge with motion vector difference (MMVD), symmetric MVD coding, bidirectional optical flow, decoder-side motion vector refinement, and bi-prediction with CU-level weights, etc.

[0029] VVC also includes new coding tools for transform, quantization, and coefficient coding, such as multiple primary transform selection with DCT2, DST7, and DCT8, secondary transform in the low-frequency region, sub-block transform of inter prediction residuals, related quantization with the maximum QP increased from 51 to 63, transform coefficient coding with signed data hiding, transform skip residual coding, etc. In addition, VVC includes new coding tools for entropy coding, such as an arithmetic coding engine with adaptive dual-window probability update, loop filter coding tools (such as loop shaping, deblocking filter with a stronger longer filter, sample adaptive offset, and adaptive loop filter). It also includes screen content coding tools, such as current picture reference with reference region limitation, 360-degree video coding tools (such as horizontal wrap-around motion compensation), and advanced syntax and parallel processing tools (such as reference picture management with direct reference picture list signaling), as well as tile groups with rectangular tile groups.

[0030] In VVC segmentation, each picture is divided into coding tree units (CTUs) similar to HEVC. Pictures can also be divided into slices, tiles, bricks, and / or sub-pictures. A CTU can be split into smaller CUs using multi-type segmentation (e.g., quadtree (QT), binary tree (BT), ternary tree (TT), etc.). In addition, there are specific rules to infer the segmentation at the picture boundary, and redundant split patterns are not allowed in nested multi-type segmentation.

[0031] For inter prediction in VVC, the merge list may include candidates such as spatial motion vector prediction (MVP) from spatially adjacent CUs, temporal MVP from co-located CUs, history-based MVP from a first-in-first-out (FIFO) table, pairwise-average MVP (e.g., using candidates already in the merge list), or a zero (0) motion vector.

[0032] After signaling a merge candidate, the merge mode with motion vector difference (MMVD) signals the MVD and the resolution index. In symmetric MVD, in the case of bi-prediction, the motion information of list 1 is derived from the motion information of list 0. In affine prediction, for different corners of a block, several motion vectors are indicated or signaled, which are used to derive the motion vectors of sub-blocks. In affine merge, the affine motion information of a block is generated based on the normal or affine motion information of adjacent blocks. In sub-block-based temporal motion vector prediction, the motion vectors of the sub-blocks of the current block are predicted from the appropriate sub-blocks (if available) indicated by the motion vectors of spatially adjacent blocks in the reference frame. In adaptive motion vector resolution (AMVR), the precision of the MVD is signaled for each CU. In bi-prediction with CU-level weights, the index indicates the weight value of the weighted average of two predicted blocks. Bi-directional optical flow (BDOF) refines the motion vector in the case of bi-prediction. BDOF generates two predicted blocks using the signaled motion vector. Then motion refinement is calculated to minimize the error between the two predicted blocks using their gradient values, and the motion refinement and gradient values are used to refine the final predicted block.

[0033] VVC includes three loop filters, including the deblocking filter and SAO (both in HEVC), and the adaptive loop filter (ALF). The filtering process order in VVC starts with the deblocking filter, then SAO, and then ALF. The SAO in VVC is the same as the SAO in HEVC.

[0034] For the luma component, the ALF selects one of 25 filters for each 4×4 block based on the classification of the direction and activity from local gradients. Additionally, two diamond filter shapes are used.

[0035] The ALF filter parameters are signaled in the adaptive parameter set (APS). In one APS, up to 25 sets of luma filter coefficients and clip-level indices, and up to 8 sets of chroma filter coefficients and clip-level indices can be signaled. To reduce the bit overhead, the filter coefficients of different classifications of the luma component can be merged. In the slice header, the indices of up to seven (7) adaptive parameter sets for the current slice are signaled.

[0036] For the chrominance component, the APS index is signaled in the slice header to indicate the chrominance filter bank used for the current slice. At the CTB level, if there is more than one chrominance filter in the APS, the filter index is signaled for each chrominance CTB. The filter coefficients are quantized with a norm equal to 128. To limit the multiplication complexity, bitstream consistency is applied such that the coefficient values at non-central positions are in the range of -27 to 27-1, including the endpoints. The central position coefficient is not signaled in the bitstream and is considered equal to 128.

[0037] At the decoder side, when ALF is enabled for a CTB, each sample R(i, j) within the CU is filtered to produce a sample value R′(i, j) as follows:

[0038]

[0039] where f(k, l) represents the decoded filter coefficient, K(x, y) is the clipping function, and c(k, l) represents the decoded clipping parameter. The variables k and l vary between -L / 2 and L / 2, where L represents the filter length. The clipping function K(x, y) = min(y, max(-y, x)) corresponds to the function Clip3(-y, y, x).

[0040] The Cross-Component Adaptive Loop Filter (CCALF) uses a linear filter to filter the luma sample values and generates a residual correction for the chrominance channel from the co-located filtered output. The filter is designed to operate in parallel with the existing luma ALF. The CCALF design uses a 3×4 diamond with eight (8) unique coefficients. When the limit is for the chrominance component of a CTU to enable chrominance ALF or CCALF, the per-pixel multiplier limit is 16 (15 for the current ALF). The dynamic range of the filter coefficients is limited to 6-bit signed.

[0041] To be consistent with the existing ALF design, the filter coefficients are signaled in the APS. Up to four (4) filters are supported, and the filter selection is indicated at the CTU level. Symmetric line selection is used at the virtual boundaries to further coordinate with the ALF. Additionally, to limit the storage required for the corrected output, the CCALF residual output is clipped to -2 BitDepthC-1 to 2 BitDepthC-1 -1, including the endpoints.

[0042] In the VVC standard, the Geometric Partitioning Merge (GEO) mode is an inter-frame coding mode that allows a Coding Unit (CU) to be flexibly partitioned into two (2) unequal-sized rectangles or two (2) non-rectangular partitions. Each of the two partitions in a GEO CU has a corresponding unidirectional Motion Vector (MV). These MVs are derived from the merge candidate indices signaled in the bitstream. VVC also supports the Merge Mode with Motion Vector Difference (MMVD) as described above. A CU using MMVD derives its MV from a base MV and a Motion Vector Difference (MVD). The base MV is determined based on one of the first two candidates in the merge list. The MVD is an offset that can be selected from a set of 32 predefined options. However, MMVD is not allowed for GEO CUs.

[0043] In VVC, GEO provides additional flexibility in partitioning but is more restrictive in terms of MV accuracy. In particular, GEO partitioning supports 64 angles and distances, while GEO MVs only support 30 MV combinations selected from six (6) merge candidates. This limitation on MV accuracy hinders GEO from achieving higher coding efficiency.

[0044] Figure 1 An example of an apparatus 100 that can be configured to perform operations in accordance with embodiments described herein is depicted. As Figure 1 shown, the apparatus includes processing circuitry 12, a memory 14, and a communication interface 16, associated with or communicating with each other. The processing circuitry 12 can communicate with the memory via a bus to transfer information between components of the apparatus. The memory can be non-transitory and can include, for example, one or more volatile and / or non-volatile memories. In other words, for example, the memory can be an electronic storage device including gates (e.g., a computer-readable storage medium) configured to store data (e.g., bits) retrievable by a machine (e.g., a computing device such as the processing circuitry). According to example embodiments of the present disclosure, the memory can be configured to store information, data, content, applications, instructions, etc., such that the apparatus can perform various functions. For example, the memory can be configured to buffer input data for processing by the processing circuitry. Additionally or alternatively, the memory can be configured to store instructions for execution by the processing circuitry.

[0045] In some embodiments, the apparatus 100 may be embodied in various computing devices. However, in some embodiments, the apparatus may be embodied as a chip or chipset. In other words, the apparatus may include one or more physical packages (e.g., chips) that include materials, components, and / or wires on a structural component (e.g., a substrate). The structural component may provide physical strength, dimensional retention, and / or electrical interaction confinement for the component circuitry included thereon. Thus, in some cases, the apparatus may be configured to implement embodiments of the present disclosure on a single chip or as a single “system on a chip”. Accordingly, in some cases, a chip or chipset may constitute the means for performing one or more operations to provide the functionality described herein.

[0046] The processing circuitry 12 may be implemented in a variety of different ways. For example, the processing circuitry may be implemented as one or more of various hardware processing components, such as a coprocessor, a microprocessor, a controller, a digital signal processor (DSP), a processing element with or without an accompanying DSP, or various other circuitry, including integrated circuits, such as an ASIC (application specific integrated circuit), an FPGA (field programmable gate array), a microcontroller unit (MCU), a hardware accelerator, a special-purpose computer chip, etc. Thus, in some embodiments, the processing circuitry may include one or more processing cores configured to execute independently. A multi-core processing circuitry may enable multi-processing within a single physical package. Additionally or alternatively, the processing circuitry may include one or more processors configured in series via a bus to enable independent execution of instructions, pipelining, and / or multi-threading.

[0047] In an example embodiment, the processing circuitry 12 may be configured to execute instructions stored in the memory 14 or accessible to the processing circuitry. Alternatively or additionally, the processing circuitry may be configured to perform hard-coded functions. Thus, whether configured by hardware or software methods or by a combination thereof, the processing circuitry can represent an entity (e.g., physically implemented as circuitry) that is capable of performing operations in accordance with embodiments of the present disclosure when correspondingly configured. Thus, for example, when the processing circuitry is implemented as an ASIC, FPGA, etc., the processing circuitry can be specifically configured as hardware for performing the operations described herein. Alternatively, as another example, when the processing circuitry is implemented as an executor of instructions, upon execution of the instructions, the instructions can specifically configure the processor to perform the algorithms and / or operations described herein. However, in some cases, the processing circuitry may be a processor of a particular device (e.g., an image or video processing system) that is configured to further configure the processing circuitry to adopt an embodiment by instructions for performing the algorithms and / or operations described herein. The processing circuitry can particularly include a clock, an arithmetic logic unit (ALU), and logic gates configured to support the operation of the processing circuitry.

[0048] The communication interface 16 can be any component, such as a device or circuitry implemented in hardware or a combination of hardware and software, that is configured to receive and / or transmit data, including media content in the form of video or image files, one or more audio tracks, etc. In this regard, the communication interface can include, for example, an antenna (or antennas) and supporting hardware and / or software to enable communication with a wireless communication network. Additionally or alternatively, the communication interface can include circuitry for interacting with the (one or more) antennas to cause signal transmission via the (one or more) antennas or to process signals received via the (one or more) antennas for reception. In some environments, the communication interface can alternatively or additionally support wired communication. Thus, for example, the communication interface can include a communication modem and / or other hardware / software to support communication via a cable, digital subscriber line (DSL), universal serial bus (USB), or other mechanism.

[0049] According to some embodiments, apparatus 100 may be configured according to an architecture for providing video encoding, decoding, and / or compression. In this regard, apparatus 100 may be configured as a video encoding device. For example, apparatus 300 may be configured to encode video according to one or more video compression standards, such as the VVC specification (e.g., "Versatile Video Coding" by B. Bross, J. Chen, S. Lin, and Y-K. Wang of JVET-Q2001-v13 in January 2020, which is incorporated herein by reference). Although some embodiments herein relate to operations associated with the VVC standard, it should be understood that the processes discussed herein may be used for any video encoding standard.

[0050] The JVET contribution JVET-Q0315 describes a triangular prediction mode with MVD. In particular, JVET-Q0315 proposes to allow MVD with the triangular prediction mode, which is a subset and precursor of GEO. The related proposed merged data syntax of JVET-Q0315 is shown in Table A below, where the enable flag tmvd_part_flag is used to indicate the use of the proposed mode and is always encoded for partition index 0. A second flag, i.e., the enable flag for partition index 1, is sent only when the enable flag for partition index 0 is true. When its enable flag is true, the MVD distance and direction are signaled for the GEO partition.

[0051]

[0052] Table A

[0053] Certain embodiments described herein provide an extension to GEO to allow MMVD-style MV refinement. This extension improves the prediction accuracy of the GEO mode by providing more flexible motion vectors for GEO partitions. However, while the flexibility in MV refinement may come with a cost in terms of overhead bits and implementation complexity, some embodiments described herein utilize a method that minimizes these costs while maximizing the encoding efficiency, as MMVD in VVC has been shown to have low complexity.

[0054] In other words, some embodiments herein combine the VVC GEO usage syntax with new syntax elements for enabling conditions. In particular, when the VVC GEO usage syntax is indicated as true and the new syntax element is indicated as false, VVC GEO without refinement is used, while when both the VVC GEO usage and the new syntax element are indicated as true, GEO with MMVD is used. The VVC GEO syntax mechanism for indicating the MVs of two GEO partitions is reused to determine the base MV. The MVD can be calculated using a mechanism similar to the MMVD syntax, for example, by looking up the MVD from a predetermined lookup table using distance syntax and direction syntax.

[0055] Figure 2 An example operation for extending GEO to allow additional motion vector refinement that can be performed by device 100 is shown. At operation 201, device 100 includes components configured to access an encoding unit that is segmented into a first partition and a second partition, such as processor 12, memory 14, etc. For example, the encoding unit can be accessed as part of an encoding process, such as an encoding process or a decoding process. In some embodiments, the encoding unit is geometrically segmented according to GEO.

[0056] At operation 202, device 100 includes components configured to determine whether an enabling value associated with the encoding unit indicates additional motion vector refinement, such as processor 12, memory 14, etc.

[0057] For example, in one embodiment, an enabling value such as an additional control flag can be used to indicate the use of additional motion vector refinement. As an example, MV refinement with two different offsets can be used (e.g., one offset for each GEO partition). In particular, the flag variable geo_mmvd_flag can be used as a usage flag and combined with the ciip_flag to indicate the use of the proposed GEO with the MMVD mode. Additionally, in this embodiment, four syntax elements can be utilized to provide information for calculating the motion vector difference of the GEO with the MMVD mode. These four syntax elements are geo_mmvd_distance_idx0, geo_mmvd_direction_idx0, geo_mmvd_distance_idx1, geo_mmvd_direction_idx1. Table B below shows these syntax elements.

[0058]

[0059] Table B

[0060] At this point, at operation 203, in accordance with determining that the enabled value associated with the coding unit indicates additional motion vector refinement, apparatus 100 includes components for calculating at least one motion vector difference for at least one of the first partition and the second partition, such as processor 12, memory 14, etc. As described above, the at least one motion vector difference is at least based on the distance value and the direction value associated with the corresponding partition.

[0061] At operation 204, apparatus 100 includes components configured to apply prediction of the coding unit in accordance with the at least one motion vector difference, such as processor 12, memory 14, etc.

[0062] In another embodiment, two syntax elements provide information for calculating the MVD value of a GEO with the MMVD mode. These two syntax elements are geo_mmvd_distance_idx and geo_mmvd_direction_idx. For example, Table C below shows these syntax values.

[0063]

[0064]

[0065] Table C

[0066] For example, in some embodiments, apparatus 300 may always apply the MVD value to a specific partition (e.g., GEO partition index 0). Similarly, in some embodiments, apparatus 300 may always apply the MVD value to GEO partition index 1. In another embodiment, apparatus 300 may apply the MVD value to the GEO partition index based on the GEO partition mode in a predetermined manner, e.g., applying the MVD value to the larger partition in a specific CU. In some embodiments, applying segmentation of a picture in accordance with the at least one motion vector difference includes applying the same motion vector difference to both the first partition and the second partition.

[0067] For example, in the present embodiment, the flag variable geo_mmvd_flag[x0][y0] specifies that the geometric segmentation merge mode with motion vector difference is used to generate the inter-frame prediction parameters of the current coding unit when it is equal to 1. Similarly, geo_mmvd_flag[x0][y0] specifies that the geometric segmentation mode with motion vector difference is not used to generate the inter-frame prediction parameters when it is equal to 0. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration relative to the top-left luminance sample of the picture. In the absence of geo_mmvd_flag[x0][y0], it is inferred to be equal to 0.

[0068] The variable geo_mmvd_distance_idx[x0][y0] specifies the index used to derive GeoMmvdDistance[0][x0][y0], as specified in Table D below. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration relative to the top-left luminance sample of the picture.

[0069]

[0070] Table D

[0071] The variable geo_mmvd_direction_idx[x0][y0] specifies the index used to derive GeoMmvdSign[x0][y0], as specified in Table E below. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coding block under consideration relative to the top-left luminance sample of the picture.

[0072]

[0073] Table E

[0074] The two components of the geometric segmentation merge plus the MVD offset GeoMmvdOffset[x0][y0] are derived as follows:

[0075] GeoMmvdOffset[x0][y0][0] = (GeoMmvdDistance[x0][y0] << 2) * GeoMmvdSign[x0][y0][0]

[0076] GeoMmvdOffset[x0][y0][1] = (GeoMmvdDistance[x0][y0] << 2) * GeoMmvdSign[x0][y0][1]

[0077] In another embodiment, two additional control flag variables can be utilized to indicate the use of the method. In this regard, when the method is enabled, MV refinement with two different offsets can be used (e.g., one offset for each GEO partition). In particular, two flag variables geo_mmvd_idx0_flag and geo_mmvd_idx1_flag are used as two usage flags, together with the ciip_flag, to indicate the use of GEOs with the MMVD mode for each GEO partition. Furthermore, in this embodiment, four syntax elements are used to provide information for calculating the MVD of GEOs with the MMVD mode. These four syntax elements are geo_mmvd_distance[0], geo_mmvd_direction[0], geo_mmvd_distance[1], and geo_mmvd_direction[1]. In this regard, Tables F and G below show two examples of the syntax elements and their semantic descriptions.

[0078]

[0079]

[0080] Table F

[0081]

[0082]

[0083] Table G

[0084] The variable geo_mmvd_idx0_flag[x0][y0] indicates that the geometric segmentation merge mode with motion vector difference is used to generate the inter-frame prediction parameters associated with the first merge candidate index of the motion compensation candidate list based on geometric segmentation when it is equal to 1. The variable geo_mmvd_idx0_flag[x0][y0] indicates that the geometric segmentation mode with motion vector difference is not used to generate the inter-frame prediction parameters associated with the first merge candidate index of the motion compensation candidate list based on geometric segmentation when it is equal to 0. The array indices x0, y0 specify the position (x0, y0) of the top-left luminance sample of the coded block under consideration relative to the top-left luminance sample of the picture. When geo_mmvd_idx0_flag[x0][y0] does not exist, it is inferred to be equal to 0.

[0085] The variable geo_mmvd_idx1_flag[x0][y0] specifies that when equal to 1, the geometric partition merge mode with motion vector difference is used to generate the inter prediction parameters associated with the second merge candidate index of the geometric partition-based motion compensation candidate list. When the variable geo_mmvd_idx1_flag[x0][y0] is equal to 0, it specifies that the geometric partition mode with motion vector difference is not used to generate the inter prediction parameters associated with the second merge candidate index of the geometric partition-based motion compensation candidate list. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture. When geo_mmvd_idx1_flag[x0][y0] does not exist, it is inferred to be equal to 0.

[0086] The variable geo_mmvd_distance_idx[PartIdx][x0][y0] specifies the index used to derive GeoMmvdDistance[PartIdx][x0][y0], as specified in Table H. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.

[0087]

[0088] Table H

[0089] The variable geo_mmvd_direction_idx[PartIdx][x0][y0] specifies the index used to derive GeoMmvdSign[PartIdx][x0][y0], as specified in Table I. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.

[0090]

[0091] Table I

[0092] When geo_mmvd_idx0_flag is equal to 1, the two components of the geometric partition merge plus the MVD offset GeoMmvdOffset[0][x0][y0] are derived as follows:

[0093] GeoMmvdOffset[0][x0][y0][0] = (GeoMmvdDistance[0][x0][y0] << 2) * GeoMmvdSign[0][x0][y0][0]

[0094] GeoMmvdOffset[0][x0][y0][1] = (GeoMmvdDistance[0][x0][y0] << 2) * GeoMmvdSign[0][x0][y0][1]

[0095] When geo_mmvd_idx0_flag is equal to 0, the two components of the geometric split-merge plus the MVD offset GeoMmvdOffset[0][x0][y0] are set to zero.

[0096] When geo_mmvd_idx1_flag is equal to 1, the two components of the geometric split-merge plus the MVD offset GeoMmvdOffset[1][x0][y0] are derived as follows:

[0097] GeoMmvdOffset[1][x0][y0][0] = (GeoMmvdDistance[1][x0][y0] << 2) * GeoMmvdSign[1][x0][y0][0]

[0098] GeoMmvdOffset[1][x0][y0][1] = (GeoMmvdDistance[1][x0][y0] << 2) * GeoMmvdSign[1][x0][y0][1]

[0099] When geo_mmvd_idx1_flag is 0, the two components of the geometric split-merge plus the MVD offset GeoMmvdOffset[1][x0][y0] are all set to zero.

[0100] Figures 3A - 3B Shows the semantic changes to the VVC specification according to the operations described herein.

[0101] In another embodiment, two additional control flags are used to indicate the use of the proposed method. When the proposed method is enabled, MV refinement with up to two different offsets is used (e.g., one offset for each GEO partition). Similar to the above embodiment, two flag variables, geo_mmvd_idx0_flag and geo_mmvd_idx1_flag, are used as two usage flags, together with the ciip_flag, to indicate the use of GEOs with the MMVD mode for each GEO partition. Different from the above embodiment where geo_mmvd_idx0_flag and geo_mmvd_idx1_flag are independent of each other, the dependency between the two flag variables is utilized. For example, geo_mmvd_idx1_flag is signaled only when geo_mmvd_idx1_flag is indicated as true. In addition, four syntax elements provide information for calculating the MVD value of GEOs with the MMVD mode. The four new syntax elements are geo_mmvd_distance_idx0, geo_mmvd_direction_idx0, geo_mmvd_distance_idx1, and geo_mmvd_direction_idx1. In this regard, Tables J and K below show two examples of the syntax elements.

[0102]

[0103] Table J

[0104]

[0105]

[0106] Table K

[0107] Figure 2A flowchart depicting a method in accordance with certain example embodiments is shown. It should be understood that each block of the flowchart and combinations of blocks in the flowchart can be implemented in various ways, such as hardware, firmware, processors, circuitry, and / or other communication devices associated with the execution of software, including one or more computer program instructions. For example, one or more of the above processes can be implemented by computer program instructions. In this regard, the computer program instructions for implementing the above processes can be stored by a memory device 34 of a device adopting embodiments of the present disclosure and executed by a processor 32. As will be appreciated, any such computer program instructions can be loaded onto a computer or other programmable device (e.g., hardware) to produce a machine, such that the resulting computer or other programmable device implements the functions specified in the flowchart blocks. These computer program instructions can also be stored in a computer-readable memory, which can direct a computer or other programmable device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture, the execution of which implements the functions specified in the flowchart blocks. The computer program instructions can also be loaded onto a computer or other programmable device to cause a series of operations to be performed on the computer or other programmable device to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide operations for implementing the functions specified in the flowchart blocks.

[0108] Accordingly, the blocks of the flowchart support combinations of components for performing the specified functions and combinations of operations for performing the specified functions to perform the specified functions. It should also be understood that one or more blocks of the flowchart, as well as combinations of blocks in the flowchart, can be implemented by a special-purpose hardware-based computer system that performs the specified functions or by a combination of special-purpose hardware and computer instructions.

[0109] Many modifications and other embodiments of the disclosure will come to mind to those skilled in the art to which this disclosure pertains, having the benefit of the teachings presented in the foregoing description and the related drawings. Therefore, it is to be understood that the disclosure is not limited to the specific embodiments disclosed, and that modifications and other embodiments are intended to be included within the scope of the appended claims.

[0110] In addition, although the example embodiments have been described above in the context of certain example combinations of elements and / or functions, it should be understood that different combinations of elements and / or functions can be provided by alternative embodiments without departing from the scope of the appended claims. In this regard, for example, combinations of elements and / or functions different from those explicitly described above are also contemplated, as may be set forth in some of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. A device for communication, comprising: means for accessing an encoded unit that is divided into a first partition and a second partition; means for determining whether an enable value associated with the encoded unit indicates additional motion vector refinement for at least one of the first partition and the second partition, the additional motion vector refinement being supplementary to a base motion vector that has been applied to the at least one of the first partition and the second partition; and in accordance with determining that the enable value associated with the encoded unit indicates additional motion vector refinement: means for determining at least one motion vector difference for at least one of the first partition and the second partition; and means for applying prediction of the encoded unit based on the at least one motion vector difference.

2. The device according to claim 1, wherein the encoded unit is geometrically partitioned according to a geometric partition merge mode.

3. The device according to claim 1 or 2, wherein the at least one motion vector difference is at least based on a distance value and a direction value associated with the corresponding partition.

4. The apparatus according to claim 1 or 2, wherein the component for applying the prediction of the coding unit according to the at least one motion vector difference comprises: means for applying the same motion vector difference to both the first partition and the second partition.

5. The apparatus according to claim 1 or 2, wherein the means for applying the prediction of the coding unit based on the at least one motion vector difference comprises: means for applying the motion vector difference only to the larger one of the first partition or the second partition.

6. The device according to claim 1 or 2, wherein determining at least one motion vector difference for at least one of the first partition and the second partition is based on corresponding enable values associated with the first partition and the second partition.

7. The device according to claim 6, wherein when the corresponding enable value of a specific partition indicates that the motion vector difference will not be used, the at least one motion vector difference is not determined for the specific partition.

8. A method for communication, comprising: accessing an encoded unit that is divided into a first partition and a second partition; determining whether an enable value associated with the encoded unit indicates additional motion vector refinement for at least one of the first partition and the second partition, the additional motion vector refinement being supplementary to a base motion vector that has been applied to the at least one of the first partition and the second partition; and in accordance with determining that the enable value associated with the encoded unit indicates additional motion vector refinement: determining at least one motion vector difference for at least one of the first partition and the second partition; and applying prediction of the encoded unit based on the at least one motion vector difference.

9. The method according to claim 8, wherein the encoded unit is geometrically partitioned according to a geometric partition merge mode.

10. The method according to claim 8 or 9, wherein the at least one motion vector difference is at least based on a distance value and a direction value associated with the corresponding partition.

11. The method according to claim 8 or 9, wherein applying the prediction of the coding unit according to the at least one motion vector difference comprises: Applying the same motion vector difference to both the first partition and the second partition.

12. The method according to claim 8 or 9, wherein applying the prediction of the coding unit according to the at least one motion vector difference comprises: Applying the motion vector difference only to the larger one of the first partition or the second partition.

13. The method according to claim 8 or 9, wherein determining at least one motion vector difference for at least one of the first partition and the second partition is based on corresponding enable values associated with the first partition and the second partition.

14. The method according to claim 13, wherein the at least one motion vector difference is not determined for a particular partition if the corresponding enable value for the particular partition indicates that the motion vector difference will not be used.

15. A computer program product comprising a non-transitory computer-readable storage medium having program code portions stored thereon, the program code portions being configured, when executed, to: access an encoded unit that is partitioned into a first partition and a second partition; determine whether an enable value associated with the encoded unit indicates additional motion vector refinement for at least one of the first partition and the second partition, the additional motion vector refinement being supplementary to a base motion vector that has been applied to the at least one of the first partition and the second partition; and in accordance with determining that the enable value associated with the encoded unit indicates additional motion vector refinement: determine at least one motion vector difference for at least one of the first partition and the second partition; and apply prediction of the encoded unit according to the at least one motion vector difference.