Sub-block merging candidates in the triangle merging mode
By adopting the triangular partition mode and geometric merging motion fields in the video compression technology, the problem of low efficiency of the double prediction mode in the inter-frame encoding block in the prior art is solved, and more efficient video compression is achieved.
Patent Information
- Application Number
- CN202080080538.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-18
- Filing Date
- 2020-12-14
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-12-14
AI Technical Summary
The existing video compression technology is difficult to effectively utilize the dual prediction mode in inter-frame encoding blocks, resulting in low compression efficiency.
A method is proposed to improve the efficiency of video compression by performing motion compensation, weighting and encoding or decoding video blocks in the encoder or decoder using sub-block merging motion fields in the triangular partition mode and geometric merge mode.
By replacing conventional merge candidates with sub-block merge candidates, the spatial and temporal redundancy of video content can be more effectively utilized, and the compression efficiency of video compression can be improved.
Smart Images

Figure CN114930819B_ABST
Abstract
Description
Technical Field
[0001] At least one of the embodiments herein generally relates to a method or an apparatus for video encoding or decoding, compression or decompression. Background Art
[0002] To achieve high compression efficiency, image and video coding schemes typically use prediction, including motion vector prediction, and transforms to exploit spatial and temporal redundancy in video content. In general, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlations, and then the difference between the original image and the predicted image (usually expressed as a prediction error or prediction residual) is transformed, quantized, and entropy encoded. To reconstruct the video, the compressed data is decoded by the inverse process corresponding to entropy coding, quantization, transform, and prediction. Summary of the invention
[0003] At least one of the embodiments of the present invention generally relates to a method or apparatus for video encoding or decoding, and more particularly to a method and apparatus for simplifying a coding mode based on a neighboring sample dependent parameter model.
[0004] According to a first aspect, a method is provided. The method comprises steps for the following operations: obtaining a unidirectional triangle candidate from a merge list for encoding a video block; extracting a unidirectional portion of the unidirectional triangle candidate; obtaining a triangle merge candidate from a candidate sub-block merge list; extracting a unidirectional seed portion of the triangle merge candidate; performing motion compensation using the unidirectional triangle candidate and the triangle merge candidate; performing weighting on the motion compensation result; signaling whether a candidate sub-block list or a regular list is used for the motion compensation result; and encoding a video block using the weighted motion compensation result.
[0005] According to a first aspect, a method is provided. The method includes steps for the following operations: parsing a video bitstream to determine whether to select a unidirectional triangle candidate from a sub-block merge list or a regular merge list; obtaining a unidirectional triangle candidate from a sub-block merge list or a regular merge list for decoding a video block; extracting a unidirectional portion of the unidirectional triangle candidate; performing weighting on the motion compensation result; performing motion compensation using the unidirectional triangle candidate; and decoding the video block using the weighted motion compensation result.
[0006] According to another aspect, a device is provided. The device includes a processor. The processor can be configured to encode a video block or decode a bitstream by executing any one of the aforementioned methods.
[0007] In another general aspect according to at least one embodiment, there is provided an apparatus comprising: means according to any one of the decoding embodiments; and at least one of the following: (i) an antenna configured to receive a signal including a video block; (ii) a band limiter configured to limit the received signal to a band including the video block; and (iii) a display configured to display an output representing the video block.
[0008] In another general aspect according to at least one embodiment, there is provided a non-transitory computer-readable medium containing data content generated according to any one of the described encoding embodiments or variations.
[0009] In another general aspect according to at least one embodiment, there is provided a signal including video data generated according to any one of the described encoding embodiments or variations.
[0010] In another general aspect according to at least one embodiment, a bitstream is formatted to include data content generated according to any one of the described encoding embodiments or variations.
[0011] In another general aspect according to at least one embodiment, there is provided a computer program product comprising instructions that, when executed by a computer, cause the computer to perform any one of the described encoding embodiments or variations.
[0012] These and other aspects, features, and advantages of the general aspects will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 Illustrates coding tree units and coding tree concepts for representing compressed HEVC pictures.
[0014] Figure 2 Illustrates an exemplary partitioning of coding tree units into coding units, prediction units, and transform units.
[0015] Figure 3 Illustrates inter prediction based on triangular partitioning.
[0016] Figure 4 Illustrates an example of single prediction motion vector selection for triangular partitioning modes.
[0017] Figure 5 Illustrates geometric segmentation decisions.
[0018] Figure 6Shows an exemplary geometric partition with an angle of 12 and a distance of 0.
[0019] Figure 7 Shows an exemplary geometric partition with an angle of 12 and a distance of 1.
[0020] Figure 8 Shows an exemplary geometric partition with an angle of 12 and a distance of 2.
[0021] Figure 9 Shows an exemplary geometric partition with an angle of 12 and a distance of 3.
[0022] Figure 10 Shows 32 angles in a geometric pattern.
[0023] Figure 11 Shows 24 angles for geometric partitions.
[0024] Figure 12 Shows the angles proposed in a method for a GEO with its corresponding width-to-height ratio.
[0025] Figure 13 Shows a standard general video compression scheme.
[0026] Figure 14 Shows a standard general video decompression scheme.
[0027] Figure 15 Shows an exemplary flowchart of a proposed sub-block triangle pattern.
[0028] Figure 16 Shows an example of motion field storage where the top sub-block partition and the bottom regular partition have a) bidirectional diagonals and b) regular / bottom unidirectional diagonals.
[0029] Figure 17 Shows a processor-based system for encoding / decoding in the context of a general description.
[0030] Figure 18 Shows an embodiment of a method in the context of a general description.
[0031] Figure 19 Shows another embodiment of a method in the context of a general description.
[0032] Figure 20 Shows an exemplary device in the context of the described aspect. Detailed Description
[0033] The embodiments described herein are in the field of video compression and generally relate to video compression and video encoding and decoding, and more particularly to the quantization step of a video compression scheme. The general aspects described are intended to provide mechanisms for manipulating the constraints in the High Efficiency Video Coding (HEVC) syntax or video coding semantics to restrict the set of possible tool combinations.
[0034] To achieve high compression efficiency, image and video coding schemes typically employ prediction including motion vector prediction and transformation to exploit the spatial and temporal redundancy in video content. Generally, intra-frame or inter-frame prediction is used to exploit the intra-frame or inter-frame correlation, and then the difference between the original image and the predicted image (usually represented as prediction error or prediction residual) is transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded through the inverse processes corresponding to entropy coding, quantization, transformation, and prediction.
[0035] In the HEVC (High Efficiency Video Coding, ISO / IEC 23008–2, ITU-T H.265) video compression standard, motion compensated temporal prediction is employed to exploit the redundancy present between consecutive pictures of a video.
[0036] To this end, a motion vector is associated with each prediction unit (PU). Each coding tree unit (CTU) is represented by a coding tree in the compressed domain. As Figure 1 shown, this is the quadtree partitioning of a CTU, where each leaf is called a coding unit (CU).
[0037] Then, each CU is given some intra-frame or inter-frame prediction parameters (prediction information). To this end, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. As Figure 2 shown, the intra-frame or inter-frame coding mode is assigned at the CU level.
[0038] In HEVC, exactly one motion vector is assigned to each PU. This motion vector is used for motion compensated temporal prediction of the considered PU. Thus, in HEVC, the motion model that links the predicted block and its reference block simply consists of a translation.
[0039] In the JVET (Joint Video Exploration Team) which proposed a new video compression standard called Joint Exploration Model (JEM), a Quadtree Binary Tree (QTBT) block partition structure has been proposed due to its high compression performance. The blocks in the Binary Tree (BT) can be divided into two equal-sized sub-blocks by splitting horizontally or vertically in the middle. Therefore, the BT blocks can have a rectangular shape with unequal width and height, which is different from the blocks in the QT where the blocks always have a square shape with equal height and width. In HEVC, the angular intra prediction direction is defined as spanning 180 degrees from 45 degrees to -135 degrees, and this angular intra prediction direction is maintained in JEM, which defines the angular direction independent of the target block shape.
[0040] To encode these blocks, intra prediction is used to provide an estimated version of the block using previously reconstructed neighboring samples. Then, the difference between the source block and the prediction is encoded. In the above classical codecs, a single line of reference samples is used at the left and top of the current block.
[0041] In HEVC (High Efficiency Video Coding, H.265), the frames of a video sequence are encoded based on a quadtree (QT) block partition structure. Based on the rate-distortion (RD) criterion, the frame is divided into square Coding Tree Units (CTUs), and all these square CTUs undergo a quadtree-based split into multiple Coding Units (CUs). Each CU is either intra predicted, i.e., each CU is predicted spatially from causally neighboring CUs, or inter predicted, i.e., each CU is predicted temporally from already decoded reference frames. In an I slice, all CUs are intra predicted, while in P slices and B slices, CUs can be either intra predicted or inter predicted. For intra prediction, HEVC defines 35 prediction modes, which include a planar mode (indexed as mode 0), a DC mode (indexed as mode 1), and 33 angular modes (indexed as modes 2 to mode 34). The angular modes are associated with prediction directions ranging from 45 degrees to -135 degrees in the clockwise direction. Since HEVC supports a quadtree (QT) block partition structure, all Prediction Units (PUs) have a square shape. Therefore, the definition of the prediction angles from 45 degrees to -135 degrees is verified from the perspective of the PU (Prediction Unit) shape. For a target prediction unit of NxN pixel size, the top reference array and the left reference array are each of size 2N+1 samples, which need to cover the aforementioned angular range of all target pixels. Considering that the height and width lengths of the PU are equal, it also makes sense that the lengths of the two reference arrays are equal.
[0042] The present invention belongs to the field of video compression. Compared with existing video compression systems, the present invention aims to improve dual prediction in inter-coded blocks. The present invention also proposes separate luminance and chrominance coding trees for inter slices.
[0043] In the HEVC video compression standard, pictures are divided into so-called coding tree units (CTUs), and the size of a coding tree unit is typically 64×64, 128×128, or 256×256 pixels. Each CTU is represented by a coding tree in the compression domain. See Figure 1 , which is the quadtree partitioning of the CTU, where each leaf is called a coding unit (CU).
[0044] Then, each CU is given some intra or inter prediction parameters (prediction information). To this end, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. See Figure 2 , and an intra or inter coding mode is assigned at the CU level.
[0045] In some video coding standards, the triangle partitioning mode (TPM) is supported for inter prediction. The triangle partitioning mode can be derived at the CU level as a residual merging mode after other merging modes including the regular merge mode, MMVD mode, sub-block merge mode, and CIIP mode.
[0046] When this mode is used, the CU is evenly divided into two triangle partitions using a diagonal split or an anti-diagonal split, as Figure 3 shown. Each triangle partition in the CU uses its own motion for inter prediction; each partition only allows single prediction, that is, each partition has one motion vector and one reference index. The single prediction motion constraint is applied to ensure that, like the regular dual prediction, each CU only needs two motion compensated predictions. The single prediction motion for each partition is derived using the process described below.
[0047] If the triangle partitioning mode is used for the current CU, the direction indicating the triangle partition (diagonal or anti-diagonal) is marked, and further two merge indices (one for each partition) are signaled, as described below. The maximum number of TPM candidates is signaled explicitly at the slice level, and the syntax binarization of the TMP merge index is specified. After predicting each of the triangle partitions, a hybrid process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edges. This is the prediction signal for the entire CU, and then the transform and quantization processes are applied to the entire CU as in other prediction modes. Finally, the motion field of the CU predicted using the triangle partitioning mode is stored in 4×4 units, as described below.
[0048] This single prediction candidate list is directly derived from the merge candidate list constructed according to the extended merge prediction process. Let n denote the index of the single prediction motion in the triangle single prediction candidate list. The LX motion vector of the nth extended merge candidate (where X is equal to the parity of n) is used as the nth single prediction motion vector of the triangle partitioning mode. These motion vectors are inFigure 4 Marked as "x" in the middle. In the absence of the corresponding LX motion vector of the nth extended merge candidate, the L(1 - X) motion vector of the same candidate is used as the single prediction motion vector for the triangular partition pattern.
[0049] There are up to 5 single prediction candidates, and the encoder must test all combinations of candidates (one for each partition) in 2 split directions. Therefore, under actual common test conditions (CTC), the maximum number of combinations tested is 40 (5 * 4 * 2), (where MaxNumTriangleMergeCand = 5 and nb_combinations = MaxNumTriangleMergeCand * (MaxNumTriangleMergeCand - 1) * 2).
[0050] As described above, the triangular partition pattern requires signaling (i) a binary flag indicating the triangular partition direction, and (ii) two merge indices after derivation when using the pattern.
[0051] The following presents a part of the typical video compression standard (VVC specification proposal 7.0) that relates to triangular partition signaling (highlighted).
[0052] 7.3.9.7 Merge Data Syntax
[0053]
[0054]
[0055] As described above, the motion fields of triangular prediction CUs are stored on 4×4 sub - blocks. Each triangular partition stores the unidirectional motion field of the corresponding candidate. In the split direction, if the candidates resolve different reference picture lists, the diagonal 4×4 sub - block stores the bidirectional motion field as a combination of two unidirectional motion fields. Otherwise, when the two candidates resolve the same reference picture list, the diagonal 4×4 sub - block stores the unidirectional motion field of the bottom candidate.
[0056] The geometric merge mode has been proposed to have 32 angles and 5 distances. The angles are quantized to be between 0 degrees and 360 degrees, where the step angle is equal to 11.25 degrees. A total of 32 angles are proposed, as Figure 10 shown. Figure 5 The description of geometric segmentation using angles and distance ρ i is depicted.
[0057] The distance ρ i is quantized to the maximum possible distance ρ max, having a fixed step size, which indicates the distance from the center of the block. For distance ρ i = 0, only the first half of the angles are available in the case of split symmetry. Figures 6 to 9 depicts the result of a geometric partition using angle 12 and distances between 0 and 3.
[0058] For distance ρ i equal to 0, symmetric angles 16 to 31 are removed because this symmetric angle corresponds to the same partition as 0 - 15, diagonal angles 4 and 12 are excluded because this diagonal angle is equivalent to two TPM partition patterns. Angles 0 and 8 are also excluded because said angles are similar to the binary partition of the CU, leaving only 12 angles for distance 0. Thus, the geometric partition can use a maximum of 140 partition patterns (12 + 32 * 4 = 140 or 14 + 32 * 4 = 142 if the triangular partition pattern is removed and the corresponding angle is added in the geometric partition pattern).
[0059] A 24 - angle scheme with only 4 distances is also proposed by removing angles near the vertical value, as Figure 11 shown. In this case, the maximum 82 pattern is used by the geometric partition (10 + 24 * 3 = 82).
[0060] Hereinafter, it will be considered that by enabling angles 4 and 12 to integrate the triangular partition pattern in the geometric merge pattern for a distance equal to 0.
[0061] An example of the syntax proposed for geometric partitioning is depicted in Table 1. The truncated binary (TB) binarization process is used to encode wedge_partition_idx (Table 1), and the mapping between wedge_partition_idx and angles and distances is shown in Table 2.
[0062]
[0063]
[0064] Table 1: Merge Data Syntax for Geometric Partitioning in Exemplary Schemes
[0065] wedge_partition_idx[x0][y0] specifies the type of geometric partition of the merge geometric pattern. The array indices x0, y0 specify the position (x0, y0) of the top - left luminance sample of the coding block being considered with respect to the top - left luminance sample of the picture.
[0066]
[0067]
[0068] Table 1: Specifications of angleIdx and distanceIdx Values Based on wedge_partition idx Values
[0069]
[0070] Table 2: Syntax Elements and Associated Binarization
[0071] The angles in GEO are replaced with angles having powers of 2 as tangents. Since the tangents of the proposed angles are numbers that are powers of 2, most multiplications can be replaced by bit shifts. With the proposed angles, one row or one column is needed to store each block size and each partitioning pattern, as Figure 12 depicted in
[0072] One problem solved by the present invention is to allow the use of sub-block merged motion fields in triangular and GEO coding modes. In at least one proposed scheme, when using the triangular partitioning pattern, the unidirectional candidates considered can be derived only from the regular merge list.
[0073] At least one embodiment proposes to consider sub-block merge candidates in triangular and GEO partitioning patterns.
[0074] The present invention covers the following general aspects:
[0075] - If sub-block merge candidates are used instead of regular merge candidates, add a flag to signal.
[0076] - Replace some regular merge candidates with sub-block merge candidates.
[0077] - Store the motion field.
[0078] - Consider all sub-block merge candidates or only SbTMVP or only affine candidates.
[0079] The present invention relates to a motion compensation circuit at least in an encoder or a decoder. The following embodiments include some of the general aspects herein.
[0080] Since GEO is an extension of the triangular partitioning pattern (or the triangle is a specific case of GEO), all descriptions are based on the triangular partitioning pattern, but can also be applied to the GEO pattern.
[0081] The two unidirectional candidates for a triangular CU (one for each triangular partition) are denoted as Cand0 and Cand1.
[0082] Each of these unidirectional candidates (Cand x ) corresponds to merge_triangle_idx x (denoted as idx x ), and is associated with a motion vector mv through a reference picture list L y (where y is given by the parity of the index idx x )x and its exponent refidx x is associated with the reference frame
[0083] In the first embodiment, a new triangle pattern dedicated to sub-block candidates is proposed.
[0084] In addition, all processes of copying the actual triangle partitioning pattern are performed to test candidates derived from the sub-block merge pattern list.
[0085] For this purpose, the encoder performs triangle partitioning pattern RDO twice. The first pass is VTM-7.0, where unidirectional triangle candidates are selected from the regular merge list. In the second pass, candidates are selected from the sub-block merge candidate list, and all remaining triangle processes remain unchanged.
[0086] New CU attributes must be added to carry the list type (sub-block or regular) used. The new CU attributes must also be signaled and are represented as subblock_merge_triangle.
[0087] On the decoder side, if subblock_merge_triangle is marked as 1, unidirectional triangle candidates are selected from the sub-block merge list; otherwise, they are selected from the regular merge list.
[0088] In the current standard, where triangle partitioning signaling (highlighted) is presented for the example (italic):
[0089] 7.3.9.7 Merge Data Syntax
[0090]
[0091]
[0092] subblock_merge_triangle[x0][y0] specifies whether the sub-block pattern list and motion compensation must be used instead of the regular pattern list and motion compensation. The array indices x0, y0 specify the position (x0, y0) of the top-left luma sample of the coding block under consideration relative to the top-left luma sample of the picture.
[0093] When subblock_merge_triangle[x0][y0] does not exist, it is indicated to be equal to 0.
[0094]
[0095] In the second embodiment, in order to limit the impact of encoder complexity and avoid modifying the syntax, some of the existing uni-directional triangle candidates may be replaced by sub-block candidates. Additionally, in this case, a sub-block partition may be combined with a regular partition.
[0096] In a variant, the first candidate of the uni-directional triangle list is selected from the regular merge list and the last candidate is selected from the sub-block merge list. As an example, the first three candidates are the same as the candidates from the regular merge list in VTM-7.0, and the last two candidates are the first two uni-directional candidates selected from the sub-block merge list.
[0097] In another variant, even candidate indices refer to candidates selected from the regular merge list and odd candidate indices refer to candidates selected from the sub-block merge list.
[0098] In another variant, compared with VTM-7.0, the number of uni-directional triangle candidates can be increased. In VTM-7.0, this number is set to 5. For example, this number can be doubled to retain all the VTM-7.0 candidates selected from the regular merge list and also retain all the VTM-7.0 candidates from the sub-block merge list.
[0099] Motion Field Storage:
[0100] As described above, the motion fields of the triangular prediction CUs are still stored on a 4×4 sub-block basis.
[0101] Each triangular partition stores the uni-directional motion of the corresponding candidate, even if it is from the regular or sub-block merge list.
[0102] In the splitting direction, if the candidates resolve different reference picture lists, the diagonal 4×4 sub-block stores the bi-directional motion field as a combination of two uni-directional motion fields. Otherwise, when the two candidates resolve the same reference picture list, the diagonal 4×4 sub-block stores the uni-directional motion field of the bottom candidate.
[0103] In a variant, when both candidates resolve the same reference picture list, the diagonal 4×4 sub-block may store the uni-directional motion field of the bottom candidate when both candidates are from the same merge list, and store the regular uni-directional motion field when the candidates are from different merge lists.
[0104] In another variant, it is also possible to store the affine CPMV (seeds) associated with the triangular partition when the candidate is from the sub-block merge list and is affine. Then, these CPMVs can be used for inheritance.
[0105] The sub-block merge list in VTM-7.0 includes 5 candidates of different natures. It can be an SbTMVP candidate, an inherited or constructed affine model.
[0106] In two previous embodiments, one can choose to select unidirectional triangle candidates in the full sub-block merge list, or only use SbTMVP candidates, all affine candidates, only inherited affine candidates, or constructed affine candidates.
[0107] The use of sub-block merge candidates may be restricted in previous embodiments. In this case, when the constraints are not met, the sub-block merge list is not considered in the unidirectional triangle candidate list.
[0108] The use of the sub-block merge list in the triangular partition mode can be restricted according to the size of the current CU.
[0109] It can be restricted according to its size (i.e., its width and height). For example, the use of the sub-block list is only allowed for CUs with a width and height greater than or equal to 8 (w>=8 && h>=8) (or (w>=16 && h>=16) (w + h>12)…).
[0110] It can also be restricted according to its area. For example, the use of the sub-block list is only allowed for CUs with an area greater than 32 (w * h>32) (or (w * h>=64) (w * h<256)…).
[0111] In the case where only affine candidates (all affine or only inherited affine candidates) are considered during the construction of the unidirectional triangle list, the use of the sub-block merge list can be restricted to a CU that has at least an affine adjacent CU, i.e., only when the sub-block merge list contains an inherited affine model.
[0112] Figure 18An embodiment of method 1800 according to general aspects described herein is shown. The method begins at start block 1801, and control proceeds to block 1810 to obtain unidirectional triangle candidates from a merge list for encoding a video block. Control proceeds from block 1810 to block 1820 to extract the unidirectional part of the unidirectional triangle candidates. Control proceeds from block 1820 to block 1830 to obtain triangle merge candidates from a sub-block merge list of candidates. Control proceeds from block 1830 to block 1840 to extract the unidirectional seed part of the triangle merge candidates. Control proceeds from block 1840 to block 1850 to perform motion compensation using the unidirectional triangle candidates and the triangle merge candidates. Control proceeds from block 1850 to block 1860 to perform weighting on the motion compensation result. Control proceeds from block 1860 to block 1870 to signal whether a sub-block list of candidates or a regular list is used for the motion compensation result. Control proceeds from block 1870 to block 1880 to encode the video block using the weighted motion compensation result.
[0113] Figure 19 An embodiment of method 1900 according to general aspects described herein is shown. The method begins at start block 1901, and control proceeds to block 1910 to parse a video bitstream to determine whether unidirectional triangle candidates are selected from a sub-block merge list or a regular merge list. Control proceeds from block 1910 to block 1920 to obtain unidirectional triangle candidates from a sub-block merge list or a regular merge list for decoding a video block. Control proceeds from block 1920 to block 1930 to extract the unidirectional part of the unidirectional triangle candidates. Control proceeds from block 1930 to block 1940 to perform weighting on the motion compensation result. Control proceeds from block 1940 to block 1950 to perform motion compensation using the unidirectional triangle candidates. Control proceeds from block 1950 to block 1960 to decode the video block using the weighted motion compensation result.
[0114] Figure 20 An embodiment of apparatus 2000 for encoding, decoding, compressing, or decompressing video data using a simplification of an encoding mode based on an adjacent sample dependency parameter model is shown. The apparatus includes a processor 2010 and may be interconnected to a memory 2020 through at least one port. Both the processor 2010 and the memory 2020 may also have one or more additional interconnections to the outside.
[0115] The processor 2010 is further configured to insert or receive information in a bitstream and perform compression, encoding, or decoding using any of the aspects.
[0116] This application describes various aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are specifically described, and at least individual characteristics are shown, usually in a manner that may sound limited. However, this is for the purpose of describing clearly and does not limit the application or scope of these aspects. In fact, all different aspects can be combined and interchanged to provide further aspects. In addition, these aspects can also be combined and interchanged with the aspects described in the previous submission.
[0117] The aspects described and contemplated in this patent application can be implemented in many different forms. Figure 13 , Figure 14 and Figure 17 Some embodiments are provided, but other embodiments are contemplated, and Figure 13 , Figure 14 and Figure 17 The discussion of does not limit the breadth of the embodiments. At least one of these aspects generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting a generated or encoded bitstream. These and other aspects can be implemented as methods, apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the methods.
[0118] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture" and "frame" can be used interchangeably. Usually, but not necessarily, the term "reconstruction" is used at the encoding end, while "decoding" is used at the decoding end.
[0119] Various methods are described herein, and each method includes one or more steps or actions for implementing the method. Unless a specific order of steps or actions is required for proper operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0120] Various methods and other aspects described in this patent application may be used to modify modules (e.g., intra-frame prediction, entropy encoding and / or decoding modules (160, 360, 145, 330)) of the video encoder 100 and decoder 200, such as Figure 13 and Figure 14 In addition, aspects of the present invention are not limited to VVC or HEVC, and may be applied, for example, to other standards and recommendations (whether pre-existing or developed in the future) and extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise specified or technically excluded, the aspects described in this application may be used alone or in combination.
[0121] Various numerical values are used in this application. The specific values are for illustrative purposes, and the aspects are not limited to these specific values.
[0122] Figure 13 Encoder 100 is illustrated. Variations of this encoder 100 are envisioned, but for clarity, encoder 100 is described below without describing all the expected variations.
[0123] Before encoding, the video sequence may undergo pre - encoding processing (101). For example, applying a color transformation to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components in the color components). Metadata may be associated with the pre - processing and attached to the bitstream.
[0124] In encoder 100, pictures are encoded by encoder elements as described below. The picture to be encoded is partitioned (102) and processed in units such as CUs, for example. Each unit is encoded using, for example, an intra - frame mode or an inter - frame mode. When a unit is encoded in the intra - frame mode, it performs intra - frame prediction (160). In the inter - frame mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) which of the intra - frame mode or the inter - frame mode is used to encode the unit, and indicates the intra - frame / inter - frame decision, for example, by a prediction mode flag. The prediction residual is calculated, for example, by subtracting (110) the prediction block from the original image block.
[0125] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, along with motion vectors and other syntax elements, are entropy - encoded (145) to output a bitstream. The encoder may skip the transformation and directly apply quantization to the untransformed residual signal. The encoder may bypass both the transformation and quantization, i.e., directly encode the residual without applying the transformation or quantization process.
[0126] The encoder decodes the encoded blocks to provide a reference for further prediction. The quantized transform coefficients are de - quantized (140) and inverse - transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. A loop filter (165) is applied to the reconstructed image to perform, for example, de - blocking effect / SAO (Sample Adaptive Offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (180).
[0127] Figure 14 A block diagram of video decoder 200 is illustrated. In decoder 200, the bitstream is decoded by decoder elements as described below. Video decoder 200 generally performs the same asFigure 13 The encoding process described above has an opposite decoding process. Encoder 100 typically also performs video decoding as part of encoding video data.
[0128] Specifically, the input to the decoder includes a video bitstream, which can be generated by video encoder 100. First, entropy decoding (230) is performed on the bitstream to obtain transform coefficients, motion vectors, and other encoded information. Picture partitioning information indicates how to partition a picture. Thus, the decoder can partition (235) the picture according to the decoded picture partitioning information. The transform coefficients are dequantized (240) and inverse-transformed (250) to decode the prediction residuals. The decoded prediction residuals and prediction blocks are combined (255) to reconstruct an image block. The prediction block can be obtained (270) from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275). A loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0129] The decoded picture can also undergo post-decoding processing (285), for example, an inverse color transformation (e.g., a transformation from YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the remapping process performed in pre-encoding processing (101). The post-decoding processing can use metadata derived in the pre-encoding processing and signaled in the bitstream.
[0130] Figure 17 A block diagram illustrating an example of a system in which various aspects and embodiments are implemented. System 1000 can be embodied as a device including the various components described below, and is configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablets, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 1000 can be embodied singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing elements and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.
[0131] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein for implementing various aspects as described, for example, in this document. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., volatile memory devices and / or non-volatile memory devices). System 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read-only memory (EEPROM), read-only memory (ROM), programmable read-only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non-limiting example, the storage device 1040 may include internal storage devices, attached storage devices (including removable and non-removable storage devices), and / or network-accessible storage devices.
[0132] System 1000 includes an encoder / decoder module 1030 configured to, for example, process data to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is well known, a device may include one or both of an encoding module and a decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a stand-alone element of System 1000 or may be incorporated within the processor 1010 as a combination of hardware and software known to those skilled in the art.
[0133] The program code to be loaded onto the processor 1010 or encoder / decoder 1030 to execute the various aspects described in this document may be stored in the storage device 1040 and subsequently loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of the various items during the execution of the processes described in this document. Such stored items may include but are not limited to input video, decoded video or partially decoded video, bitstreams, matrices, variables, and intermediate or final results of processing equations, formulas, operations, and operation logic.
[0134] In some embodiments, the memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be the memory 1020 and / or the storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In several embodiments, the external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM is used as the working memory for video encoding and decoding operations, such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by the Joint Video Exploration Team (JVET)).
[0135] Inputs to the elements of the system 1000 can be provided by various input devices as shown in block 1130. Such input devices include, but are not limited to: (i) a radio frequency (RF) section that receives, for example, RF signals transmitted over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 17 Other examples not shown, including composite video.
[0136] In various embodiments, the input device of block 1130 has corresponding input processing elements associated therewith as known in the art. For example, the RF section may be associated with elements suitable for: (i) selecting a desired frequency (also referred to as selecting a signal, or band-limiting a signal band to one band), (ii) down-converting the selected signal, (iii) again band-limiting to a narrower band to select a signal band that may be referred to as a channel in some embodiments, (iv) demodulating the down-converted and band-limited signal, (v) performing error correction, and (vi) de-multiplexing to select a desired data packet stream. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, filters, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section may include a tuner that performs various of these functions, including, for example, down-converting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive an RF signal transmitted via a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, e.g., inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.
[0137] In addition, the USB and / or HDMI terminals may include corresponding interface processors for connecting system 1000 to other electronic devices across the USB and / or HDMI connections. It should be understood that various aspects of input processing (e.g., Reed-Solomon error correction) may be implemented, as needed, e.g., within a separate input processing IC or within processor 1010. Similarly, aspects of USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within processor 1010. The demodulated stream, error-corrected stream, and de-multiplexed stream are provided to various processing elements, including, for example, processor 1010 as well as encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.
[0138] The various elements of system 1000 may be disposed within an integrated housing in which the various elements may be interconnected using suitable connection arrangements (e.g., internal buses known in the art, including inter-integrated circuit (I2C) buses, wiring, and printed circuit boards) and data may be transmitted therebetween.
[0139] System 1000 includes a communication interface 1050 capable of communicating with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0140] In various embodiments, data is streamed or otherwise provided to system 1000 using a wireless network such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers). The Wi-Fi signals of these embodiments are received via the communication channel 1060 and the communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or a router that provides access to an external network including the Internet for allowing streaming applications and other cloud-based communications. Other embodiments use a set-top box to provide streaming data to system 1000, and the set-top box delivers data via an HDMI connection of the input block 1130. Still other embodiments use an RF connection of the input block 1130 to provide streaming data to system 1000. As described above, various embodiments provide data in a non-streaming manner. In addition, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.
[0141] System 1000 may provide output signals to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touchscreen display, an organic light-emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 may be used for a television, a tablet, a notebook, a cellular phone (mobile phone), or other devices. The display 1100 may also be integrated with other components (e.g., as in a smart phone) or be separate (e.g., an external monitor for a laptop). In various examples of embodiments, other peripheral devices 1120 include one or more of an independent digital video disc (or digital versatile disc, both terms are DVR), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120 that provide functions based on the output of system 1000. For example, a disc player performs the function of playing the output of system 1000.
[0142] In various embodiments, control signals are transmitted between system 1000 and display 1100, speaker 1110, or other peripheral devices 1120 using signaling such as AV.Link, Consumer Electronics Control (CEC), or other communication protocols that enable device-to-device control with or without user intervention. The output devices may be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, the output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. Display 1100 and speaker 1110 may be integrated into a single unit with other components of system 1000 in an electronic device such as a television. In various embodiments, display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0143] Alternatively, if the RF portion of input 1130 is part of a separate set-top box, display 1100 and speaker 1110 are optionally separate from one or more of the other components. In various embodiments where display 1100 and speaker 1110 are external components, output signals may be provided via dedicated output connections including, for example, an HDMI port, a USB port, or a COMP output.
[0144] These embodiments may be implemented by processor 1010 or by computer software implemented by hardware or a combination of hardware and software. As a non-limiting example, these embodiments may be implemented by one or more integrated circuits. As a non-limiting example, memory 1020 may be of any type suitable for the technical environment and may be implemented using any appropriate data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. As a non-limiting example, processor 1010 may be of any type suitable for the technical environment and may encompass one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0145] Various embodiments participate in decoding. As used in this application, "decoding" may include, for example, performing all or part of a process on a received encoded sequence to produce a final output suitable for display. In various embodiments, such processes include one or more processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include or alternatively include processes performed by the decoders of the various embodiments described in this application.
[0146] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in yet another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" specifically refers to a subset of operations or more generally to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0147] Various specific implementations participate in encoding. In a manner similar to the discussion of "decoding" above, "encoding" as used in this application can cover, for example, all or part of the process performed on an input video sequence to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, such as partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also include or alternatively include processes performed by the encoders of the various specific implementations described in this application.
[0148] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in yet another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" specifically refers to a subset of operations or more generally to a broader encoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0149] Note that the grammatical elements used herein are descriptive terms. Thus, they do not exclude the use of other grammatical element names.
[0150] When the drawings are presented as flowcharts, it should be understood that they also provide block diagrams of the corresponding apparatus. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide flowcharts of the corresponding method / process.
[0151] Various embodiments may refer to parametric models or rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often taking into account constraints on computational complexity. It can be measured by a rate-distortion optimization (RDO) metric or by least mean square (LMS), mean average error (MAE), or other such measures. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. There are different ways to solve the rate-distortion optimization problem. For example, these methods can be based on extensive testing of all coding options (including all considered modes or coding parameter values) and a complete evaluation of their coding costs as well as the associated distortion of the reconstructed signal after encoding and decoding. Faster methods can also be used to reduce the coding complexity, especially for the calculation of an approximate distortion based on the predicted or prediction residual signal rather than the reconstructed residual signal. A hybrid of these two methods can also be used, such as by using approximate distortion for only some of the possible coding options and full distortion for other coding options. Other methods only evaluate a subset of the possible coding options. More generally, many methods employ any one of various techniques to perform the optimization, but the optimization does not necessarily involve a complete evaluation of both the coding cost and the associated distortion.
[0152] The specific implementations and aspects described herein can be implemented, for example, in a method or process, a device, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of a specific implementation (e.g., only as a method), the specific implementation of the discussed features can be implemented in other forms (e.g., a device or a program). A device can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor that generally refers to a processing device, which includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. The processor also includes communication devices, such as, for example, a computer, a mobile phone, a portable / personal digital assistant (“PDA”), and other devices that facilitate information communication among end users.
[0153] Reference to “one embodiment” or “an embodiment” or “one specific implementation” or “a specific implementation” and other variations thereof means that the particular features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearance of the phrase “in one embodiment” or “in an embodiment” or “in one specific implementation” or “in a specific implementation” and any other variations that occur throughout this application do not necessarily all refer to the same embodiment.
[0154] Additionally, this application may relate to “determining” various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.
[0155] In addition, the present application may relate to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, computing information, determining information, predicting information, or estimating information.
[0156] Additionally, the present application may relate to "receiving" various information. Like "accessing", receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Further, "receiving" is typically involved in one way or another during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, computing information, determining information, predicting information, or estimating information.
[0157] It should be understood that, for example, in the case of "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following " / " and "and / or" and "at least one" is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the case of "A, B, and / or C" and "at least one of A, B, and C", such phrases are intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or the selection of the first-listed option and the second-listed option (A and B), or the selection of the first-listed option and the third-listed option (A and C), or the selection of the second-listed option and the third-listed option (B and C), or the selection of all three options (A and B and C). As will be apparent to those of ordinary skill in the art and related fields, this can be extended to as many items as are listed.
[0158] Moreover, as used herein, the term "signal" (verb) means, among other things, to indicate something to a corresponding decoder. For example, in some embodiments, an encoder signals a particular one of a plurality of transforms, coding modes, or flags. Thus, in one embodiment, the same transform, parameters, or mode is used on both the encoder side and the decoder side. For example, the encoder may transmit (explicit signaling) a particular parameter to the decoder such that the decoder may use the same particular parameter. Conversely, if the decoder already has the particular parameter among others, signaling may be used without transmitting (implicit signaling) simply to allow the decoder to know and select the particular parameter. By avoiding transmitting any actual functionality, bit savings are achieved in various embodiments. It should be understood that signaling may be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the term "signal", the term "signal" (noun) may also be used herein.
[0159] It will be apparent to those of ordinary skill in the art that a particular implementation may generate various signals formatted to carry information such as may be stored or transmitted. The information may include, for example, instructions for performing a method or data generated by one of the particular implementations. For example, a signal may be formatted to carry a bitstream of the embodiment. Such a signal may be formatted, for example, as an electromagnetic wave (e.g., using a radio frequency portion of the spectrum) or a baseband signal. Formatting may include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal may be, for example, analog or digital information. It is known that signals may be transmitted over various different wired or wireless links. A signal may be stored on a processor-readable medium.
[0160] We have described a number of embodiments, across various claim categories and types. The features of these embodiments may be provided individually or in any combination. In addition, embodiments may include one or more of the following features, devices, or aspects, individually or in any combination, across various claim categories and types:
[0161] · A process or apparatus for jointly encoding and decoding a digital video image using a combination of a triangular partitioning mode and a geometric merge mode.
[0162] · A process or apparatus for encoding or decoding a digital video image using a sub-block merged motion field in a triangular and geometric coding mode.
[0163] · A process or apparatus including adding a flag to the syntax or semantics of a signal when using a sub-block merge candidate rather than a regular candidate of the foregoing process or apparatus.
[0164] ·A process or device that includes replacing some regular merge candidates with sub-block merge candidates in the above process or device in terms of syntax or semantics.
[0165] ·A process or device that includes storing the syntax or semantics of at least one motion field in the aforementioned process or device.
[0166] ·A process or device that includes considering all sub-block merge candidates or only sub-block temporal motion vector prediction or only affine prediction in the aforementioned process or device in terms of syntax or semantics.
[0167] ·A bitstream or signal that includes one or more of the described syntax elements or variants thereof.
[0168] ·A bitstream or signal that includes transmitting information generated according to any one of the described embodiments in terms of syntax.
[0169] ·Creation and / or transmission and / or reception and / or decoding according to any one of the described embodiments.
[0170] ·A method, process, device, medium storing instructions, medium storing data, or signal according to any one of the described embodiments.
[0171] ·Inserting a syntax element in signaling, which enables a decoder to determine a tool in a manner corresponding to the encoding method used by an encoder.
[0172] ·Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variants thereof.
[0173] ·A television, set-top box, cellular phone, tablet computer, or other electronic device that performs a transformation method according to any one of the described embodiments.
[0174] ·A television, set-top box, cellular phone, tablet computer, or other electronic device that performs a transformation method according to any one of the described embodiments and determines and displays the resulting image (e.g., using a monitor, screen, or other type of display).
[0175] ·A television, set-top box, cellular phone, tablet computer, or other electronic device that selects, band-limits, or tunes (e.g., using a tuner) a channel according to any one of the described embodiments to receive a signal including an encoded image and performs a transformation method.
[0176] ·A television, set-top box, cellular phone, tablet computer, or other electronic device that receives a signal including an encoded image via air (e.g., using an antenna) and performs a transformation method.
Claims
1. A method, the method comprising: Obtaining a unidirectional triangular candidate from a merge list for encoding a video block; Extracting a unidirectional part of the unidirectional triangular candidate; Obtaining a triangular merge candidate from a candidate sub-block merge list, wherein the sub-block merge list in a triangular partitioning mode is restricted based on the size of the video block; Extracting a unidirectional seed part of the triangular merge candidate; Performing motion compensation using the unidirectional triangular candidate and the triangular merge candidate; Performing weighting on the motion compensation result; Signaling whether a candidate sub-block list or a regular list is used for the motion compensation result; And Encoding the video block using the weighted motion compensation result.
2. A device, the device comprising: A processor configured to: Obtain a unidirectional triangular candidate from a merge list for encoding a video block; Extract a unidirectional part of the unidirectional triangular candidate; Obtain a triangular merge candidate from a candidate sub-block merge list, wherein the sub-block merge list in a triangular partitioning mode is restricted based on the size of the video block; Extract a unidirectional seed part of the triangular merge candidate; Perform motion compensation using the unidirectional triangular candidate and the triangular merge candidate; Perform weighting on the motion compensation result; Signal whether a candidate sub-block list or a regular list is used for the motion compensation result; And Encode the video block using the weighted motion compensation result.
3. A method, the method comprising: Parsing a video bitstream to determine whether to select a unidirectional triangular candidate from a sub-block merge list or in a regular merge list; Obtaining a unidirectional triangular candidate from a sub-block merge list or a regular merge list for decoding a video block, wherein the sub-block merge list in a triangular partitioning mode is restricted based on the size of the video block; Extracting a unidirectional part of the unidirectional triangular candidate; Performing motion compensation using the unidirectional triangular candidate; Performing weighting on the motion compensation result; and Decoding the video block using the weighted motion compensation result.
4. A device, the device comprising: A processor configured to: Parse a video bitstream to determine whether to select a unidirectional triangular candidate from a sub-block merge list or in a regular merge list; Obtain a unidirectional triangular candidate from a sub-block merge list or a regular merge list for decoding a video block, wherein the sub-block merge list in a triangular partitioning mode is restricted based on the size of the video block; Extract a unidirectional part of the unidirectional triangular candidate; Perform motion compensation using the unidirectional triangular candidate; Perform weighting on the motion compensation result; and Decode the video block using the weighted motion compensation result.
5. The method according to claim 1 or 3, wherein the candidates include sub-block temporal motion vector predictors, inherited candidates, and constructed affine model candidates.
6. The method according to claim 1 or 3, wherein the use of the sub-block merge list is restricted based on the size of the coding unit containing the video block, the size of the coding unit, or the coding unit area.
7. The method according to claim 1 or 3, wherein if only affine model candidates are considered during the construction of the unidirectional triangle list, the sub-block merge list uses coding units that are limited to having at least affine neighboring coding units.
8. The method according to claim 1 or 3, wherein the unidirectional triangle candidates are replaced by sub-block candidates, and one sub-block partition is combined with a regular partition.
9. The method according to claim 1 or 3, further comprising storing a 4×4 motion field.
10. The method according to claim 1 or 3, wherein when the candidates refer to the same reference picture list, the diagonal 4×4 sub-block stores the unidirectional motion field of the bottom candidate when the two candidates are from the same merge list, and stores the regular unidirectional motion field when the candidates are from different merge lists.
11. The method according to claim 1 or 3, further comprising storing the affine control point motion vectors associated with the triangle partition when the candidate is from the sub-block merge list and is affine.
12. The apparatus according to claim 2 or 4, wherein the candidates include sub-block temporal motion vector predictors, inherited candidates, and constructed affine model candidates.
13. The apparatus according to claim 2 or 4, wherein the use of the sub-block merge list is restricted based on the size of the coding unit containing the video block, the size of the coding unit, or the coding unit area.
14. The apparatus according to claim 2 or 4, wherein if only affine model candidates are considered during the construction of the unidirectional triangle list, the sub-block merge list uses coding units that are limited to having at least affine neighboring coding units.
15. The apparatus according to claim 2 or 4, wherein the unidirectional triangle candidates are replaced by sub-block candidates, and one sub-block partition is combined with a regular partition.
16. The apparatus according to claim 2 or 4, further comprising storing a 4×4 motion field.
17. The apparatus according to claim 2 or 4, wherein when the candidates refer to the same reference picture list, the diagonal 4×4 sub-block stores the unidirectional motion field of the bottom candidate when the two candidates are from the same merge list, and stores the regular unidirectional motion field when the candidates are from different merge lists.
18. The apparatus according to claim 2 or 4, further comprising storing the affine control point motion vectors associated with the triangle partition when the candidate is from the sub-block merge list and is affine.
19. An apparatus, the apparatus comprising: The apparatus according to claim 4; and at least one of the following: (i) an antenna configured to receive a signal, the signal including a video block; (ii) a band limiter configured to limit the received signal to a band including the video block; and (iii) a display configured to display an output representing the video block.
20. A non-transitory computer-readable medium storing instructions that, when executed by a computer, cause the computer to perform the method according to claim 1 or 3.
21. A computer program product comprising instructions which, when executed by a computer, cause the computer to perform the method according to claim 1 or 3.