Method and apparatus for video encoding and decoding for triangle prediction
By employing a triangle prediction unit for geometric partitioning in video encoding and decoding, the method addresses the challenge of efficiently handling high-resolution video data, achieving improved encoding efficiency and video quality.
Patent Information
- Application Number
- JP2023113440
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-03-13
- Filing Date
- 2023-07-11
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2040-03-12
AI Technical Summary
Existing video encoding and decoding technologies face challenges in efficiently compressing and decompressing high-quality video data, particularly as video resolutions advance from HD to 4K and 8K, leading to increased data volumes and computational complexity.
The implementation of a method for video encoding and decoding that utilizes a triangle prediction unit, which partitions video images into geometric shapes for improved motion compensation prediction, thereby enhancing encoding efficiency and video quality.
This approach allows for more efficient compression and decompression of video data, maintaining high image quality even at high resolutions, by optimizing prediction processes through geometric partitioning.
Smart Images

Figure 0007700175000001 
Figure 0007700175000002 
Figure 0007700175000003
Abstract
Description
Cross - reference to related applications
[0001] This application claims priority to U.S. Provisional Application No. 62 / 817,537, filed on March 12, 2019, with the title "Video Encoding and Decoding by Triangle Prediction" and U.S. Provisional Application No. 62 / 817,852, filed on March 13, 2019, with the title "Video Encoding and Decoding by Triangle Prediction", and the entire specifications of these patent applications are incorporated herein by reference.
Technical Field
[0002] This application generally relates to video encoding, decoding, and compression, and in particular, but not limited to, methods and apparatuses for motion - compensated prediction by a triangle prediction unit (i.e., a special case of a geometric partition prediction unit) in video encoding and decoding.
Background Art
[0003] Various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, video game consoles, smartphones, video conferencing devices, and video streaming devices support digital video. These electronic devices receive, transmit, encode, decode, and store digital video data by performing video compression / decompression. Digital video devices implement video encoding and decoding technologies described in standards defined by Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), MPEG - 2, MPEG - 4, ITU - T H.263, ITU - T H.264 / MPEG - 4, Part 10, Advanced Video Coding (AVC), ITU - T H.265 / High Efficiency Video Coding (HEVC), and extensions of such standards.
[0004] Video encoding and decoding generally utilize prediction methods (e.g., inter prediction, intra prediction) based on the redundancy present in video images or sequences. One of the important goals of video encoding and decoding technology is to compress video data into a lower bitrate form while avoiding or minimizing the degradation of video quality. As evolving video services become available, encoding and decoding technologies with better encoding and decoding efficiency are needed.
[0005] Video compression typically includes performing spatial (intra-frame) prediction and / or temporal (inter-frame) prediction to reduce or eliminate the redundancy inherent in video data. In block-based video encoding, a video frame is partitioned into one or more slices that contain multiple video blocks called coding tree units (CTUs). Each CTU contains one coding unit (CU) or may be recursively divided into smaller CUs until a predefined minimum CU size is reached. Each CU (also called a leaf CU) contains one or more transform units (TUs) and one or more prediction units (PUs). Each CU can be encoded in intra, inter, or IBC mode. Video blocks within an intra-coded (I) slice in a video frame are encoded using spatial prediction with respect to reference samples in adjacent blocks within the same video frame. Video blocks within an inter-coded (P or B) slice in a video frame use spatial prediction with respect to reference samples in adjacent blocks within the same video frame or temporal prediction with respect to reference samples in other previous and / or future reference video frames.
[0006] In spatial prediction or temporal prediction using previously symbolized reference blocks, such as adjacent blocks, a prediction block for the current video block to be coded is obtained. The process of finding the reference block can be realized by a block matching algorithm. Residual data indicating the pixel difference between the current block to be coded and the prediction block is called a residual block or prediction error. An inter-coded block is coded according to a motion vector indicating a reference block in the reference frame forming the prediction block and the residual block. The process of determining the motion vector is usually called motion estimation. An intra-coded block is coded by an intra-prediction mode and the residual block. For further compression, the residual block is transformed from the pixel domain to a transform domain, such as the frequency domain, resulting in residual transform coefficients that are then quantized. Then, the transform coefficients first arranged and quantized in a two-dimensional matrix are scanned to generate a one-dimensional transform coefficient vector, which is then entropy-coded into a video bitstream to achieve further compression.
[0007] Then, the coded video bitstream is stored in a computer-readable storage medium (e.g., flash memory) and accessed by another electronic device with digital video capabilities, or directly transmitted to this electronic device via wire or wirelessly. And this electronic device, for example, analyzes this coded video bitstream to obtain syntax elements from this bitstream, and reconstructs digital video data from this coded video stream into the original format based on at least a part of the syntax elements obtained from this bitstream, thereby performing video decompression (a process opposite to the above-described video compression), and reproducing this reconstructed digital video data on the display of the electronic device.
[0008] As the quality of digital video advances from high definition to 4K×2K and / or 8K×4K, the amount of video data to be encoded / decoded is increasing exponentially. It has always been a challenge to encode / decrypt video data efficiently while maintaining the image quality of the decoded video data.
[0009] At the Joint Video Experts Team (JVET) meeting, the first draft of Versatile Video Coding (VVC) and the coding method of VVC Test Model 1 (VTM1) were defined. A quadtree with a nested multi-type tree with binary and ternary split coding block structures was determined to be included as the first new coding feature of VVC. Since then, the reference software VTM for executing the encoding / decoding method and the draft VVC decoding process have been developed during the JVET meeting.
Summary of the Invention
[0010] The present disclosure generally describes an example of a technique related to motion compensation prediction by a triangle prediction unit, which is a special case of geometric partition prediction in video encoding / decoding.
[0011] According to a first aspect of the present disclosure, a method for video encoding / decoding is provided, including partitioning a video image into a plurality of coding units (CUs) partitioned into two PUs, at least one of which further includes at least one geometric shape prediction unit (PU); constructing a first merge list including a plurality of candidates each including one or more motion vectors; and obtaining a single prediction merge list for the PU in triangle prediction mode, including a plurality of single prediction merge candidates each including one motion vector of a corresponding candidate in the first merge list.
[0012] According to a second aspect of the present disclosure, there is provided a computing device including a processor and a memory configured to store instructions executable by the processor, wherein when the processor executes the instructions, the video image is partitioned into a plurality of coding units (CUs) partitioned into two PUs, at least one of which further includes a prediction unit (PU) of at least one geometric shape, and a first merge list including a plurality of candidates each including one or more motion vectors is constructed, and a single prediction merge list for the PU is derived in a triangular partitioning mode, the single prediction merge list including a plurality of single prediction merge candidates each including one motion vector of a corresponding candidate in the first merge list.
[0013] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing instructions which, when executed by a processor, cause the processor to partition a video image into a plurality of coding units (CUs) partitioned into two PUs, at least one of which further includes a prediction unit (PU) of at least one geometric shape, construct a first merge list including a plurality of candidates each including one or more motion vectors, and derive a single prediction merge list for the PU in a triangular partitioning mode, the single prediction merge list including a plurality of single prediction merge candidates each including one motion vector of a corresponding candidate in the first merge list.
Brief Description of the Drawings
[0014] A more specific description of examples of the present disclosure is given by reference to specific examples shown in the accompanying drawings. These drawings merely illustrate several examples and, considered not to be limiting in scope, these examples will be described with additional specificity and detail by using the accompanying drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Best Mode for Carrying Out the Invention
[0015] Hereinafter, reference will be made in detail to specific embodiments in which examples are shown in the accompanying drawings. In the following detailed description, a plurality of non-limiting specific details are set forth in order to facilitate understanding of the gist described herein. However, it is obvious to those skilled in the art that various modifications can be realized. For example, it is obvious to those skilled in the art that the gist described herein can be implemented in many types of electronic devices having digital video functions.
[0016] References in this specification to "one embodiment", "an embodiment", "an example", "a certain embodiment", "a certain example" or similar expressions mean that the particular feature, structure or characteristic described is included in at least one embodiment or example. Features, structures, elements or characteristics described in connection with one or some embodiments are applicable to other embodiments as well, unless explicitly stated otherwise.
[0017] Throughout this disclosure, terms such as "first", "second", "third", etc. are used only for the purpose of referring to related elements, such as devices, components, configurations, steps, etc., and do not mean any spatial or chronological order unless the context clearly indicates otherwise. For example, "a first device" and "a second device" refer to two separately formed devices, or two parts, components, or operating states of the same device, and can be arbitrarily named.
[0018] As used herein, terms such as "(if)... then" or "(if)... then", "(if)... and" can be understood to mean "when..." or "in response to..." depending on the context. These terms may not mean that the related limitations or features are conditional or optional when they appear in the claims.
[0019] The terms "module", "sub-module", "circuit", "sub-circuit", "circuit system", "sub-circuit system", "unit" or "sub-unit" include memory (shared, dedicated, or grouped) that stores code or instructions executable by one or more processors. A module may include one or more circuits that do or do not store code or instructions. A module or circuit can include one or more components that are directly or indirectly connected. These components can be physically connected to each other, adjacent to each other, or not.
[0020] A unit or module may be implemented entirely in software, entirely in hardware, or in a combination of hardware and software. In a complete software implementation, for example, a unit or module can include functionally related code blocks or software components that are directly or indirectly linked to each other to perform a particular function.
[0021] FIG. 1 shows a block diagram of an exemplary block-based hybrid video encoder 100 that can be used in combination with many video coding / decoding standards based on block-based processing. In encoder 100, a video frame is partitioned into a plurality of video blocks for processing. For each particular video block, a prediction is formed based on an inter-prediction approach or an intra-prediction approach. In inter-prediction, one or more predictors are formed by motion estimation and motion compensation based on pixels from a previously reconstructed frame. In intra-prediction, a predictor is formed based on the reconstructed pixels in the current frame. Through mode decision, an optimal predictor for predicting the current block can be selected.
[0022] The prediction residual representing the difference between the current video block and its predictor is sent to the conversion circuit 102. Then, for entropy reduction, the conversion coefficients are sent from the conversion circuit 102 to the quantification circuit 104. Next, the quantified coefficients are supplied to the entropy encoding circuit 106 to generate a compressed video bit stream. As shown in FIG. 1, prediction-related information 110 such as video block partition information, motion vectors, reference image indexes, and intra prediction modes from the inter prediction circuit and / or the intra prediction circuit 112 is also supplied via the entropy encoding circuit 106 and stored in the compressed video bit stream 114.
[0023] In the encoder 100, a decoder-related circuit is also required to reconstruct pixels for prediction purposes. First, the prediction residual is reconstructed via inverse quantization 116 and the inverse conversion circuit 118. This reconstructed prediction residual is combined with the block predictor 120 to generate the unfiltered reconstructed pixels of the current video block.
[0024] Spatial prediction (also called "intra prediction") predicts the current video block using pixels from already encoded adjacent blocks (also called reference samples) within the same video frame as the current video block.
[0025] Temporal prediction (also called "inter prediction") predicts the current block using reconstructed pixels from already encoded video images. Temporal prediction reduces the temporal redundancy inherent in the video signal. The temporal prediction signal for a particular CU is typically signaled by one or more motion vectors (MVs) indicating the amount and direction of motion between the current CU and its temporal reference. Also, when multiple reference images are supported, one reference image index is additionally transmitted to identify from which reference image in the reference image storage the temporal prediction signal is.
[0026] After spatial and / or temporal prediction, the intra / inter mode decision circuit 121 in the encoder 100 selects an optimal prediction mode, for example, based on a rate distortion optimization method. Next, the block predictor 120 is subtracted from the current video block; the obtained prediction residual is decorrelated by the transform circuit 102 and the quantization circuit 104. The obtained quantized residual coefficients are inverse quantized by the inverse quantization circuit 116 and inverse transformed by the inverse transform circuit 118 to form a reconstruction residual, which is then added to the prediction block to form the reconstructed signal of this CU. Further, this reconstructed CU is placed in the reference image storage unit of the picture buffer 117, and before being used for the encoding / decoding of future video blocks, an in-loop filter 115 such as a deblocking filter, a sample adaptive offset (SAO), and / or an adaptive in-loop filter (ALF) can be used for this reconstructed CU. To form the output video bitstream 114, the encoding mode (inter or intra), the prediction mode information, the motion information, and the quantized residual coefficients are all sent to the entropy encoding unit 106, and further compressed and packed to form a bitstream.
[0027] For example, the current versions of AVC, HEVC, and VVC provide a deblocking filter. In HEVC, an additional in-loop filter called SAO (sample adaptive offset) is defined to further improve the encoding efficiency. In the current version of the VVC standard, another additional in-loop filter called ALF (adaptive loop filter) is being actively studied and is likely to be included in the final standard.
[0028] These in-loop filter operations are selectable. Performing these operations improves the encoding efficiency and visual quality. They can be turned off according to the decision by the encoder 100 to save computational complexity.
[0029] In addition, when these filter options are turned on by the encoder 100, intra prediction is usually based on the unfiltered reconstructed pixels, while inter prediction is based on the filtered reconstructed pixels.
[0030] FIG. 2 is a block diagram showing an exemplary block-based video decoder 200 that can be used in combination with many video encoding / decoding standards. This decoder 200 is similar to the reconstruction-related part present in the encoder 100 of FIG. 1. In the decoder 200, the input video bitstream 201 is first decoded through entropy decoding 202 to derive the quantized coefficient levels and prediction-related information. Next, the quantized coefficient levels are processed through inverse quantization 204 and inverse transform 206 to obtain the reconstructed prediction residuals. The block predictor mechanism implemented in the intra / inter mode selection unit 212 is configured to perform intra prediction 208 or motion compensation 210 based on the decoded prediction information. The reconstructed prediction residuals from the inverse transform 206 and the prediction output generated by the block predictor mechanism are added by the adder 214 to obtain a set of unfiltered reconstructed pixels.
[0031] Before the reconstructed block is stored in the image buffer 213 that functions as a reference image memory unit, it can further pass through the in-loop filter 209. The reconstructed video in the image buffer 213 can be sent to drive a display device or used to predict future video blocks. When the in-loop filter 209 is turned on, a filtering operation is performed on these reconstructed pixels to derive the final reconstructed video output 222.
[0032] The above-described video encoding / decoding standards such as VVC, JEM, HEVC, and MPEG-4, Part10 are conceptually similar. For example, they all use block-based processing. The block partitioning schemes in some of the standards will be described in detail below.
[0033] HEVC is based on a motion-compensated transform coding method using hybrid blocks. The basic unit for compression is called a coding tree unit (CTU). For the 4:2:0 chroma format, the maximum CTU size is defined as 64×64 luma pixels and two 32×32 chroma pixel blocks. Each CTU contains one coding unit (CU), or is recursively divided into four smaller CUs until it reaches a predefined minimum CU size. Each CU (also called a leaf CU) contains a tree of one or more prediction units (PUs) and one or more transform units (TUs).
[0034] Generally, except for monochromatic content, a CTU contains one luma coding tree block (CTB) and two corresponding chroma CTBs, a CU contains one luma coding block (CB) and two corresponding chroma CBs, a PU contains one luma prediction block (PB) and two corresponding chroma PBs, and a TU may contain one luma transform block (TB) and two corresponding chroma TBs. However, the minimum TB size is 4×4 for both luma and chroma (i.e., 2×2 chroma TBs are not supported in the 4:2:0 color format), and each intra-chroma CB always has only one intra-chroma PB regardless of the number of intra-luma PBs in the corresponding intra-luma CB.
[0035] In an intra CU, the luminance CB can be predicted by one or four luminance PBs, and two chrominance CBs are usually predicted by one chrominance PB each. Here, each luminance PB has one intra luminance prediction mode, and the two chrominance PBs share one intra chrominance prediction mode. Furthermore, in an intra CU, the TB size cannot be larger than the PB size. For each PB, intra prediction is applied to predict the samples of each TB within the PB from the neighboring reconstructed samples of the TB. For each PB, in addition to 33 directional intra prediction modes, the DC mode and the planar mode are also supported to predict flat regions and gradually changing regions respectively.
[0036] For each inter PU, it is possible to select one from three prediction modes including inter, skip, and merge. Generally speaking, the motion vector competition (MVC) method is introduced to select motion candidates from a specific set of candidates including spatial and temporal motion candidates. By multiple references to motion estimation, the optimal reference in two possible reconstructed reference image lists (i.e., List0 and List1) can be found. In the inter mode (referred to as the AMVP mode representing advanced motion vector prediction), the inter prediction indicator (List0, List1, or bidirectional prediction), the reference index, the motion candidate index, the motion vector difference (MVD), and the prediction residual are transmitted. For the skip mode and the merge mode, only the merge index is transmitted, and the current PU inherits the inter prediction indicator, the reference index, and the motion vector from the adjacent PU pointed to by the encoded merge index. In the case of a skip-encoded CU, the residual signal is also omitted.
[0037] The Common Exploration Test Model (JEM) is built on top of the HEVC Test Model. The basic encoding and decoding flow of HEVC has not been changed in JEM. However, the design elements of the most important modules, including block structure, intra and inter prediction, residue transformation, loop filter, and entropy encoding / decoding modules, have been somewhat changed, and additional encoding tools have been added. JEM includes the following new encoding features.
[0038] In HEVC, a CTU is divided into CUs by a quadtree structure shown as an encoding tree to adapt to various local characteristics. The decision on whether to encode an image region using inter-image (temporal) or intra-image (spatial) prediction is made at the CU level. Each CU can be further divided into one, two, or four PUs depending on the PU split type. Within one PU, the same prediction process is applied, and related information is sent to the decoder based on the PU. After obtaining a residual block by applying a prediction process based on the PU split type, the CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the encoding tree of the CU. One of the main features of the HEVC structure is the existence of multiple partitioning concepts including CUs, PUs, and TUs.
[0039] Figure 3 is a schematic diagram showing a quadtree plus binary tree (QTBT) structure according to an embodiment of the present disclosure.
[0040] The QTBT structure removes the concept of multiple partition types, that is, it removes the separation of the concepts of CU, PU, and TU, and supports more flexibility in the CU partition shape. In the QTBT block structure, the CU can take on a square or rectangular shape. As shown in Figure 3, the Coding Tree Unit (CTU) is first partitioned by a quadtree structure. The quadtree leaf nodes can be further partitioned by a binary tree structure. There are two types of binary tree partitions: symmetric horizontal partition and symmetric vertical partition. The binary tree leaf nodes are called Coding Units (CUs), and their segmentation is used for prediction and transformation processing without further partitioning. This means that the CU, PU, and TU have the same block size in the QTBT coding block structure. In JEM, the CU consists of coding blocks (CBs) of different color components. For example, in the case of P slices and B slices in the 4:2:0 chroma format, one CU may contain one luminance CB and two chroma CBs. The CU may consist of a single-component CB. For example, in the case of I slices, one CU may contain only one luminance CB, or only two chroma CBs.
[0041] The following parameters are defined for the QTBT partitioning method. - CTU size: It is the size of the quadtree root node, which is the same concept as in HEVC; - MinQTSize: The minimum allowable quadtree leaf node size; - MaxBTSize: The maximum allowable binary tree root node size; - MaxBTDepth: The maximum allowable binary tree depth; - MinBTSize: The minimum allowable binary tree leaf node size.
[0042] In an example of the QTBT partitioning structure, the CTU size has 128×128 luminance samples with two corresponding 64×64 chrominance sample blocks (4:2:0 chrominance format), MinQTSize is 16×16, MaxBTSize is 64×64, MinBTSize (both width and height) is 4×4, and MaxBTDepth is set to 4. The quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from 16×16 (i.e., MinQTSize) to 128×128 (i.e., CTU size). When the quadtree leaf node is 128×128, since its size exceeds MaxBTSize (i.e., 64×64), the quadtree leaf node is not further divided by a binary tree. Otherwise, the quadtree leaf node can be further partitioned by a binary tree. Therefore, the quadtree leaf node is also a binary tree root node and has a binary tree depth of 0. When the binary tree depth reaches MaxBTDepth (i.e., 4), no further division is considered. When the binary tree node has a width equal to MinBTSize (i.e., 4), no further horizontal division is considered. Similarly, when the binary tree node has a height equal to MinBTSize, no further vertical division is considered. The quadtree leaf nodes are further processed by prediction processing and transformation processing without further partitioning. In JEM, the maximum CTU size is 256×256 luminance samples.
[0043] FIG. 3 shows an example of block partitioning by the QTBT method and the corresponding tree display. Solid lines indicate quadtree partitioning, and dotted lines indicate binary tree partitioning. As shown in FIG. 3, the coding tree unit (CTU) 300 is first partitioned by a quadtree structure, and three of the four quadtree leaf nodes 302, 304, 306, 308 are further partitioned by either a quadtree structure or a binary tree structure. For example, the quadtree leaf node 306 is further partitioned by quadtree partitioning. The quadtree leaf node 304 is further partitioned into two leaf nodes 304a and 304b by binary tree partitioning. Also, the quadtree leaf node 302 is further partitioned by binary tree partitioning. At each split (i.e., non-leaf) node of the binary tree, one flag indicating the split type (i.e., horizontal or vertical) is signaled, where 0 indicates a horizontal split and 1 indicates a vertical split. For example, in the case of the quadtree leaf node 304, 0 is signaled to indicate a horizontal split, and in the case of the quadtree leaf node 302, 1 is signaled to indicate a vertical split. In quadtree partitioning, since the block is always split both horizontally and vertically to generate four sub-blocks of the same size, there is no need to indicate the split type.
[0044] Also, the QTBT method supports the ability to have separate QTBT structures for luminance and chrominance. Currently, for P slices and B slices, the luminance CTB and chrominance CTB within one CTU share the same QTBT structure. However, for I slices, the luminance CTB is partitioned into CUs by one QTBT structure, and the chrominance CTB is partitioned into chrominance CUs by another QTBT structure. This means that a CU within an I slice consists of an encoded block of the luminance component or encoded blocks of two chrominance components, while a CU within a P slice or B slice consists of encoded blocks of all three color components.
[0045] At the meeting of the Joint Video Experts Team (JVET), JVET defined the first draft of Versatile Video Coding (VVC) and the VVC Test Model 1 (VTM1) coding method. It was determined that the quadtree with a nested multi-type tree of binary and ternary split coding block structures is included as the first new coding feature of VVC.
[0046] In VVC, due to the picture partitioning structure, the input video is divided into blocks called Coding Tree Units (CTUs). A CTU is divided into Coding Units (CUs) by a quadtree with a nested multi-type tree structure, where the leaf coding units (CUs) define regions that share the same prediction mode (e.g., intra or inter). Here, the term "unit" defines a region of the picture that includes all components. The term "block" is for defining a region that includes a specific component (e.g., luminance), and may be different in spatial position when a chroma sampling format such as 4:2:0 is considered. Partitioning of an image into CTUs
[0047] Figure 4 is a schematic diagram showing an example of an image divided into CTUs according to an embodiment of the present disclosure.
[0048] In VVC, an image is divided into a series of CTUs, where the concept of CTU is the same as that of the CTU in HEVC. For an image with three sample arrays, a CTU consists of an N×N luminance sample block and two corresponding chroma sample blocks. Figure 4 shows an example of an image 400 divided into CTU402.
[0049] The maximum allowable size of the luminance block in a CTU is specified as 128×128 (however, the maximum size of the luminance transform block is 64×64). CTU partitioning by a tree structure
[0050] Figure 5 is a schematic diagram showing the multi-type tree split mode according to an embodiment of the present disclosure.
[0051] In HEVC, a CTU is divided into CUs by a quadtree structure shown as an encoding tree so as to adapt to various local characteristics. Encoding an image region using either inter-picture (temporal) prediction or intra-picture (spatial) prediction is determined at the leaf CU level. Each leaf CU can be further divided into one, two, or four PUs depending on the PU partition type. Within one PU, the same prediction process is applied and related information is sent to the decoder based on the PU. After obtaining a residual block by applying a prediction process based on the PU partition type, a leaf CU can be partitioned into transform units (TUs) according to another quadtree structure similar to the encoding tree of this CU. One of the main features of the HEVC structure is that there are multiple partitioning concepts including CUs, PUs, and TUs.
[0052] In VVC, a quadtree with a nested multi-type tree using binary and ternary segmentation structures replaces the concept of multiple partition unit types, that is, it removes the separation of the concepts of CU, PU, and TU (except for the case where the size of the maximum transform length is too large for the CU), and supports more flexibility in the CU partition shape. In the coding tree structure, the CU can take a square or rectangular shape. The coding tree unit (CTU) is first partitioned by a quadtree structure. Next, the leaf nodes of this quadtree can be further partitioned by a multi-type tree structure. As shown in FIG. 5, the multi-type tree structure has four split types: vertical binary split 502 (SPLIT_BT_VER), horizontal binary split 504 (SPLIT_BT_HOR), vertical ternary split 506 (SPLIT_TT_VER), and horizontal ternary split 508 (SPLIT_TT_HOR). The leaf nodes of the multi-type tree are called coding units (CUs), and as long as the maximum transform length of the CU is not too large, this segmentation is used for prediction processing and transform processing without further partitioning. This means that in most cases, in a quadtree with a nested multi-type tree coding block structure, the CU, PU, and TU have the same block size. The exception is when the maximum transform support length is smaller than the width or height of the color components of the CU. In VTM1, a CU consists of coding blocks (CBs) of different color components. For example, one CU contains one luminance CB and two chrominance CBs (except when the video is monochrome, that is, there is only one color component). Partitioning of a CU into multiple prediction units
[0053] In VVC, for each CU partitioned based on the above structure, the prediction of the block content can be performed for the entire CU block or in a sub-block manner described below. Such an operation unit of prediction is called a prediction unit (or PU).
[0054] In the case of intra prediction (or intra-frame prediction), usually, the size of the PU is the same as that of the CU. In other words, the prediction is performed for the entire CU block. In the case of inter prediction (or inter-frame prediction), the size of the PU can be made smaller than that of the CU. In other words, the CU may be divided into a plurality of PUs to be predicted.
[0055] Examples where the PU size is smaller than the CU size include the affine prediction mode, the advanced temporal level motion vector prediction (ATMVP) mode, and the triangular prediction mode.
[0056] In the affine prediction mode, it is possible to divide the CU into 4×4 PUs that are the objects of prediction. A motion vector can be derived for each 4×4 PU, and motion compensation can be performed accordingly for this 4×4 PU. In the ATMVP mode, it is possible to divide the CU into 8×8 PUs that are one or more objects of prediction. A motion vector can be derived for each 8×8 PU, and motion compensation can be performed accordingly for this 8×8 PU. In the triangular prediction mode, it is possible to divide the CU into two triangular-shaped prediction units. A motion vector is derived for each PU, and motion compensation is performed accordingly. The triangular prediction mode is supported in inter prediction. The details of the triangular prediction mode are shown as follows. Triangle prediction mode (or triangle partitioning mode)
[0057] FIG. 6 is a schematic diagram showing the division of a CU into triangular prediction units according to an embodiment of the present disclosure.
[0058] The concept of the triangular prediction mode introduces triangular partitions for motion compensation prediction. The triangular prediction mode is also referred to as the triangular prediction unit mode or the triangular partition mode. As shown in FIG. 6, CU602, or 604, is divided into two triangular prediction units PU1 and PU2 in the diagonal or anti-diagonal direction (i.e., divided from the upper left corner to the lower right corner as shown in CU602, or divided from the upper right corner to the lower left corner as shown in CU604). Each triangular prediction unit within the CU is inter-predicted using its own single prediction motion vector and reference frame index derived from a single prediction candidate list. After predicting these triangular prediction units, an adaptive weighting process is performed on the diagonal edges. Next, the transformation process and quantization process are applied to the entire CU. Note that this mode is only applicable to the current VVC skip mode and merge mode. As shown in FIG. 6, the CU is shown as a square block, but the triangular prediction mode may also be applicable to non-square (i.e., rectangular) shaped CUs.
[0059] The single prediction candidate list contains one or more candidates, and each candidate can be a motion vector. Therefore, throughout this disclosure, the terms "single prediction candidate list", "single prediction motion vector candidate list", and "single prediction merge list" are applied interchangeably, and the terms "single prediction merge candidate list" and "single prediction motion vector" can be applied interchangeably. Single prediction motion vector candidate list
[0060] FIG. 7 is a schematic diagram showing the positions of adjacent blocks according to an embodiment of the present disclosure.
[0061] In one example, the single prediction motion vector candidate list can include from 2 to 5 single prediction motion vector candidates. In another example, other numbers are also possible. It is derived from adjacent blocks. The single prediction motion vector candidate list is derived from 7 adjacent blocks including 5 spatially adjacent blocks (1 to 5) and 2 blocks at the same temporal position (6 to 7) as shown in FIG. 7. The motion vectors of these 7 adjacent blocks are collected in the first merge list. Next, a single prediction candidate list is formed based on the motion vectors of the first merge list in a predetermined order. Based on that order, the single prediction motion vector from the first merge list is first put into the single prediction motion vector candidate list, then the reference picture list 0 or L0 motion vector of the dual prediction motion vector, and the reference picture list 1 or L1 motion vector of the dual prediction motion vector, and then the averaged motion vector of the L0 and L1 motion vectors of the dual prediction motion vector is successively put into the list. At that point, if the number of candidates is less than the target number (5 in the current VVC), zero motion vectors are added to the list to meet the target number.
[0062] For each of the triangular PUs, a predictor is derived based on its motion vector. Note that the derived predictor covers a larger area than the actual triangular PU so that there is an overlapping area of the two predictors along the shared diagonal edge of the two triangular PUs. To derive the final prediction of the CU, a weighting process is applied to the diagonal edge area between the two predictors. Currently, the weighting factors applied to the luminance samples and the chrominance samples are {7 / 8, 6 / 8, 5 / 8, 4 / 8, 3 / 8, 2 / 8, 1 / 8} and {6 / 8, 4 / 8, 2 / 8}, respectively. Normal merge mode motion vector candidate list
[0063] According to the current VVC, in the normal merge mode where the entire CU is predicted without being divided into a plurality of PUs, the motion vector candidate list or the merge candidate list is created by a different procedure from the case of the triangular prediction mode.
[0064] First, as shown in FIG. 8 which is a schematic diagram showing the positions of spatial merge candidates according to an embodiment of the present disclosure, spatial motion vector candidates are selected based on motion vectors from adjacent blocks. In deriving the spatial merge candidates for the current block 802, up to four merge candidates are selected from candidates at positions as shown in FIG. 8. The order of derivation is A1→B1→B0→A0→(B2). The position B2 is considered only when any of the PUs at positions A1, B1, B0, A0 is unavailable or is intra-coded / decoded.
[0065] Next, temporal merge candidates are derived. In deriving the temporal merge candidates, scaled motion vectors are derived based on PUs at the same position belonging to an image with the smallest difference in Picture Order Count (POC) from the current image within a specific reference picture list. The reference picture list used for deriving PUs at the same position is explicitly signaled in the slice header. The scaled motion vectors for the temporal merge candidates are obtained as indicated by the dotted line in FIG. 9 showing motion vector scaling for the temporal merge candidates according to an embodiment of the present disclosure. The scaled motion vectors for the temporal merge candidates are scaled from the motion vectors of PUcol_PU at the same position using the POC distances tb and td, where tb is defined as the POC difference between the reference picture curr_ref of the current picture curr_pic and the current picture curr_pic, and td is defined as the POC difference between the reference picture col_ref of the picture col_pic at the same position and the picture col_pic at the same position. The reference picture index for the temporal merge candidates is set to 0. The actual implementation of the scaling process is described in the HEVC draft specification. In the case of a B slice, two motion vectors are obtained, one for the reference picture list 0 and the other for the reference picture list 1, and combined to create a bi-prediction merge candidate.
[0066] FIG. 10 is a schematic diagram showing candidate positions for temporal merge candidates according to an embodiment of the present disclosure.
[0067] The positions of PUs at the same location are selected from two candidate positions C3 and H, as shown in FIG. 10. If the PU at position H is unavailable, intra-coded / decoded, or outside the current CTU, then position C3 is used for deriving the temporal merge candidate. Otherwise, position H is used for deriving the temporal merge candidate.
[0068] After both the spatial motion vector and the temporal motion vector are inserted into the merge candidate list as described above, merge candidates based on history are added. So-called merge candidates based on history include motion vectors from previously coded / decoded CUs that are held in an individual motion vector list and managed based on specific rules.
[0069] After the history-based candidates are inserted, if the merge candidate list is not full, pairwise-averaged motion vector candidates are further added to this list. As the name indicates, this type of candidate is composed of averaging the candidates already in the current list. More specifically, based on a certain order or rule, two candidates are taken from the merge candidate list one by one at a time, and the averaged motion vector of the two candidates is added to the current list.
[0070] After the pairwise-averaged motion vector is inserted, if the merge candidate list is still not full, zero motion vectors are added until the list is full. Creation of a first merge list for triangle prediction by a normal merge list creation process
[0071] In the current VVC, the triangular prediction mode shares some common points with the normal merge prediction mode in the overall procedure of forming predictors. For example, in both prediction modes, it is necessary to create a merge list based on at least the adjacent spatial motion vectors of the current CU and the motion vectors at the same position. On the other hand, the triangular prediction mode also has some differences from the normal merge prediction mode.
[0072] For example, in both the triangular prediction mode and the normal merge prediction mode, it is necessary to create a merge list, but the detailed procedures for obtaining such a list are different.
[0073] These differences require additional logic, which incurs additional costs for codec implementation. The procedures and logic for creating the merge list can be integrated and shared between the triangular prediction mode and the normal merge prediction mode.
[0074] In one example, when forming a unidirectional prediction (also called single prediction) merge list for the triangular prediction mode, before adding a new motion vector to the merge list, the new motion vector is completely truncated with respect to the motion vectors already in the list. In other words, the new motion vector is compared with each motion vector already in the single prediction merge list and added to this list only if it is different from all the motion vectors in the merge list. Otherwise, the new motion vector is not added to this list.
[0075] According to an example of the present disclosure, in the triangular prediction mode, the unidirectional prediction merge list can be obtained or created from a normal merge mode motion vector candidate list called the normal merge list.
[0076] More specifically, in order to create a merge candidate list for the triangular prediction mode, first, a first merge list is created based on the merge list creation process for normal merge prediction. The first merge list includes a plurality of candidates each of which is a motion vector. Next, using the motion vectors in this first merge list, a unidirectional prediction merge list for the triangular prediction mode is further created or derived.
[0077] Note that for the first merge list created in this case, it is possible to select a list size different from the list size for the general merge mode or the normal merge mode. In an example of the present disclosure, the first merge list has the same size as the list for the general merge mode. In another example of the present disclosure, the created first merge list has a list size different from the list for the general merge mode. Creation of a unidirectional prediction merge list from the first merge list
[0078] According to an example of the present disclosure, the unidirectional prediction merge list for the triangular prediction mode can be created or derived from the first merge list based on one of the following methods.
[0079] In an example of the present disclosure, in order to create or derive this unidirectional prediction merge list, first, the predicted list 0 motion vector of the candidates in the first merge list is checked and selected for the unidirectional prediction merge list. After this process, if this unidirectional prediction merge list is not full (for example, the number of candidates in this list is still less than the target number), the predicted list 1 motion vector of the candidates in the first merge list is checked and selected for the unidirectional prediction merge list. If the unidirectional prediction merge list is still not full, the predicted list 0 zero vector is added to this unidirectional prediction merge list. If the unidirectional prediction merge list is still not full, the predicted list 1 zero vector is added to this unidirectional prediction merge list.
[0080] In another example of the present disclosure, for each candidate in the first merge list, its prediction list 0 motion vector and prediction list 1 motion vector are added to the unidirectional prediction merge list in an interleaved manner. More specifically, for each candidate in the first merge list, if the candidate is a unidirectional prediction motion vector, it is directly added to the unidirectional prediction merge list. Otherwise, if the candidate is a bidirectional prediction motion vector in the first merge list, first its prediction list 0 motion vector is added to the unidirectional prediction merge list, and then its prediction list 1 motion vector is added. After all motion vector candidates in the first merge list have been checked and added, if the unidirectional prediction merge list is not yet full, it is possible to add a unidirectional prediction zero motion vector. For example, for each reference frame index, it is possible to separately add the prediction list 0 zero motion vector and the prediction list 1 zero motion vector to the unidirectional prediction merge list until this list is full.
[0081] In yet another example of the present disclosure, first, the unidirectional prediction motion vectors from the first merge list are selected into the unidirectional prediction merge list. If the unidirectional prediction merge list is not full after this process, for each bidirectional prediction motion vector in the first merge list, first its prediction list 0 motion vector is added to the unidirectional prediction merge list, and then its prediction list 1 motion vector is added. After this process, if the unidirectional prediction merge list is still not full, it is possible to add a unidirectional prediction zero motion vector. For example, for each reference frame index, it is possible to separately add the prediction list 0 zero motion vector and the prediction list 1 zero motion vector to the unidirectional prediction merge list until this list is full.
[0082] In the above description, when adding a unidirectional prediction motion vector to the unidirectional prediction merge list, in order to confirm that the newly added motion vector is different from the motion vectors already in the unidirectional prediction merge list, it is possible to execute a motion vector truncation process. Such a motion vector truncation process can also be partially executed to reduce complexity. For example, check the newly added motion vector only for some, rather than all, of the motion vectors already in the unidirectional prediction merge list. In an extreme case, the motion vector truncation (i.e., the motion vector comparison operation) is not executed in this process. Creation of a unidirectional prediction merge list from the first merge list based on an image prediction configuration
[0083] In an example of the present disclosure, a single prediction merge list can be adaptively created based on whether the current image uses backward prediction. For example, the single prediction merge list can be created in different ways depending on whether the current image uses backward prediction. That all Picture Order Count (POC) values of the reference images are not greater than the POC value of the current image means that the current image does not use backward prediction.
[0084] In one example of the present disclosure, if the current image does not use backward prediction or it is determined that the current image does not use backward prediction, first, the candidate prediction list 0 motion vectors in the first merge list are checked and selected into the uni-directional prediction merge list, and then the candidate prediction list 1 motion vectors of those are selected into this uni-directional prediction merge list. Also, if this uni-directional prediction merge list is not yet full, it is possible to add a single prediction zero motion vector. Otherwise, if the current image uses backward prediction, the prediction list 0 motion vectors and prediction list 1 motion vectors of each candidate in the first merge list are checked and can be selected into the uni-directional prediction merge list in an interleaved manner as described above, that is, the prediction list 0 motion vector of the first candidate in the first merge list is added, then the prediction list 1 motion vector of the first candidate is added, then the prediction list 0 motion vector of the second candidate is added, followed by the prediction list 1 motion vector of the second candidate, and so on. At the end of this process, if the uni-directional prediction merge list is not yet full, it is possible to add a single prediction zero vector.
[0085] In another example of the present disclosure, when the current image does not use backward prediction, first, the prediction list 1 motion vectors of the candidates in the first merge list are checked and selected into the uni-directional prediction merge list, and the prediction list 0 motion vectors of those candidates are selected into the uni-directional prediction merge list. Also, if the uni-directional prediction merge list is not yet full, it is possible to add a single prediction zero motion vector. Otherwise, when the current image uses backward prediction, the prediction list 0 motion vectors and the prediction list 1 motion vectors of each candidate in the first merge list are checked and can be selected into the uni-directional prediction merge list in an interleaved manner as described above, that is, the prediction list 0 motion vector of the first candidate in the first merge list is added, then the prediction list 1 motion vector of the first candidate is added, then the prediction list 0 motion vector of the second candidate is added, followed by the prediction list 1 motion vector of the second candidate, and so on. At the end of the process, if the uni-directional prediction merge list is not yet full, it is possible to add a single prediction zero vector.
[0086] In yet another example of the present disclosure, when the current image does not use backward prediction, first, only the prediction list 0 motion vectors of the candidates in the first merge list are checked and selected into the uni-directional prediction merge list. Also, if the uni-directional prediction merge list is not yet full, it is possible to add a single prediction zero motion vector. Otherwise, when the current image uses backward prediction, the prediction list 0 motion vectors and the prediction list 1 motion vectors of each candidate in the first merge list are checked and selected into the uni-directional prediction merge list in an interleaved manner as described above, that is, the prediction list 0 motion vector of the first candidate in the first merge list is added, then the prediction list 1 motion vector of the first candidate is added, then the prediction list 0 motion vector of the second candidate is added, followed by the prediction list 1 motion vector of the second candidate, and so on. At the end of the process, if the uni-directional prediction merge list is not yet full, it is possible to add a single prediction zero vector.
[0087] In yet another example of the present disclosure, when the current image does not use backward prediction, only the predicted list 1 motion vector of the candidate in the first merge list is first checked and selected into the uni-directional prediction merge list. Also, when the uni-directional prediction merge list is not yet full, it is possible to add a single predicted zero motion vector. Otherwise, when the current image uses backward prediction, the predicted list 0 motion vector and the predicted list 1 motion vector of each candidate in the first merge list are checked and can be selected into the uni-directional prediction merge list in an interleaved manner as described above, that is, the predicted list 0 motion vector of the first candidate in the first merge list is added, then the predicted list 1 motion vector of the first candidate is added, then the predicted list 0 motion vector of the second candidate is added, followed by the predicted list 1 motion vector of the second candidate, and so on. At the end of the process, when the uni-directional prediction merge list is not yet full, it is possible to add a single predicted zero vector. Use of the first merge list for triangle prediction without creation of a unidirectional prediction merge list
[0088] In the above example, the uni-directional prediction merge list for triangular prediction is created by selecting motion vectors from the first merge list into the uni-directional prediction merge list. However, in practice, the method can be implemented in different forms regardless of whether the uni-directional prediction (or single prediction) merge list is physically formed. In one example, the first merge list can be directly used without physically creating the uni-directional prediction merge list. For example, the list 0 and / or list 1 motion vectors of each candidate in the first merge list can be easily indexed based on a certain order and directly accessed from the first merge list.
[0089] For example, the first merge list can be obtained from a decoder or other electronic device / component. In other examples, after creating a first merge list that includes a plurality of candidates, each of which is one or more motion vectors, based on the merge list creation process for normal merge prediction, a unidirectional prediction merge list is not created. Instead, a unidirectional merge candidate for the triangular prediction mode is derived using a predefined index list that includes a plurality of reference indexes, each of which is a reference to the motion vector of a candidate in the first merge list. The index list can be regarded as the representation of a unidirectional prediction merge list for triangular prediction, and the unidirectional prediction merge list includes at least a subset of the candidates in the first merge list corresponding to the reference indexes. Note that the order of the indexes can follow any of the selection orders described in the example of creating a unidirectional prediction merge list. In practice, such an index list can be realized in various forms. For example, it may be explicitly realized as a list. In other examples, it may be realized or obtained by a specific logic or program function without explicitly creating any list.
[0090] In certain examples of the present disclosure, the index list can be adaptively determined based on whether the current image uses backward prediction. For example, the reference indexes in the index list may be arranged according to whether the current image uses backward prediction, that is, based on the comparison result between the picture order count (POC) of the current image and the POC of the reference image. If the POC values of all reference images are below the POC value of the current image, it means that the current image is not using backward prediction.
[0091] In one example of the present disclosure, when the current picture does not use backward prediction, the candidate prediction list 0 motion vectors in the first merge list are used as unidirectional prediction merge candidates indexed according to the same index order as in the first merge list. That is, when it is determined that the POC of the current picture is greater than each of the POCs of the reference pictures, the reference index is arranged according to the same order as the list 0 motion vectors of the candidates in the first merge list. Otherwise, when the current picture uses backward prediction, the list 0 motion vectors and list 1 motion vectors of each candidate in the first merge list are used as unidirectional prediction merge candidates indexed based on an interleaved pattern of the list 0 motion vector of the first candidate in the first merge list, the list 1 motion vector of the first candidate, the list 0 motion vector of the second candidate, then the list 1 motion vector of the second candidate, and so on. That is, when it is determined that the POC of the current picture is less than at least one of the POCs of the reference pictures, the reference index is arranged according to the interleaved pattern of the list 0 motion vectors and list 1 motion vectors of the candidates that are bidirectional prediction motion vectors in the first merge list. When the candidates in the first merge list are unidirectional motion vectors, the zero motion vector is indexed as an unidirectional prediction merge candidate following the motion vector of that candidate. This provides two unidirectional motion vectors as unidirectional prediction merge candidates regardless of whether each candidate in the first merge list is a bidirectional prediction motion vector or an unidirectional prediction motion vector when the current picture uses backward prediction.
[0092] In another example of the present disclosure, when the current image does not use backward prediction, the prediction list 0 motion vectors of the candidates in the first merge list are used as unidirectional prediction merge candidates indexed according to the same index order as in the first merge list. Otherwise, when the current image uses backward prediction, the list 0 motion vectors and list 1 motion vectors of each candidate in the first merge list are used as unidirectional prediction merge candidates indexed based on the above-described interleaving method of the list 0 motion vector of the first candidate in the first merge list, the list 1 motion vector of the first candidate, the list 0 motion vector of the second candidate, then the list 1 motion vector of the second candidate, and so on. When the candidate in the first merge list is a unidirectional motion vector, a specific motion offset is added to the motion vector, and it is indexed as a unidirectional prediction merge candidate following the motion vector of this candidate.
[0093] Therefore, when the candidate in the first merge list is a unidirectional motion vector, if it is determined that the POC of the current image is smaller than at least one of the POCs of the reference images, the reference index is arranged according to the interleaving method of the motion vectors of each candidate in the first merge list, and the zero motion vector or the sum of this motion vector and the offset.
[0094] In the above process, when checking for new motion vectors to be added to the unidirectional prediction merge list, truncation can be performed either completely or partially. When performed partially, it means that the new motion vector is compared not with all but only with some of the motion vectors already in the single prediction merge list. In an extreme case, motion vector truncation (i.e., motion vector comparison processing) is performed in this process.
[0095] Also, when forming a single prediction merge list, it is possible to adaptively perform motion vector truncation based on whether the current image uses backward prediction. For example, in an example of the present disclosure regarding index list determination based on an image prediction configuration, if the current image does not use backward prediction, the motion vector truncation process is performed completely or partially. If the current image is using backward prediction, the motion vector truncation process is not performed. Selection of a single prediction merge candidate for the triangle prediction mode
[0096] In addition to the above-described examples, other methods for creating a single prediction merge list or selecting a single prediction merge candidate are disclosed.
[0097] In an example of the present disclosure, when a first merge list for the normal merge mode is created, it is possible to select a single prediction merge candidate for triangular prediction according to the following rules.
[0098] For a motion vector candidate in the first merge list, only one of its list 0 motion vector and list 1 motion vector is used for triangular prediction;
[0099] For a specific motion vector candidate in the first merge list, if the merge index value of this motion vector candidate in the list is even, its list 0 motion vector is used for triangular prediction if available, and this motion vector candidate is in the list 0 If there is no motion vector, its list 1 The motion vector is used for triangular prediction; and
[0100] For a specific motion vector candidate in the first merge list, if the merge index value of this motion vector candidate in the list is odd, its list 1 motion vector is used for triangular prediction if available, and if there is no list 1 motion vector for this motion vector candidate, its list 0 motion vector is used for triangular prediction.
[0101] Figure 11A shows an example of single prediction motion vector (MV) selection (or single prediction merge candidate selection) for the triangular prediction mode. In this example, the first five merge MV candidates derived in the first merge list are indexed from 0 to 4. Each row has two columns representing the list 0 motion vector and the list 1 motion vector of the candidates in the first merge list, respectively. Each candidate in this list can be either a single prediction or a dual prediction. In the case of a single prediction candidate, only one of the list 0 motion vector and the list 1 motion vector can be present, not both. In the case of a dual prediction candidate, both the list 0 motion vector and the list 1 motion vector are present. In Figure 11A, for each merge index, the motion vector marked with "x" is used first for triangular prediction if available. If the motion vector marked with "x" is not available, the unmarked motion vector corresponding to the same merge index is used next for triangular prediction.
[0102] The above concept can be extended to other examples. Figure 11B shows another example of single prediction motion vector (MV) selection for the triangular prediction mode. According to Figure 11B, the rules for selecting a single prediction merge candidate for triangular prediction are as follows.
[0103] For the motion vector candidates in the first merge list, only one of its list 0 motion vector and list 1 motion vector is used for triangular prediction;
[0104] For a specific motion vector candidate in the first merge list, if the merge index value of this motion vector candidate in the list is even, its list 1 motion vector is used for triangular prediction if available, and if this motion vector candidate has no list 1 motion vector, its list 0 motion vector is used for triangular prediction; and
[0105] For a particular motion vector candidate in the first merge list, if the merge index value of this motion vector candidate in the list is odd, then the list 0 motion vector, if available, is used for triangular prediction, and if there is no list 0 motion vector for this motion vector candidate, then the list 1 motion vector is used for triangular prediction.
[0106] In one example, other different orders are defined and used to select a single prediction merge candidate for triangular prediction from those motion vector candidates in the first merge list. More specifically, for a particular motion vector candidate in the first merge list, the determination of whether its list 0 motion vector or list 1 motion vector, if available, is used for triangular prediction does not necessarily depend on the parity of the index value of the candidate in the first merge list as described above. For example, the following rules may be used.
[0107] For a motion vector candidate in the first merge list, only one of its list 0 motion vector and list 1 motion vector is used for triangular prediction;
[0108] Based on a certain predefined pattern, for multiple motion vector candidates in the first merge list, their list 0 motion vectors, if available, are used for triangular prediction, and if there is no list 0 motion vector, the corresponding list 1 motion vectors are used for triangular prediction;
[0109] Based on the same predefined pattern, for the remaining motion vector candidates in the first merge list, their list 1 motion vectors, if available, are used for triangular prediction, and if there is no list 1 motion vector, the corresponding list 0 motion vectors are used for triangular prediction;
[0110] Figures 12A through 12D show examples of single prediction motion vector (MV) selection for the triangular prediction mode. For each merge index, the motion vector marked with an "x" is used first for triangular prediction if available. If the motion vector marked with an "x" is not available, the unmarked motion vector corresponding to the same merge index is used next for triangular prediction.
[0111] In Figure 12A, for the first three motion vector candidates in the first merge list, first their list 0 motion vectors are checked. The corresponding list 1 motion vectors are used for triangular prediction only if the list 0 motion vectors are not available. For the fourth and fifth motion vector candidates in the first merge list, first the list 1 motion vectors are checked. The corresponding list 0 motion vectors are used for triangular prediction only if the list 1 motion vectors are not available. Figures 12B through 12D show three other patterns when selecting a single prediction merge candidate from the first merge list. The examples shown in the figures are not limiting and there are additional examples. For example, horizontally and / or vertically mirrored versions of those patterns shown in Figures 12A through 12D may be used.
[0112] The selected single prediction merge candidate is indexed and can be directly accessed from the first merge list; or these selected single prediction merge candidates can be put into a single prediction merge list for triangular prediction. The derived single prediction merge list includes a plurality of single prediction merge candidates, and each single prediction merge candidate includes one motion vector of the corresponding candidate in the first merge list. According to an example of the present disclosure, each candidate in the first merge list includes at least one of a list 0 motion vector and a list 1 motion vector, and each single prediction merge candidate can be only one of the list 0 motion vector and the list 1 motion vector of the corresponding candidate in the first merge list. Each single prediction merge candidate is associated with an integer-valued merge index. The list 0 motion vector and the list 1 motion vector are selected based on a preset rule of the single prediction merge candidate.
[0113] In one example, for each single prediction merge candidate having an even merge index value, the list 0 motion vector of the corresponding candidate having the same merge index in the first merge list is selected as the single prediction merge candidate; and for each single prediction merge candidate having an odd merge index value, the list 1 motion vector of the corresponding candidate having the same merge index in the first merge list is selected. In another example, for each single prediction merge candidate having an even merge index value, the list 1 motion vector of the corresponding candidate having the same merge index in the first merge list is selected; and for each single prediction merge candidate having an odd merge index value, the list 0 motion vector of the corresponding candidate having the same merge index in the first merge list is selected.
[0114] In yet another example, for each single prediction merge candidate, if it is determined that a list 1 motion vector of the corresponding candidate in the first merge list is available, then the list 1 motion vector is selected as the single prediction merge candidate; and when it is determined that the list 1 motion vector is not available, a list 0 motion vector of the corresponding candidate in the first merge list is selected.
[0115] In yet another example, for each single prediction merge candidate having a merge index value within a first range, a list 0 motion vector of the corresponding candidate in the first merge list is selected as the single prediction merge candidate; and for each single prediction merge candidate having a merge index value within a second range, a list 1 motion vector of the corresponding candidate in the first merge list is selected.
[0116] In the process described above, motion vector truncation can also be performed in the same way. Such truncation can be performed completely or partially. When performed partially, it means that the new motion vector is compared with not all but only some of the motion vectors already included in the single prediction merge list. It also means that not all but only some of the new motion vectors need to be checked for truncation before being used as merge candidates for triangular prediction. One specific example is that only the second motion vector is checked for truncation against the first motion vector before the second motion vector is used as a merge candidate for triangular prediction, and all other motion vectors are not checked for truncation. In an extreme case, motion vector truncation (i.e., motion vector comparison processing) is not performed in this process.
[0117] The methods for forming a single prediction merge list in the present disclosure have been described with respect to the triangular prediction mode, but these methods are applicable to other prediction modes of the same kind. For example, in a more general geometric partitioning prediction mode where a CU is partitioned into two PUs along a line that is not an exact diagonal, the two PUs can have geometric shapes such as a triangle, a wedge, or a trapezoid. In such cases, the prediction for each PU is formed in a manner similar to the triangular prediction mode, and the methods described herein are equally applicable.
[0118] FIG. 13 is a block diagram showing an apparatus for video encoding / decoding according to an embodiment of the present disclosure. The apparatus 1300 may be a terminal such as a mobile phone, a tablet computer, a digital broadcast terminal, a tablet device, or a personal digital assistant.
[0119] As shown in FIG. 13, the apparatus 1300 may include one or more of a processing unit 1302, a memory 1304, a power supply unit 1306, a multimedia unit 1308, an audio unit 1310, an input / output (I / O) interface 1312, a sensor unit 1314, and a communication unit 1316.
[0120] The processing unit 1302 generally controls the overall operation of the apparatus 1300, such as operations related to display, phone call initiation, data communication, camera operation, and recording operation. The processing unit 1302 may include one or more processors 1320 for executing instructions to implement all or part of the steps of the above methods. Further, the processing unit 1302 may include one or more modules that contribute to the interaction between the processing unit 1302 and other components. For example, the processing unit 1302 may include a multimedia module for contributing to the interaction between the multimedia unit 1308 and the processing unit 1302.
[0121] Memory 1304 is configured to store different types of data to support the operation of device 1300. Examples of such data include instructions for any application or method operating on device 1300, contact data, phone book data, messages, images, videos, and the like. Memory 1304 is implemented by any type of volatile or non-volatile storage device or a combination thereof, and memory 1304 may be a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or a compact disk.
[0122] Power supply unit 1306 supplies power to each component of device 1300. Power supply unit 1306 may include a power management system, one or more power supplies, and other components related to generating, managing, and distributing power for device 1300.
[0123] The multimedia unit 1308 includes a screen that provides an output interface between the device 1300 and the user. In one example, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen that receives input signals from the user. This touch panel may include one or more touch sensors for sensing touches, slides, and gestures on this touch panel. The touch sensors can detect not only the boundaries of touch or slide operations but also the duration and pressure associated with the touch or slide operations. In one example, the multimedia unit 1308 may include a front camera and / or a rear camera. When the device 1300 is in an operation mode such as an imaging mode or a video mode, the front camera and / or the rear camera can receive external multimedia data.
[0124] The audio unit 1310 is configured to output and / or input audio signals. For example, the audio unit 1310 includes a microphone (MIC). The microphone is configured to receive external audio signals when the device 1300 is in an operation mode such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 1304 or transmitted via the communication unit 1316. In one example, the audio unit 1310 further includes a speaker for outputting audio signals.
[0125] The I / O interface 1312 provides an interface between the processing unit 1302 and the peripheral interface module. The above-mentioned peripheral interface module may be a keyboard, a click wheel, buttons, etc. These buttons include, but are not limited to, a home button, a volume button, a start button, and a lock button.
[0126] The sensor unit 1314 includes one or more sensors for providing state evaluation in different modes of the device 1300. For example, the sensor unit 1314 can detect the on / off state of the device 1300 and the relative positions of components. For example, the components are the display and keypad of the device 1300. The sensor unit 1314 can also detect changes in the position of the device 1300 or components of the device 1300, the presence or absence of user contact on the device 1300, the orientation or acceleration / deceleration of the device 1300, and temperature changes of the device 1300. The sensor unit 1314 may include a proximity sensor configured to detect the presence of nearby objects without physical contact. The sensor unit 1314 may further include an optical sensor such as a CMOS or CCD image sensor used in imaging applications. In one example, the sensor unit 1314 may further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0127] The communication unit 1316 is configured to facilitate wired or wireless communication between the device 1300 and other devices. The device 1300 can access a wireless network based on communication standards such as WiFi, 4G, or a combination thereof. In one example, the communication unit 1316 receives a notification signal or notification-related information from an external notification management system via a notification channel. In one example, the communication unit 1316 may further include a near-field communication (NFC) module for facilitating short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0128] In one example, the device 1300 may be implemented by one or more of the following: an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a controller, a microcontroller, a microprocessor, or other electronic elements for executing the above method.
[0129] A non-transitory computer-readable storage medium may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flash memory, a hybrid drive or a solid state hybrid drive (SSHD), a read-only memory (ROM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, etc.
[0130] FIG. 14 is a flowchart showing an exemplary process of video encoding / decoding for motion compensation prediction by triangle prediction according to an embodiment of the present disclosure.
[0131] In step 1402, the processor 1320 partitions a video image into a plurality of coding units (CUs), at least one of which is further partitioned into two prediction units (PUs). These two PUs can include at least one geometric shape PU. For example, this geometric shape PU can include a pair of triangular PUs, a pair of wedge-shaped PUs, or other geometric shape PUs.
[0132] In step 1404, the processor 1320 constructs a first merge list including a plurality of candidates each including one or more motion vectors. For example, the processor 1320 can construct the first merge list based on a merge list construction process for normal merge prediction. The processor 1320 can also obtain the first merge list from another electronic device or storage unit.
[0133] In step 1406, the processor 1320 obtains or derives a single prediction merge list for the triangular-shaped PU. Here, the single prediction merge list includes a plurality of single prediction merge candidates, and each single prediction merge candidate includes one motion vector of the corresponding candidate in the first merge list.
[0134] In one example, an apparatus for video coding is provided. The apparatus includes a processor 1320 and a memory 1304 configured to store instructions executable by the processor. Here, the processor is configured to execute a method as shown in FIG. 14 when the instructions are executed.
[0135] In another example, a non-transitory computer-readable storage medium 1304 storing instructions is provided. When these instructions are executed by a processor 1320, they cause the processor to execute a method as shown in FIG. 14.
[0136] The description of the present disclosure is presented for purposes of illustration and is not intended to be exhaustive or limited to the present disclosure. Many modifications, variations, and alternative implementations will be apparent to those skilled in the art having the benefit of the teachings presented in the foregoing description and the related drawings.
[0137] Embodiments are selected and described in order to explain the principles of the present disclosure and enable those skilled in the art to understand the disclosure for various implementations, make the best use of various implementations for the basic principles and specific applications for which various changes are expected. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed implementations, and that modifications and other implementations are also included within the scope of the present disclosure.
Claims
Claim 1 Partitioning a plurality of coding units (CUs) into two PUs, where at least one video image is further partitioned into at least two prediction units (PUs) including geometric shapes; Constructing a first merge list including a plurality of candidates each including one or more motion vectors; Directly obtaining from within the first merge list a plurality of single prediction merge candidates each including one motion vector of a corresponding candidate within the first merge list without creating a single prediction merge candidate list; comprising each candidate within the first merge list includes at least one of a list 0 motion vector and a list 1 motion vector, and each single prediction merge candidate includes only one of the list 0 motion vector and the list 1 motion vector of each candidate within the first merge list; each single prediction merge candidate is associated with an integer-valued merge index, and the list 0 motion vector and the list 1 motion vector are selected based on a preset rule for the single prediction merge candidate; each single prediction merge candidate having an even merge index value includes the list 0 motion vector of the corresponding candidate within the first merge list if it is determined that the list 0 motion vector of the corresponding candidate is available, or includes the list 1 motion vector of the corresponding candidate within the first merge list if it is determined that the list 0 motion vector of the corresponding candidate is not available, or each single prediction merge candidate having an odd merge index value includes the list 1 motion vector of the corresponding candidate within the first merge list if it is determined that the list 1 motion vector of the corresponding candidate is available, or includes the list 0 motion vector of the corresponding candidate within the first merge list if it is determined that the list 1 motion vector of the corresponding candidate is not available, a method for video coding. Claim 2 The method according to claim 1, wherein to obtain a plurality of single prediction merge candidates, the list 0 motion vector and / or the list 1 motion vector of each candidate within the first merge list are indexed based on a specific order. Claim 3 The method according to claim 1, wherein each single prediction merge candidate having a merge index value includes a list 0 motion vector or a list 1 motion vector of a corresponding candidate having the same merge index within the first merge list.
4. One or more processors, A memory configured to store instructions executable by the one or more processors, Including, The one or more processors, when executing the instructions, are configured to execute the method for video encoding according to any one of claims 1 to 3, and an apparatus for video encoding.
5. A non-transitory computer-readable storage medium storing instructions, The instructions, when executed by one or more processors, cause the one or more processors to execute the method according to any one of claims 1 to 3 to generate a video bitstream and store the video bitstream in the non-transitory computer-readable storage medium.
6. A computer program storing instructions, The instructions, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 3.
7. Executing the method for video encoding according to any one of claims 1 to 3 to generate a bitstream; Transmitting the bitstream to a decoding device, and a method for transmitting a bitstream.
Citation Information
Patent Citations
Image encoding device, image encoding method, image encoding program, image decoding device, image decoding method, and image decoding program
WO2020184459A1