Systems and methods for signaling a motion merge mode in video coding and decoding
By obtaining the conventional merge identifier of the encoding unit in the video decoding method, determining whether to use the conventional merge mode or the MMVD mode, and constructing a merge list to select appropriate motion vector candidates, the problem of inefficient signaling transmission motion merge mode in the prior art is solved, and higher encoding efficiency and visual quality are achieved.
Patent Information
- Application Number
- CN202410636938.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-31
- Filing Date
- 2019-12-30
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2039-12-30
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in signaling motion merging mode, resulting in poor encoding efficiency and visual quality.
By obtaining the conventional merge identifier of the encoding unit in the video decoding method, it is determined whether to use the conventional merge mode or the merge mode (MMVD) with motion vector difference (MMVD), and a merge list is constructed to select appropriate motion vector candidates.
The semantic signaling efficiency of merging related modes is improved, and the encoding efficiency and visual quality of video encoding and decoding are improved.
Smart Images

Figure CN118433377B_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with the application number 201980087346.0 and the invention name "Systems and Methods for Signaling Motion Merge Modes in Video Coding and Decoding". This invention patent application is the national stage application in China of the international patent application PCT / US2019 / 068977 filed on December 30, 2019. This international patent application is based on and claims the priority of the provisional application No. 62 / 787,230 filed on December 31, 2018, and the entire content of this application is incorporated herein by reference. Technical Field
[0002] This application relates to video coding and compression. More specifically, this application relates to systems and methods for signaling motion merge modes in video coding and decoding. Background Art
[0003] Various video coding and decoding techniques can be used to compress video data. Video coding and decoding is performed according to one or more video coding and decoding standards. For example, video coding and decoding standards include Versatile Video Coding (VVC), Joint Exploration Test Model (JEM), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), Moving Picture Experts Group (MPEG) coding, etc. Video coding and decoding typically utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.), which utilize the redundancy present in video images or sequences. An important goal of video coding and decoding techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. Summary of the Invention
[0004] Examples of the present disclosure provide a method for improving the efficiency of semantic signaling of merge-related modes.
[0005] According to a first aspect of the present disclosure, there is provided a method for video decoding, including: obtaining, from a decoder, a regular merge flag for a coding unit (CU), the coding unit being coded in a merge mode and a merge-related mode; when the regular merge flag is 1, indicating that a regular merge mode or a merge mode with motion vector difference (MMVD) is used by the CU, constructing a single merge list for the CU, where the single merge list includes regular motion vector candidates and MMVD motion vector candidates, the regular motion vector candidates and MMVD motion vector candidates are selected by a regular merge index to indicate the candidates being used, and the single merge list is constructed for both the regular merge mode and MMVD; and when the regular merge flag is zero, indicating that the regular merge mode is not used by the CU, and further receiving a mode flag to indicate that an associated merge-related mode is used when a constraint condition of the mode flag is satisfied.
[0006] According to a second aspect of the present disclosure, there is provided a computing device, including: one or more processors; a memory coupled to the one or more processors; and a plurality of programs stored in the memory, the plurality of programs, when executed by the one or more processors, cause the computing device to execute the above video decoding method.
[0007] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing a bitstream and a plurality of programs, wherein the plurality of programs, when executed by one or more processing units, cause the computing device to execute the above video decoding method to decode the bitstream.
[0008] According to a fourth aspect of the present disclosure, there is provided a computer program product including a plurality of program codes for execution by a computing device having one or more processing units, wherein the plurality of program codes, when executed by the one or more processing units, cause the computing device to execute the above video decoding method.
[0009] It should be understood that the foregoing general description and the following detailed description are merely exemplary and not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings incorporated in and constituting a part of this specification illustrate examples consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0011] Figure 1 is a block diagram of an encoder according to an example of the present disclosure.
[0012] Figure 2 is a block diagram of a decoder according to an example of the present disclosure.
[0013] Figure 3 is a flowchart illustrating a method for deriving a constructed affine merge candidate according to an example of the present disclosure.
[0014] Figure 4 is a flowchart illustrating a method for determining whether an identification constraint condition is satisfied according to an example of the present disclosure.
[0015] Figure 5A is a diagram illustrating an MMVD search point according to an example of the present disclosure.
[0016] Figure 5B is a diagram illustrating an MMVD search point according to an example of the present disclosure.
[0017] Figure 6A is an affine motion model based on control points according to an example of the present disclosure.
[0018] Figure 6B is an affine motion model based on control points according to an example of the present disclosure.
[0019] Figure 7 is a diagram illustrating an affine motion vector field (MVF) of each sub-block according to an example of the present disclosure.
[0020] Figure 8 is a diagram illustrating the position of an inherited affine motion predictor according to an example of the present disclosure.
[0021] Figure 9 is a diagram illustrating the inheritance of control point motion vectors according to an example of the present disclosure.
[0022] Figure 10 is a diagram illustrating the position of candidate orientations according to an example of the present disclosure.
[0023] Figure 11 is a diagram illustrating spatial neighboring blocks used by sub-block based temporal motion vector prediction (SbTMVP) according to an example of the present disclosure.
[0024] FIG. 12A is a diagram illustrating a sub-block based temporal motion vector prediction (SbTMVP) process according to an example of the present disclosure.
[0025] Figure 12B is a diagram illustrating a sub-block based temporal motion vector prediction (SbTMVP) process according to an example of the present disclosure.
[0026] Figure 13A is a diagram illustrating a triangular partition according to an example of the present disclosure.
[0027] Figure 13B is a diagram illustrating a triangular partition according to an example of the present disclosure.
[0028] Figure 14 is a diagram illustrating a computing environment coupled to a user interface according to an example of the present disclosure. DETAILED DESCRIPTION
[0029] Reference will now be made in detail to example embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, where like numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations set forth in the following description of example embodiments do not represent all implementations consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with aspects of the present disclosure recited in the appended claims.
[0030] The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a", "an" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein is intended to mean and include any and all possible combinations of one or more of the associated listed items.
[0031] It should be understood that although the terms "first", "second", "third", etc. may be used herein to describe various information, such information should not be limited by these terms. These terms are only used to distinguish one category of information from another. For example, without departing from the scope of this disclosure, the first information may be regarded as the second information; and similarly, the second information may also be regarded as the first information. As used herein, depending on the context, the term "if" may be understood to mean "when" or "at the time of" or "in response to a judgment".
[0032] Video coding and decoding system.
[0033] Conceptually, video coding standards are similar. For example, many use block-based processing and share a similar video coding block diagram to achieve video compression.
[0034] In this embodiment of the present disclosure, several methods are proposed to improve the efficiency of semantic signaling of merge-related modes. It should be noted that the proposed methods can be applied independently or in combination.
[0035] Figure 1 A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-frame mode determination 116, block predictor 140, adder 128, transform 130, quantization 132, prediction-related information 142, intra prediction 118, picture buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, loop filter 122, entropy coding 138, and bitstream 144.
[0036] In an example embodiment of the encoder, a video frame is partitioned into blocks for processing. For each given video block, a prediction is formed based on either inter-frame prediction or intra-frame prediction. In inter-frame prediction, a predictor can be formed based on pixel points from a previously reconstructed frame through motion estimation and motion compensation. In intra-frame prediction, a predictor can be formed based on the reconstructed pixel points in the current frame. Through mode determination, the best predictor can be selected to predict the current block.
[0037] The prediction residual (i.e., the difference between the current block and its predictor) is sent to the transform module. The transform coefficients are then sent to the quantization module for entropy reduction. The quantized coefficients are fed to the entropy coding module to generate the compressed video bitstream. As Figure 1 shown, prediction-related information from the inter-frame and / or intra-frame prediction modules (such as block partition information, motion vectors, reference picture indices, and intra-frame prediction modes, etc.) also passes through the entropy coding module and is saved into the bitstream.
[0038] In the encoder, a decoder-related module is also needed to reconstruct the pixel points for prediction purposes. First, the prediction residual is reconstructed through inverse quantization and inverse transformation. This reconstructed prediction residual is combined with the block predictor to generate the unfiltered reconstructed pixel points of the current block.
[0039] To improve the coding efficiency and visual quality, loop filters are usually used. For example, the deblocking filter is available in AVC, HEVC, and the current VVC. In HEVC, an additional loop filter called SAO (Sample Adaptive Offset) is defined to further improve the coding efficiency. In the latest VVC, another loop filter called ALF (Adaptive Loop Filter) is being actively studied and has a great chance of being incorporated into the final standard.
[0040] Figure 2 A typical decoder 200 block diagram is shown. The decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transformation 216, adder 218, intra / inter mode selection 220, intra prediction 222, memory 230, loop filter 228, motion compensation 224, picture buffer 226, prediction-related information 234, and video output 232.
[0041] In the decoder, first, the bitstream is decoded by the entropy decoding module to derive the quantized coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by the inverse quantization and inverse transformation modules to obtain the reconstructed prediction residual. Based on the decoded prediction information, the block predictor is formed through an intra prediction or motion compensation process. The unfiltered reconstructed pixel points are obtained by summing the reconstructed prediction residual and the block predictor. When the loop filter is turned on, a filtering operation is performed on these pixel points to derive the final reconstructed video for output.
[0042] Figure 3 An example method for deriving the constructed affine merge candidates according to the present disclosure is shown.
[0043] In step 310, obtain a regular merge flag for a coding unit (CU) from a decoder, where the coding unit is coded in a merge mode and a merge-related mode.
[0044] In step 312, when the regular merge flag is 1, indicate that the regular merge mode or the merge mode with motion vector difference (MMVD) is used by the CU, construct a motion vector merge list for the CU, and use a regular merge index to indicate the candidate being used.
[0045] In step 314, when the regular merge flag is zero, indicate that the regular merge mode is not used by the CU, and further receive a mode flag to indicate that a related merge-related mode is used when the constraint conditions of the mode flag are satisfied.
[0046] Figure 4 An example method for determining whether identity constraint conditions are satisfied according to the present disclosure is shown.
[0047] In step 410, obtain a coded block from a decoder, where the coded block has a width and a height.
[0048] In step 412, determine by the decoder whether both the width of the coded block and the height of the coded block are not equal to 4.
[0049] In step 414, determine by the decoder that the width of the coded block is not equal to 8 or the height of the coded block is not equal to 4.
[0050] In step 416, determine by the decoder that the width of the coded block is not equal to 4 or the height of the coded block is not equal to 8.
[0051] In step 418, determine by the decoder that the regular merge flag is not set.
[0052] Versatile Video Coding (VVC).
[0053] At the 10th JVET meeting (April 10 - 20, 2018, San Diego, USA), JVET defined a draft initial version of the Versatile Video Coding (VVC) and VVC Test Model 1 (VTM1) coding methods. The decision included using a quadtree with nested multi-type trees of binary and ternary splits coding block structure as the initial new coding feature of VVC. Since then, a reference software VTM for implementing the coding method and the draft VVC decoding process has been developed during JVET meetings.
[0054] The picture partitioning structure divides the input video into blocks called coding tree units (CTUs). A quadtree with a nested multi-type tree structure is used to divide a CTU into coding units (CUs), where a leaf coding unit (CU) defines a region that shares the same prediction mode (e.g., intra or inter). In this document, the term "unit" is defined to cover the image region of all components; the term "block" is used to define a region that covers a specific component (e.g., luminance), and when considering chrominance sampling formats such as 4:2:0, the term "block" may differ in spatial position.
[0055] Extended merge mode in VVC.
[0056] In VTM3, the merge candidate list is constructed by sequentially including the following five types of candidates:
[0057] 1. Spatial MVP from spatial neighbor CUs
[0058] 2. Temporal MVP from collocated CUs
[0059] 3. History-based MVP from the FIFO table
[0060] 4. Paired-average MVP
[0061] 5. Zero MV.
[0062] The size of the merge list is signaled in the slice header, and the maximum allowed size of the merge list is 6 in VTM3. For each CU code in the merge mode, the index of the best merge candidate is encoded using truncated unary binary (TUB). Context is used to encode the first binary bit (bin) of the merge index, and bypass coding is used for the other binary bits. In the following context of this disclosure, this extended merge mode is also referred to as the regular merge mode, since its concept is the same as the merge mode used in HEVC.
[0063] Merge mode with MVD (MMVD).
[0064] In addition to the merge mode in which the implicitly derived motion information is directly used for the prediction sample generation of the current CU, a merge mode with motion vector difference (MMVD) is introduced in VVC. The MMVD flag is signaled immediately after the skip flag and the merge flag to specify whether the MMVD mode is used for the CU.
[0065] In MMVD, after a merge candidate is selected, the merge candidate is further refined by MVD information signaled. The further information includes a merge candidate identifier, an index for specifying a motion amplitude, and an index for indicating a motion direction. In the MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. The merge candidate identifier is signaled to specify which one is used.
[0066] The distance index specifies the motion amplitude information and indicates a predefined offset from a starting point. As shown in FIG. 5 (described below), the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1:
[0067] Table 1 - Relationship between distance index and predefined offset
[0068]
[0069] The direction index indicates the direction of the MVD relative to the starting point. The direction index can indicate four directions, as shown in Table 2. Note that the meaning of the MVD sign may vary according to the information of the starting MV. When the starting MV is a uni - directional prediction MV or a bi - directional prediction MV and both lists point to the same side of the current picture (i.e., both POCs of the two reference pictures are greater than the POC of the current picture, or both are less than the POC of the current picture), the sign specification in Table 2 is added to the sign of the MV offset of the starting MV. When the starting MV is a bi - directional prediction MV and the two MVs point to different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), the sign specification in Table 2 is added to the sign of the MV offset of the list 0 MV component of the starting MV, and the sign of the list 1 MV has the opposite value.
[0070] Table 2 - Signs of MV offsets specified by direction index
[0071]
[0072]
[0073] Figure 5A FIG. shows a diagram illustrating MMVD search points for a first list (L0) reference according to the present disclosure.
[0074] Figure 5B FIG. shows a diagram illustrating MMVD search points for a second list (L1) reference according to the present disclosure.
[0075] Affine motion compensation prediction.
[0076] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). In the real world, there are many kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VTM3, block-based affine transformation motion compensation prediction is applied. Figure 6A and 6B As shown in (and described below), the affine motion field of a block is described by the motion information of two control points (4 parameters) or three control point motion vectors (6 parameters).
[0077] Figure 6A A control point based affine motion model for a 4-parameter affine model according to the present disclosure is shown.
[0078] Figure 6B A control point based affine motion model for a 6-parameter affine model according to the present disclosure is shown.
[0079] For the 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as:
[0080]
[0081] For the 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as:
[0082]
[0083] Among them (m v0x ,m v0y ) is the motion vector of the upper left control point, (m v1x ,m v1y ) is the motion vector of the upper right control point, and (m v2x, m v2y ) is the motion vector of the lower left control point.
[0084] To simplify motion compensated prediction, block-based affine transform prediction is applied. To derive the motion vector for each 4×4 luminance sub-block, the following formula is used to calculate Figure 7 The motion vector of the center sample of each sub-block shown in (described below) is obtained and the motion vector is rounded to 1 / 16 fractional precision. Then, a motion compensated interpolation filter is applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size of the chroma component is also set to 4×4. The MV of a 4×4 chroma sub-block is calculated as the average of the MVs of the four corresponding 4×4 luminance sub-blocks.
[0085] Figure 7 The affine motion vector field (MVF) of each sub-block according to the present disclosure is shown.
[0086] Similar to what is done for translational motion inter - prediction, there are also two affine motion inter - prediction modes: affine merge mode and affine AMVP mode.
[0087] Affine merge prediction.
[0088] The AF_MERGE mode can be applied to CUs where both the width and height are greater than or equal to 8. In this mode, the CPMV of the current CU is generated based on the motion information of spatially neighboring CUs. There can be up to five CPMVP candidates, and an index is signaled to indicate which candidate will be used for the current CU. The following three types of CPVM candidates are used to form the affine merge candidate list:
[0089] 6. Inherited affine merge candidates extrapolated from the CPMV of neighboring CUs
[0090] 7. Constructed affine merge candidate CPMVP derived using the translational MV of neighboring CUs
[0091] 8. Zero MV.
[0092] In VTM3, there are up to two inherited affine candidates derived from the affine motion models of neighboring blocks, one from the left neighboring CU and one from the upper neighboring CU. The candidate blocks are as Figure 8 shown (described below). For the left predictor, the scan order is A0 -> A1, and for the upper predictor, the scan order is B0 -> B1 -> B2. Only the first inherited candidate from each side is selected. No pruning check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vectors are used to derive CPMVP candidates in the affine merge list of the current CU. As Figure 9 shown (described below), if the lower - left neighbor block A is coded in affine mode, the motion vectors v 2 、v 3 and v 4 of the upper - left, upper - right, and lower - left corners of the CU containing block A are obtained. When block A is coded using a 4 - parameter affine model, two CPMVs of the current CU are calculated based on v 2 and v 3 . In the case where block A is coded using a 6 - parameter affine model, three CPMVs of the current CU are calculated based on v 2 、v 3 and v 4 .
[0093] Figure 8 Shows the position of the inherited affine motion predictor according to the present disclosure.
[0094] Figure 9 Shows the inheritance of control point motion vectors according to the present disclosure.
[0095] The constructed affine candidates mean that candidates are constructed by combining the translational motion information of the neighbors of each control point. The motion information of the control points is derived from Figure 10 the specified spatial and temporal neighbors shown below (described below). CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV 1 , check the B2->B3->A2 block and use the MV of the first available block. For CPMV 2 , check the B1->B0 block, and for CPMV 3 , check the A1->A0 block. For TMVP, it is used as CPMV 4 (if it is available).
[0096] Figure 10 Shows the position of the candidate orientation for the constructed affine merge mode according to the present disclosure.
[0097] After obtaining the MVs of the four control points, affine merge candidates are constructed based on the corresponding motion information. The following combinations of control point MVs are used to construct the following items in order:
[0098] {CPMV 1 , CPMV 2 , CPMV 3}, {CPMV 1 , CPMV 2 , CPMV 4}, {CPMV 1 , CPMV 3 , CPMV 4}, {CPMV 2 , CPMV 3 , CPMV 4}, {CPMV 1 , CPMV 2}, {CPMV 1 , CPMV 3}.
[0099] Combinations of 3 CPMVs constitute 6-parameter affine merge candidates, and combinations of 2 CPMVs constitute 4-parameter affine merge candidates. To avoid the motion scaling process, if the reference indices of the control points are different, the relevant combinations of control point MVs are discarded.
[0100] After checking the inherited affine merge candidates and the constructed affine merge candidates, if the list is still not full, zero MVs are inserted at the end of the list.
[0101] Sub - block - based Temporal Motion Vector Prediction (SbTMVP).
[0102] VTM supports the sub - block - based temporal motion vector prediction (SbTMVP) method. Similar to the temporal motion vector prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated picture to improve the motion vector prediction and merge mode of the CUs in the current picture. The same collocated picture used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in the following two main aspects:
[0103] 1. TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub - CU level;
[0104] 2. Although TMVP obtains the temporal motion vector from the collocated block (the collocated block is the bottom - right or center block relative to the current CU) in the collocated picture, SbTMVP applies a motion shift before obtaining the temporal motion information from the collocated picture, where the motion shift is obtained from the motion vector of one of the spatial neighboring blocks of the current CU.
[0105] The SbTVMP process is illustrated as (described below) in Figure 11 、 Figure 12A and Figure 12B SbTMVP predicts the motion vectors of the sub - CUs within the current CU in two steps. In the first step, the spatial neighbors in Figure 11 are checked in the order of A1, B1, B0, and A0. Once the first spatial neighboring block with a motion vector using the collocated picture as its reference picture is identified, that motion vector is selected as the motion shift to be applied. If no such motion is identified from the spatial neighbors, the motion shift is set to (0, 0).
[0106] Figure 11 shows the spatial neighboring blocks used by the sub - block - based temporal motion vector prediction (SbTMVP). SbTMVP is also known as Alternative Temporal Motion Vector Prediction (ATMVP).
[0107] In the second step, the motion shifts identified in step 1 (i.e., the coordinates added to the current block) are applied to obtain sub-CU level motion information (motion vectors and reference indices) from the collocated pictures, as shown in FIGS. 12A and 12B. The examples in FIGS. 12A and 12B assume that the motion shift is set to the motion of block A1. Then, for each sub-CU, the motion information of its corresponding block (the smallest motion grid covering the central sample) in the collocated picture is used to derive the motion information of the sub-CU. After the motion information of the collocated sub-CUs is identified, it is converted into the motion vectors and reference indices of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference pictures of the temporal motion vectors with those of the current CU.
[0108] FIG. 12A shows the SbTMVP process for collocated pictures in VVC when deriving the sub-CU motion field by applying the motion shift from spatial neighbors and scaling the motion information from the corresponding collocated sub-CUs.
[0109] Figure 12B Shows the SbTMVP process for the current picture in VVC when deriving the sub-CU motion field by applying the motion shift from spatial neighbors and scaling the motion information from the corresponding collocated sub-CUs.
[0110] In VTM3, a sub-block based merge list containing a combination of SbTVMP candidates and affine merge candidates is used for signaling the sub-block based merge mode. In the following context, the sub-block merge mode is used. The SbTVMP mode is enabled / disabled by sequence parameter set (SPS) identification. If the SbTMVP mode is enabled, the SbTMVP predictor is added as the first entry in the list of sub-block based merge candidates, followed by the affine merge candidates. The size of the sub-block based merge list is signaled in the SPS, and the maximum allowed size of the sub-block based merge list in VTM3 is 5.
[0111] The sub-CU size used in SbTMVP is fixed to 8×8, and similar to what is done for the affine merge mode, the SbTMVP mode only applies to CUs where both the width and height are greater than or equal to 8.
[0112] The encoding logic for additional SbTMVP merge candidates is the same as that for other merge candidates, i.e., for each CU in a P or B slice, an additional RD check is performed to decide whether to use the SbTMVP candidate.
[0113] Combined Inter-Intra Prediction (CIIP).
[0114] In VTM3, when a CU is encoded in merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), additional signaling is sent to indicate whether the combined inter / intra prediction (CIIP) mode is applied to the current CU.
[0115] To form the CIIP prediction, the intra prediction mode is first derived from two additional semantic elements. Up to four possible intra prediction modes can be used: directional angular prediction (DC), planar prediction (PLANAR), horizontal prediction (HORIZONAL), or vertical prediction (VERTICAL). Then, the inter prediction and intra prediction signals are derived using the regular intra / inter decoding process. Finally, a weighted average of the inter and intra prediction signals is performed to obtain the CIIP prediction.
[0116] Intra prediction mode derivation.
[0117] In CIIP mode, up to 4 intra prediction modes (including DC, PLANAR, HORIZONAL, and VERTICAL modes) can be used to predict the luma component. If the CU shape is very wide (i.e., the width is more than twice the height), the HORIZONAL mode is not allowed. If the CU shape is very narrow (i.e., the height is more than twice the width), the VERTICAL mode is not allowed. In these cases, only 3 intra prediction modes are allowed.
[0118] The CIIP mode uses the 3 most probable modes (MPM) for intra prediction. The CIIP MPM candidate list is formed as follows:
[0119] - The left and top neighboring blocks are set as A and B respectively
[0120] - The intra prediction modes of block A and block B, denoted as intraModeA and intraModeB respectively, are derived as follows:
[0121] ○ Let X be A or B
[0122] ○ If 1) block X is not available; or 2) block X is not predicted using the CIIP mode or intra mode; 3) block B is outside the current CTU, then set intra mode X to DC
[0123] ○ Otherwise, 1) if the intra prediction mode of block X is DC or PLANAR, set the intra mode X to DC or PLANAR; or 2) if the intra prediction mode of block X is a "vertically - like" angular mode (greater than 34), set the intra mode X to VERTICAL; or 3) if the intra prediction mode of block X is a "horizontally - like" angular mode (less than or equal to 34), set the intra mode X to HORIZONAL
[0124] - If the intra mode A and the intra mode B are the same:
[0125] о If the intra mode A is PLANAR or DC, set the three MPMs in the order of {PLANAR, DC, VERTICAL} to {PLANAR, DC, VERTICAL},
[0126] о Otherwise, set the three MPMs in the order of {intra mode A, PLANAR, DC} to {intra mode A, PLANAR, DC}
[0127] - Otherwise (the intra mode A and the intra mode B are different):
[0128] о Set the first two MPMs in the order of {intra mode A, intra mode B} to {intra mode A, intra mode B}
[0129] о Check the uniqueness of PLANAR, DC, and VERTICAL in the order of PLANAR, DC, and VERTICAL against the first two MPM candidate modes; once a unique mode is found, add it as the third MPM.
[0130] If the CU shape is very wide or very narrow (as defined above), infer that the MPM flag is 1 and no signaling is required. Otherwise, signal the MPM flag to indicate whether the CIIP intra prediction mode is one of the CIIP MPM candidate modes.
[0131] If the MPM flag is 1, further signal the MPM index to indicate which one of the MPM candidate modes is used in the CIIP intra prediction. Otherwise, if the MPM flag is 0, set the intra prediction mode to the "default" mode in the MPM candidate list. For example, if the PLANAR mode is not in the MPM candidate list, PLANAR is the default mode and the intra prediction mode is set to PLANAR. Since 4 possible intra prediction modes are allowed in CIIP and the MPM candidate list contains only 3 intra prediction modes, one of the 4 possible modes must be the default mode.
[0132] For the chrominance component, the DM mode is always applied without additional signaling; that is, the chrominance uses the same prediction mode as the luminance.
[0133] The intra prediction mode of the CIIP-encoded CU will be saved and used in the intra mode encoding of future neighboring CUs.
[0134] Joint inter-intra prediction signal.
[0135] Use the same inter prediction process applied to the regular merge mode to derive the inter prediction signal P in the CIIP mode inter ; and use the CIIP intra prediction mode to derive the intra prediction signal P after the regular intra prediction process intra . Then, use weighted averaging to combine the intra and inter prediction signals, where the weight values depend on the intra prediction mode and where the samples are located in the coded block, as follows:
[0136] - If the intra prediction mode is DC or PLANAR mode, or if the block width or height is less than 4, equal weights are applied to the intra prediction and inter prediction signals.
[0137] - Otherwise, the weights are determined based on the intra prediction mode (in this case, either HORIZONTAL mode or VERTICAL mode) and the sample positions in the block. Take the HORIZONTAL prediction mode as an example (the weights for the VERTICAL mode are derived in a similar way but in the orthogonal direction). Denote W as the width of the block and H as the height of the block. First, divide the coded block into four equal-area parts, each with a size of (W / 4)xH. Starting from the part closest to the intra prediction reference samples and ending at the part farthest from the intra prediction reference samples, the weights wt of each of the 4 regions are set to 6, 5, 3, and 2 respectively. Use the following formula to derive the final CIIP prediction signal:
[0138] P CIIP = ((8 - wt) * P inter + wt * P intra ) >> 3.
[0139] Triangular partitioning for inter prediction.
[0140] In VTM3, a new triangular partitioning mode was introduced for inter prediction. The triangular partitioning mode is only applied to CUs that are 8x8 or larger and are encoded in skip or merge mode. For CUs that meet these conditions and have the merge flag enabled, a CU-level flag is signaled to indicate whether the triangular partitioning mode is applied.
[0141] When using this mode, the CU is evenly divided into two triangular partitions using diagonal partitioning or anti-diagonal partitioning ( Figure 13A and Figure 13B , as described below). Each triangular partition in the CU is inter-frame predicted using its own motion; only unidirectional prediction is allowed for each partition, that is, each partition has one motion vector and one reference index. The unidirectional prediction motion constraint conditions are applied to ensure that, as with conventional bidirectional prediction, only two motion compensation predictions are required for each CU.
[0142] If the CU level flag indicates that the current CU is encoded using the triangular partitioning mode, an index in the range [0, 39] is further signaled. Using this triangular partitioning index, the direction (diagonal or anti-diagonal) of the triangular partition and the motion of each partition can be obtained through a look-up table. After predicting each triangular partition, a hybrid process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edges. This is the prediction signal for the entire CU, and as in other prediction modes, the transform and quantization processes will be applied to the entire CU. Finally, the motion field of the CU predicted using the triangular partitioning mode is stored in 4x4 units.
[0143] Figure 13A Inter-frame prediction based on triangular partitioning according to the present disclosure is shown.
[0144] Figure 13B Inter-frame prediction based on triangular partitioning according to the present disclosure is shown.
[0145] Context Adaptive Binary Arithmetic Coding (CABAC).
[0146] Context Adaptive Binary Arithmetic Coding (CABAC) is a form of entropy coding used in the H.264 / MPEG-4 AVC and High Efficiency Video Coding (HEVC) standards and VVC. CABAC is based on arithmetic coding, with some innovations and changes to adapt it to the requirements of video coding standards:
[0147] · It encodes binary symbols, which maintains low complexity and allows probability modeling for the more frequent bits in any symbol.
[0148] · The probability model is adaptively selected based on local context, allowing better probability modeling because the coding patterns usually have good local correlation.
[0149] · It uses multiplication-free range division through the use of quantized probability ranges and probability states.
[0150] For different contexts, CABAC has multiple probability models. It first converts all non-binary symbols into binary. Then, for each binary digit (or bit), the encoder selects which probability model to use and then uses information from nearby elements to optimize the probability estimate. Finally, arithmetic coding is applied to compress the data.
[0151] Context modeling provides an estimate of the conditional probability of the encoded symbols. Using a suitable context model, the given inter-symbol redundancy can be exploited by switching between different probability models according to the symbols already encoded in the neighborhood of the current symbol to be encoded.
[0152] The encoding of data symbols involves the following stages.
[0153] · Binaryization: CABAC uses binary arithmetic coding, which means that only binary decisions (1 or 0) are encoded. Non-binary value symbols (e.g., transform coefficients or motion vectors) are "binaryized" or converted into binary codes before arithmetic coding. This process is similar to the process of converting data symbols into variable-length codes, but the binary codes are further encoded (by the arithmetic encoder) before transmission.
[0154] · Repeat each stage for each binary digit (or "bit") of the binaryized symbol.
[0155] · Context model selection: A "context model" is a probability model for one or more binary digits of the binaryized symbol. This model can be selected from the available models depending on the statistics of the most recently encoded data symbols. The context model stores the probabilities that each binary digit is "1" or "0".
[0156] · Arithmetic coding: The arithmetic encoder encodes each binary digit according to the selected probability model. Note that for each binary digit (corresponding to "0" and "1"), there are only two sub-ranges.
[0157] · Probability update: Update the selected context model based on the actual encoded value (e.g., if the binary digit value is "1", increment the frequency count of "1").
[0158] Figure 14 A computing environment 1410 coupled to a user interface 1460 is shown. The computing environment 1410 can be part of a data processing server. The computing environment 1410 includes a processor 1420, a memory 1440, and an input / output interface 1450.
[0159] Processor 1420 generally controls the overall operation of computing environment 1410, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1420 may include one or more processors to execute instructions to perform all or some of the steps in the methods described above. Additionally, processor 1420 may include one or more modules that facilitate interaction between processor 1420 and other components. The processor may be a central processing unit (CPU), a microprocessor, a single-chip machine, a GPU, etc.
[0160] Memory 1440 is configured to store various types of data to support the operation of computing environment 1410. Examples of such data include instructions for any application or method operating on computing environment 1410, MRI data sets, image data, etc. Memory 1440 may be implemented by using any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0161] Input / output interface 1450 provides an interface between processor 1420 and peripheral interface modules (such as a keyboard, click wheel, buttons, etc.). These buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. Input / output interface 1450 may be coupled to an encoder and a decoder.
[0162] In one embodiment, a non-transitory computer-readable storage medium is also provided, which includes a plurality of programs, such as the programs included in memory 1440, that can be executed by processor 1420 in computing environment 1410 to perform the methods described above. For example, the non-transitory computer-readable storage medium may be a ROM, a RAM, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0163] A non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, where the plurality of programs, when executed by the one or more processors, cause the computing device to perform the methods described above for motion prediction.
[0164] In one embodiment, the computing environment 1410 may be implemented using one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0165] Method according to the present disclosure.
[0166] As described above, in VTM-3.0, the merge mode is further classified into five categories, including conventional merge, merge mode with MVD (MMVD), sub-block merge (including affine merge and sub-block based temporal motion vector prediction), combined inter-intra prediction (CIIP) merge, and triangular partition merge. The semantics of the merge mode signaling in the current VVC are illustrated in the following table.
[0167] Table 3 Semantics of merge-related modes in the current VVC
[0168]
[0169]
[0170] In summary, in the current VVC, the semantics (associated identifiers) signaled to indicate the corresponding merge mode are as described below.
[0171] Table 4 Signaling of merge-related modes in the current VVC
[0172] MMVD identifier Sub-block identifier CIIP identifier Triangle identifier MMVD 1 - - - Sub-block 0 1 - - CIIP 0 0 1 - Triangle 0 0 0 1 Conventional 0 0 0 0 。
[0173] It has been observed that more than 50% of the merge modes are conventional merge modes. However, in VTM-3.0, the codeword for the conventional merge mode is the longest among the five different merge modes, which is not an efficient design in terms of semantic parsing. In the current VVC, the skip mode has a semantic design similar to the merge mode except for the CIIP mode for skipping. However, the same observation has been made in the skip mode.
[0174] Semantics for conventional merge.
[0175] As mentioned above, among several merge-related modes including (conventional merge, MMVD, sub-block merge, CIIP, and triangular merge), the scheme of the conventional merge mode in the current VVC is the most frequently used. In an embodiment of the present disclosure, explicit identification for the conventional merge mode is signaled to indicate whether the conventional merge mode is used. As shown in the following table, a conventional identification (or referred to as the conventional merge identification) is explicitly signaled into the bitstream, and all signaling of related identifications is modified accordingly. CABAC is used to context-encode the conventional merge identification. In one scheme, only one context is used to encode the conventional merge identification. In another scheme, multiple context models are used to encode the conventional merge identification, and the selection of the context model is based on encoded information such as the conventional merge identification of neighboring blocks or the size of the current CU.
[0176] Example of signaling for merge-related modes in the proposed scheme in Table 5
[0177] Conventional identifier MMVD identifier Sub-block identifier CIIP identifier Conventional 1 - - - MMVD 0 1 - - Sub-block 0 0 1 - CIIP 0 0 0 1 Triangle 0 0 0 0 。
[0178] In the current VVC, the constraint conditions for enabling merge-related modes are different, and thus the signaling of the identification for each merge-related mode is also different, as summarized below.
[0179] Table 6 Constraint conditions for enabling / signaling merge-related modes
[0180]
[0181] Therefore, the signaling of the conventional merge identification should also consider the different constraint conditions applied to each identification signaling. For example, when the block size is 4x4, 8x4, or 4x8, only the conventional merge mode and MMVD are valid. Under these conditions (block size of 4x4, 8x4, or 4x8), only the conventional merge identification is signaled; when the conventional merge identification is equal to 1, the conventional merge mode is used; otherwise, when the conventional merge identification is equal to 0, MMVD is used. The following illustrates an example based on the semantics of the current VVC working draft.
[0182] Example of semantics in the proposed scheme in Table 7
[0183]
[0184]
[0185] In this example, it should be noted that the regular merge flag is explicitly signaled into the bitstream. However, the regular merge flag can be signaled in any orientation, and the regular merge flag does not have to be the first orientation as described above. In yet another scenario, the regular merge flag is signaled, but is signaled after the MMVD and sub-block merge flags.
[0186] Integrate the relevant merge modes into the regular merge mode.
[0187] In an embodiment of the present disclosure, MMVD, CIIP, and triangle are merged into the regular merge mode. In this scenario, all MMVD candidates, CIIP candidates, and triangle merge candidates are considered regular merge candidates, and a regular merge index is used to indicate the candidate being used. Therefore, the size of the regular merge candidate list needs to be expanded accordingly. In one example, a regular merge index equal to N (N can be any positive integer and is less than the maximum size of the regular merge candidate list) means that the MMVD mode is selected, and further semantics are signaled / received to indicate which MMVD candidate is being used. The same scenario is also applied to the CIIP and triangle merge modes.
[0188] In yet another example, CIIP and triangle are merged into the regular merge mode. In this scenario, all CIIP candidates and triangle merge candidates are considered regular merge candidates, and a regular merge index is used to indicate the candidate being used. Therefore, the size of the regular merge candidate list needs to be expanded accordingly.
[0189] Constraint alignment.
[0190] As mentioned above, the constraints for enabling different merge-related modes are different. In an embodiment of the present disclosure, the constraints for enabling different merge modes and signaling related flags are more aligned. In one example, the constraints are modified as illustrated in the following table.
[0191] Table 8 Modified constraints for enabling / signaling merge-related modes
[0192]
[0193] In yet another example of the present disclosure, the constraints are modified as illustrated in the following table.
[0194] Table 9 Modified constraints for enabling / signaling merge-related modes
[0195]
[0196] In yet another example of the present disclosure, the constraints are modified as illustrated in the following table. In this scenario, it should be noted that when the block width = 128 or the block height = 128, the identification of CIIP is still signaled. When the block width = 128 or the block height = 128, the identification of CIIP is constrained to always be zero because intra prediction does not support these conditions.
[0197] Table 10 Modified Constraints for Enabling / Signaling Merge-Related Modes
[0198] Constraint condition Conventional No constraint condition MMVD No constraint condition Sub-block Block width > 8 and block height > 8 CIIP Block width > 8 and block height > 8 Triangle Block width > 8 and block height > 8 。
[0199] In yet another example of the present disclosure, the constraints are modified as illustrated in the following table. In this scenario, it should be noted that when the block width = 128 or the block height = 128, the identification of CIIP is still signaled. When the block width = 128 or the block height = 128, the identification of CIIP is constrained to always be zero because intra prediction does not support these conditions.
[0200] Table 11 Modified Constraints for Enabling / Signaling Merge-Related Modes
[0201] Constraint condition Conventional No constraint condition MMVD No constraint condition Sub-block Block width > 8 and block height > 8 CIIP (Block width x Block height) >= 64 Triangle (Block width x Block height) >= 64 。
[0202] Switch the order of the CIIP identification and the triangular merge identification.
[0203] Switch the signaling order of the CIIP identification and the triangular merge identification because it has been observed that the triangular merge mode is used more frequently.
Claims
1. A video decoding method, comprising: obtaining, from a decoder, a regular merge flag for a coding unit (CU), the coding unit being coded as one of a regular merge mode and a merge-related mode, wherein the merge-related mode includes at least one of a merge mode with motion vector difference (MMVD), a sub-block merge mode, a combined inter-intra prediction (CIIP) mode, and a triangular merge mode; when the regular merge flag is 1, indicating that the regular merge mode or the merge mode with motion vector difference (MMVD) is used by the CU, constructing a single merge list for the CU, wherein the single merge list includes regular motion vector candidates and MMVD motion vector candidates, and the regular motion vector candidates and MMVD motion vector candidates are selected by a regular merge index to indicate the candidates to be used, wherein the single merge list is constructed for both the regular merge mode and MMVD; and when the regular merge flag is 0, indicating that the regular merge mode is not used by the CU, further receiving a mode flag, the mode flag indicating that a related merge-related mode is used when a constraint condition of the mode flag is satisfied.
2. The method according to claim 1, wherein, the mode flag is an identifier of the merge mode with motion vector difference (MMVD), and the constraint conditions of the MMVD flag include: obtaining, from a decoder, a coded block, wherein the coded block has a width and a height; determining, by the decoder, whether both the width of the coded block and the height of the coded block are not equal to 4; determining, by the decoder, that the width of the coded block is not equal to 8 or the height of the coded block is not equal to 4; determining, by the decoder, that the width of the coded block is not equal to 4 or the height of the coded block is not equal to 8; and determining, by the decoder, that the regular merge flag is not set; or wherein the method further includes: when the MMVD flag is equal to 1, receiving, by the decoder, an MMVD merge flag, an MMVD distance index, and an MMVD direction index; or wherein the method further includes receiving a sub-block flag, and the constraint conditions of the sub-block flag include: obtaining, from a decoder, a coded block, wherein the coded block has a width and a height; determining, by the decoder, whether the maximum number (MaxNumSubblockMergeCand) of sub-block-based merge MVP candidates is greater than zero; determining, by the decoder, whether the width of the coded block is greater than or equal to 8; and determining, by the decoder, whether the height of the coded block is greater than or equal to 8; or wherein the mode flag is a combined inter-intra prediction (CIIP) flag, and the constraint conditions of the CIIP flag include: obtaining, from a decoder, a coded block, wherein the coded block has a width and a height; determining, by the decoder, whether sps_mh_intra_enabled_flag is set; determining, by the decoder, whether cu_skip_flag is equal to 0; determining, by the decoder, whether the width of the coded block multiplied by the height of the coded block is greater than or equal to 64; determining, by the decoder, whether the width of the coded block is less than 128; and determining, by the decoder, whether the height of the coded block is less than 128; or Wherein, the mode identifier is a triangular identifier, and the constraints of the triangular identifier include: Determining, by a decoder, whether the regular merge identifier is not set; Determining, by a decoder, whether the identifier of the merge mode with motion vector difference (MMVD) is not set; Determining, by a decoder, whether the merge sub-block identifier is not set; and Determining, by a decoder, whether the mode identifier combined inter-intra prediction (CIIP) identifier is not set.
3. The method according to claim 1, further comprising: When the constraints of the mode identifier are satisfied, receiving, by the decoder, the regular merge identifier before receiving the mode identifier.
4. The method according to claim 3, further comprising: When the constraints of the identifier of the merge mode with motion vector difference (MMVD) are satisfied, receiving, by the decoder, the regular merge identifier before receiving the identifier of the merge mode with motion vector difference (MMVD); or Wherein the method further comprises: When the constraints of the combined inter-intra prediction (CIIP) flag are satisfied, receiving, by the decoder, the regular merge identifier before receiving the combined inter-intra prediction (CIIP) identifier.
5. The method according to claim 1, further comprising: When the constraints of the mode identifier are satisfied, receiving, by the decoder, the regular merge identifier after receiving the mode identifier.
6. The method according to claim 5, further comprising: When the constraints of the sub-block merge mode identifier are satisfied, receiving, by the decoder, the regular merge identifier after receiving the sub-block merge mode identifier.
7. The method according to claim 1, further comprising: Receiving, by the decoder, the regular merge identifier using context adaptive binary arithmetic coding (CABAC) with one context; or Wherein the method further comprises: receiving, by the decoder, the regular merge identifier using context adaptive binary arithmetic coding (CABAC) with multiple context models, and the selection of the context models is based on coding information, wherein the coding information includes the size of the current CU.
8. A computing device, comprising: One or more processors; A memory coupled to the one or more processors; and Multiple programs stored in the memory, which, when executed by the one or more processors, cause the computing device to execute the method according to any one of claims 1-7.
9. A non-transitory computer-readable storage medium storing a bitstream and multiple programs, wherein the multiple programs, when executed by one or more processing units, cause a computing device to execute the method according to any one of claims 1-7 to decode the bitstream.
10. A computer program product comprising multiple program codes for being executed by a computing device having one or more processing units, wherein the multiple program codes, when executed by the one or more processing units, cause the computing device to execute the method according to any one of claims 1-7.
11. A method for sending a bitstream, comprising: Sending a bitstream to a decoding device, wherein the bitstream is decoded according to the method according to any one of claims 1-7.