Systems and methods for signaling motion merge mode in video coding
By improving the semantic signaling of the merge-related mode in video coding, and constructing a merge list using conventional merge identifiers and MMVD, the problem of low efficiency in signaling transmission motion merging mode in existing technologies is solved, achieving more efficient video coding and better visual quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-30
- Publication Date
- 2026-03-24
AI Technical Summary
Existing video encoding and decoding technologies are inefficient in signaling transmission motion merging mode, making it difficult to effectively utilize redundant information in video images or sequences for compression.
By improving the semantic signaling of merge-related patterns, using the merge list constructed with regular merge identifiers and motion vector difference merge patterns (MMVD), and using merge-related patterns when specific constraints are met, the selection and encoding process of motion vectors is optimized.
It improves video encoding efficiency and visual quality, reduces bit rate requirements, and enhances encoding efficiency.
Smart Images

Figure CN118921476B_ABST
Abstract
Description
[0001] This application is a divisional application of patent application No. 201980087346.0 entitled "System and Method for Signaling Transmission Motion Combining Mode in Video Encoding and Decoding," which is the Chinese national phase application of international patent application PCT / US2019 / 068977 filed on December 30, 2019. This international patent application is based on and claims priority to provisional application No. 62 / 787,230 filed on December 31, 2018, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to video encoding / decoding and compression. More specifically, this application relates to systems and methods for signaling motion merging modes in video encoding / decoding. Background Technology
[0003] Various video codec techniques can be used to compress video data. Video codecs are performed according to one or more video codec standards. For example, video codec standards include Universal Video Codec (VVC), Joint Explore Test Model (JEM), High Efficiency Video Codec (H.265 / HEVC), High-Level Video Codec (H.264 / AVC), and Moving Picture Experts Group (MPEG) coding. Video codecs typically utilize prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. A key goal of video codec techniques is to compress video data to a lower bitrate while avoiding or minimizing degradation in video quality. Summary of the Invention
[0004] The examples disclosed herein provide a method for improving the efficiency of semantic signaling for merging related patterns.
[0005] According to a first aspect of this disclosure, a method for video encoding is provided, comprising: signaling a regular merging identifier for a coding unit (CU), the coding unit being encoded as a regular merging mode and a merging-related mode; when the regular merging identifier is signaled to be 1, indicating that a regular merging mode or a merging mode with motion vector difference (MMVD) is used by the CU; constructing a single merging list for the CU, wherein the single merging list includes regular motion vector candidates and MMVD motion vector candidates, the regular motion vector candidates and MMVD motion vector candidates being selected by a regular merging index to indicate the candidates being used, wherein the single merging list is constructed for both the regular merging mode and MMVD; and when the regular merging identifier is signaled to be zero, indicating that the regular merging mode is not used by the CU, and further signaling a mode identifier to indicate that a associated merging-related mode is used when a constraint condition of the mode identifier is met; wherein the method further comprises: when the regular merging identifier is signaled to be 1, determining whether to signal an MMVD merging identifier based on the value of an MMVD identifier.
[0006] According to a second aspect of this disclosure, a computing device is provided, comprising: one or more processors; a memory coupled to the one or more processors; and a plurality of programs stored in the memory, wherein, when executed by the one or more processors, the computing device causes the computing device to perform the method encoded above.
[0007] According to a third aspect of this disclosure, a non-transitory computer-readable storage medium is provided, storing a plurality of programs, wherein when executed by one or more processing units of a computing device, the computing device performs an encoding according to the above-described encoding method to generate a bit stream and store the bit stream in the non-transitory computer-readable storage medium.
[0008] It should be understood that the foregoing general description and the following detailed description are merely examples and not limitations of this disclosure. Attached Figure Description
[0009] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate examples consistent with this disclosure and, together with this description, serve to explain the principles of this disclosure.
[0010] Figure 1 This is a block diagram of an encoder based on an example of this disclosure.
[0011] Figure 2 This is a block diagram of a decoder based on an example of this disclosure.
[0012] Figure 3This is a flowchart illustrating a method for deriving constructed affine merge candidates according to an example of this disclosure.
[0013] Figure 4 This is a flowchart illustrating a method for determining whether an identification constraint is satisfied, according to an example of this disclosure.
[0014] Figure 5A This is a diagram illustrating an MMVD search point according to an example of this disclosure.
[0015] Figure 5B This is a diagram illustrating an MMVD search point according to an example of this disclosure.
[0016] Figure 6A It is an example of a control point-based affine motion model based on this disclosure.
[0017] Figure 6B It is an example of a control point-based affine motion model based on this disclosure.
[0018] Figure 7 This is a diagram illustrating the affine motion vector field (MVF) of each sub-block according to an example of this disclosure.
[0019] Figure 8 This is a diagram illustrating the position of an affine motion predictor inherited according to an example of this disclosure.
[0020] Figure 9 This is a diagram illustrating the inheritance of control point motion vectors according to an example of this disclosure.
[0021] Figure 10 It is a diagram illustrating the location of the candidate orientation according to an example of this disclosure.
[0022] Figure 11 This is an illustration of a spatially neighboring block used by sub-block-based temporal motion vector prediction (SbTMVP) according to an example of this disclosure.
[0023] Figure 12 This is a diagram illustrating a sub-block-based temporal motion vector prediction (SbTMVP) process according to an example of this disclosure.
[0024] Figure 13A This is a diagram illustrating a triangular partition according to an example of this disclosure.
[0025] Figure 13B This is a diagram illustrating a triangular partition according to an example of this disclosure.
[0026] Figure 14 This is a diagram illustrating a computing environment coupled with a user interface according to an example of this disclosure. Detailed Implementation
[0027] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the drawings, wherein the same numerals in different drawings denote the same or similar elements unless otherwise indicated. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with this disclosure. Instead, they are merely examples of apparatuses and methods consistent with aspects of this disclosure recited in the appended claims.
[0028] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure. As used in this disclosure and the appended claims, the singular forms “a,” “an,” and “the” are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein is intended to represent and include any one or all possible combinations of one or more of the associated listed items.
[0029] It should be understood that although the terms “first,” “second,” “third,” etc., may be used herein to describe various types of information, such information should not be limited by these terms. These terms are used only to distinguish one category of information from another. For example, without departing from the scope of this disclosure, first information may be considered second information; and similarly, second information may be considered first information. As used herein, depending on the context, the term “if” may be understood to mean “when,” “in,” or “in response to a judgment.”
[0030] Video encoding and decoding system.
[0031] Conceptually, video codec standards are similar. For example, many people use block-based processing and share similar video codec block diagrams to achieve video compression.
[0032] In this embodiment of the disclosure, several methods are proposed to improve the efficiency of semantic signaling for merging related patterns. It should be noted that the proposed methods can be applied independently or in combination.
[0033] Figure 1 A typical encoder 100 is shown. The encoder 100 has a video input 110, motion compensation 112, motion estimation 114, intra / inter-frame mode determination 116, block predictor 140, adder 128, transform 130, quantization 132, prediction related information 142, intra-frame prediction 118, image buffer 120, inverse quantization 134, inverse transform 136, adder 126, memory 124, loop filter 122, entropy coding 138, and bitstream 144.
[0034] In an example embodiment of the encoder, video frames are partitioned into blocks for processing. For each given video block, a prediction is formed based on either inter-frame prediction or intra-frame prediction. In inter-frame prediction, the predictor can be formed based on pixels from previously reconstructed frames, through motion estimation and motion compensation. In intra-frame prediction, the predictor can be formed based on reconstructed pixels in the current frame. The optimal predictor can be selected to predict the current block through mode determination.
[0035] The prediction residual (i.e., the difference between the current block and its predictor) is sent to the transform module. The transform coefficients are then sent to the quantization module for entropy reduction. The quantization coefficients are fed into the entropy coding module to generate a compressed video bitstream. Figure 1 As shown, prediction-related information (such as block partitioning information, motion vectors, reference image indexes, and intra-prediction modes) from the inter-frame and / or intra-frame prediction modules are also processed by the entropy coding module and saved into the bitstream.
[0036] The encoder also requires a decoder-related module to reconstruct the pixels used for prediction. First, the prediction residual is reconstructed through inverse quantization and inverse transform. This reconstructed prediction residual is then combined with a block predictor to generate the unfiltered reconstructed pixels for the current block.
[0037] To improve coding efficiency and visual quality, loop filters are commonly used. For example, deblocking filters are available in AVC, HEVC, and the current VVC. In HEVC, an additional loop filter called SAO (Sample Adaptive Offset) is defined to further improve coding efficiency. In the latest VVC, another loop filter called ALF (Adaptive Loop Filter) is under active investigation, and this loop filter has a high chance of being included in the final standard.
[0038] Figure 2 A typical block diagram of decoder 200 is shown. Decoder 200 has a bitstream 210, entropy decoding 212, inverse quantization 214, inverse transform 216, adder 218, intra / inter-frame mode selection 220, intra-frame prediction 222, memory 230, loop filter 228, motion compensation 224, picture buffer 226, prediction related information 234, and video output 232.
[0039] In the decoder, the bitstream is first decoded by an entropy decoding module to derive quantization coefficient levels and prediction-related information. Then, the quantized coefficient levels are processed by an inverse quantization and inverse transform module to obtain the reconstructed prediction residual. Based on the decoded prediction information, a block predictor is formed through intra-frame prediction or motion compensation. The unfiltered reconstructed pixels are obtained by summing the reconstructed prediction residual and the block predictor. With the loop filter enabled, these pixels are filtered to derive the final reconstructed video for output.
[0040] Figure 3 An example method for deriving constructed affine merging candidates according to this disclosure is shown.
[0041] In step 310, a regular merge identifier for a coding unit (CU) is obtained from the decoder, the coding unit being encoded as a merge mode and a merge-related mode.
[0042] In step 312, when the regular merge identifier is 1, it indicates that the regular merge mode or the merge mode with motion vector difference (MMVD) is used by the CU, a motion vector merge list is constructed for the CU and the regular merge index is used to indicate the candidates used.
[0043] In step 314, when the regular merge identifier is zero, it indicates that the regular merge mode is not used by the CU, and a mode identifier is further received to indicate that the associated merge-related mode is used when the constraints of the mode identifier are met.
[0044] Figure 4 An example method for determining whether an identifier constraint is satisfied, according to this disclosure, is shown.
[0045] In step 410, an encoded block is obtained from the decoder, wherein the encoded block has a width and a height.
[0046] In step 412, the decoder determines whether the width and height of the encoded block are both not equal to 4.
[0047] In step 414, the decoder determines that the width of the encoded block is not equal to 8 or the height of the encoded block is not equal to 4.
[0048] In step 416, the decoder determines that the width of the encoded block is not equal to 4 or the height of the encoded block is not equal to 8.
[0049] In step 418, the decoder determines that the regular merge identifier is not set.
[0050] Universal Video Codec (VVC).
[0051] At the 10th JVET meeting (April 10-20, 2018, San Diego, USA), JVET defined an initial draft of the Universal Video Coding (VVC) and VVC Test Model 1 (VTM1) coding methods. It was decided that a quadtree with nested multi-type trees using binary and ternary splits of the coding block structure would be included as the initial new coding feature for VVC. Subsequently, reference software VTM was developed during the JVET meeting to implement the coding method and the draft VVC decoding process.
[0052] The image partitioning structure divides the input video into blocks called coding tree units (CTUs). CTUs are further divided into coding units (CUs) using a quadtree with a nested multi-type tree structure, where leaf coding units (CUs) define regions that share the same prediction mode (e.g., intra-frame or inter-frame). In this document, the term "unit" defines an image region covering all components; the term "block" is used to define a region covering a specific component (e.g., luma), and the term "block" may differ in spatial location when considering chroma sampling formats such as 4:2:0.
[0053] Extended merge pattern in VVC.
[0054] In VTM3, the list of merged candidates is constructed by including the following five types of candidates in sequence:
[0055] 1. Spatial MVP from Spatial Neighbor CU
[0056] 2. Time MVP from co-located CU
[0057] 3. Historical MVPs from FIFO tables
[0058] 4. Paired average MVP
[0059] 5. Zero MV.
[0060] The size of the merge list is sent via signaling in the slice header, and the maximum allowed size of the merge list is 6 in VTM3. For each CU code in the merge mode, the index of the best merge candidate is encoded using truncated unary binary representation (TU). The first binary bit (bin) of the merge index is encoded using context, and bypass encoding is used for the remaining binary bits. In the following context of this disclosure, this extended merge mode is also referred to as the regular merge mode because its concept is the same as the merge mode used in HEVC.
[0061] Merge pattern with MVD (MMVD).
[0062] In addition to the merging pattern that directly uses implicitly derived motion information to generate prediction samples for the current CU, VVC also introduces a merging pattern with motion vector difference (MMVD). Immediately after sending the skip flag and the merge flag, signaling sends the MMVD flag to specify whether the MMVD pattern should be used for the CU.
[0063] In MMVD, after a merge candidate is selected, it is further refined via MVD information sent through signaling. This further information includes a merge candidate identifier, an index specifying the motion amplitude, and an index indicating the motion direction. In MMVD mode, one of the first two candidates in the merge list is selected as the MV basis. The merge candidate identifier is sent via signaling to specify which one is used.
[0064] The distance index specifies motion amplitude information and indicates a predefined offset from the starting point. As shown in Figure 5 (described below), the offset is added to the horizontal or vertical component of the starting MV. The relationship between the distance index and the predefined offset is specified in Table 1:
[0065] Table 1 - Relationship between Distance Index and Predefined Offset
[0066]
[0067] The direction index indicates the direction of the MVD relative to the starting point. The direction index can represent four directions, as shown in Table 2. It is important to note that the meaning of the MVD symbol may vary depending on the information of the starting MV. When the starting MV is a unidirectional or bidirectional prediction MV, and both lists point to the same side of the current image (i.e., both reference images have a POC greater than the current image's POC, or both have a POC less than the current image's POC), the symbols in Table 2 specify the sign of the MV offset added to the starting MV. When the starting MV is a bidirectional prediction MV, and the two MVs point to different sides of the current image (i.e., one reference image has a POC greater than the current image's POC, and the other reference image has a POC less than the current image's POC), the symbols in Table 2 specify the sign of the MV offset added to the list 0 MV component of the starting MV, and the signs for list 1 MV have the opposite values.
[0068] Table 2 - Signs of MV Offsets Specifyed by Direction Index
[0069] Directional IDX 00 01 10 11 X-axis + – N / A N / A y-axis N / A N / A + – .
[0070] Figure 5AA diagram illustrating MMVD search points for a first list (L0) reference according to this disclosure is shown.
[0071] Figure 5B A diagram illustrating MMVD search points for a second list (L1) reference, according to this disclosure, is shown.
[0072] Affine motion compensation prediction.
[0073] In HEVC, only the translational motion model is applied to motion compensation prediction (MCP). However, in the real world, many types of motion exist, such as zooming in / out, rotation, perspective motion, and other irregular motions. In VTM3, block-based affine transformation motion compensation prediction is applied. For example... Figure 6A and 6B As shown (described below), the affine motion field of a block is described by motion information from two control points (4 parameters) or three control point motion vectors (6 parameters).
[0074] Figure 6A A control-point-based affine motion model for a 4-parameter affine model is shown according to this disclosure.
[0075] Figure 6B A control-point-based affine motion model for a 6-parameter affine model is shown according to this disclosure.
[0076] For a 4-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as follows:
[0077]
[0078] For a 6-parameter affine motion model, the motion vector at the sample position (x, y) in the block is derived as follows:
[0079]
[0080] Where (m v0x ,m v0y ) is the motion vector of the upper left control point, (m v1x ,m v1y ) is the motion vector of the upper right control point, and (m v2x ,m v2y ) is the motion vector of the lower left control point.
[0081] To simplify motion compensation prediction, a block-based affine transformation prediction is applied. To derive the motion vector for each 4×4 lumen sub-block, the following formula is used: Figure 7The motion vector of the center sample of each sub-block shown (described below) is calculated, and the motion vector is rounded to 1 / 16 fractional precision. A motion-compensated interpolation filter is then applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size for the chroma component is also set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks.
[0082] Figure 7 The affine motion vector field (MVF) for each sub-block according to this disclosure is shown.
[0083] Similar to the inter-frame prediction for translational motion, there are two other inter-frame prediction modes for affine motion: affine merging mode and affine AMVP mode.
[0084] Affine merging prediction.
[0085] The AF_MERGE mode can be applied to CUs whose width and height are both greater than or equal to 8. In this mode, the CPVM of the current CU is generated based on the motion information of spatially neighboring CUs. Up to five CPVM candidates can exist, and a signaling transmission index indicates which candidate will be used for the current CU. The following three types of CPVM candidates are used to form the affine merging candidate list:
[0086] 6. Affine merging candidates inherited from the CPMV extrapolated from the neighboring CU
[0087] 7. Affine merging candidate CPMVP derived using the translation MV of neighboring CUs
[0088] 8. Zero MV.
[0089] In VTM3, there are at most two inherited affine candidates derived from the affine motion model of neighboring blocks: one from the left neighboring CU and one from the upper neighboring CU. Candidate blocks are as follows: Figure 8 As shown in the diagram (described below). For the left predictor, the scan order is A0->A1, and for the upper predictor, the scan order is B0->B1->B2. Only the first inherited candidate from each side is selected. No pruning check is performed between the two inherited candidates. When a neighboring affine CU is identified, its control point motion vector is used to derive CPMVP candidates in the affine merging list of the current CU. Figure 9As shown (described below), if the lower-left neighboring block A is encoded using an affine model, the motion vectors v2, v3, and v4 of the top-left, top-right, and bottom-left corners of the CU containing block A are obtained. When block A is encoded using a 4-parameter affine model, the two CPMVs of the current CU are calculated based on v2 and v3. In the case where block A is encoded using a 6-parameter affine model, the three CPMVs of the current CU are calculated based on v2, v3, and v4.
[0090] Figure 8 The position of the affine motion predictor inherited according to this disclosure is shown.
[0091] Figure 9 The inheritance of control point motion vectors according to this disclosure is shown.
[0092] Constructing an affine candidate means building the candidate by combining the translational motion information of each control point's neighbors. The motion information of the control points comes from... Figure 10 The CPMV is derived from the specified spatial and temporal neighbors shown below. k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV1, check the B2->B3->A2 block and use the MV of the first available block. For CPMV2, check the B1->B0 block, and for CPMV3, check the A1->A0 block. For TMVP, it is used as CPMV4 (if it is available).
[0093] Figure 10 The positions of candidate orientations for the constructed affine merging pattern are shown according to this disclosure.
[0094] After obtaining the motion values (MVs) of the four control points, affine merging candidates are constructed based on the corresponding motion information. The following combinations of control point MVs are used to construct the following items in sequence:
[0095] {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3}.
[0096] Combinations of three CPMVs constitute a 6-parameter affine merging candidate, and combinations of two CPMVs constitute a 4-parameter affine merging candidate. To avoid motion scaling, combinations of control point MVs are discarded if the reference indices of the control points are different.
[0097] After checking the inherited affine merge candidates and the constructed affine merge candidates, if the list is still not full, zero MV is inserted at the end of the list.
[0098] Sub-block-based temporal motion vector prediction (SbTMVP).
[0099] VTM supports the Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP) method. Similar to Temporal Motion Vector Prediction (TMVP) in HEVC, SbTMVP uses the motion field in the collocated image to improve motion vector prediction and merging patterns for CUs in the current image. The same collocated image used by TMVP is used for SbTVMP. SbTMVP differs from TMVP in two main aspects:
[0100] 1. TMVP predicts motion at the CU level, but SbTMVP predicts motion at the sub-CU level;
[0101] 2. Although TMVP obtains the temporal motion vector from the juxtaposed block in the juxtaposed image (the juxtaposed block is the lower right or center block relative to the current CU), SbTMVP applies a motion shift before obtaining the temporal motion information from the juxtaposed image. The motion shift is obtained from the motion vector of one of the spatially neighboring blocks of the current CU.
[0102] The SbTVMP process is as follows: Figure 11 , Figure 12 A and Figure 12 The diagram in Figure B (described below) illustrates this. SbTMVP predicts the motion vectors of sub-CUs within the current CU in two steps. In the first step, the vectors are checked in the order of A1, B1, B0, and A0. Figure 11 Spatial neighbors. Once a first spatial neighbor block with a motion vector that uses a juxtaposed image as its reference image is identified, that motion vector is selected as the motion shift to apply. If no such motion is identified from the spatial neighbors, the motion shift is set to (0,0).
[0103] Figure 11 The image shows a spatially neighboring block used by Sub-Block-Based Temporal Motion Vector Prediction (SbTMVP). SbTMVP is also known as Alternate Temporal Motion Vector Prediction (ATMVP).
[0104] In the second step, the motion shift identified in step 1 (i.e., the coordinates added to the current block) is applied to obtain sub-CU level motion information (motion vectors and reference indices) from the juxtaposed image, such as... Figure 12 As shown in A and 12B. Figure 12The examples in A and 12B assume that the motion shift is set as the motion of block A1. Then, for each sub-CU, the motion information of the sub-CU is derived using the motion information of its corresponding block (the smallest motion grid covering the center sample) in the juxtaposed image. After the motion information of the juxtaposed sub-CU is identified, it is converted into the motion vector and reference index of the current sub-CU in a manner similar to the TMVP process of HEVC, where temporal motion scaling is applied to align the reference image of the temporal motion vector with those of the current CU.
[0105] Figure 12 A illustrates the SbTMVP process in VVC for a juxtaposed image when deriving the motion field of a sub-CU by applying motion shifts from spatial neighbors and scaling motion information from the corresponding juxtaposed sub-CU.
[0106] Figure 12 B illustrates the SbTMVP process in VVC for the current image when deriving the motion field of a sub-CU by applying motion shifts from spatial neighbors and scaling motion information from the corresponding juxtaposed sub-CU.
[0107] In VTM3, a sub-block-based merge list containing a combination of SbTVMP and affine merge candidates is used for signaling in sub-block-based merge mode. Sub-block merge mode is used in the following contexts: SbTVMP mode is enabled / disabled via the Sequence Parameter Setting (SPS) flag. If SbTVMP mode is enabled, the SbTVMP predictor is added as the first entry in the sub-block-based merge candidate list, followed by the affine merge candidate. The signaling in the SPS transmits the size of the sub-block-based merge list, and the maximum allowed size of the sub-block-based merge list in VTM3 is 5.
[0108] The sub-CU size used in SbTMVP is fixed at 8×8, and as with the affine merge mode, the SbTMVP mode only applies to CUs whose width and height are both greater than or equal to 8.
[0109] The coding logic for the additional SbTMVP merge candidate is the same as that for other merge candidates, that is, for each CU in the P or B slice, an additional RD check is performed to determine whether to use the SbTMVP candidate.
[0110] Joint Inter-Frame Intra-Frame Prediction (CIIP).
[0111] In VTM3, when a CU is encoded in merge mode, if the CU contains at least 64 luma samples (i.e., the CU width multiplied by the CU height is equal to or greater than 64), the signaling sends an additional identifier to indicate whether the Joint Inter-Frame / Intra-Frame Prediction (CIIP) mode is applied to the current CU.
[0112] To generate CIIP predictions, the intra-prediction mode is first derived from two additional semantic elements. Up to four possible intra-prediction modes can be used: Direction Angle Prediction (DC), Planar Prediction, Horizontal Prediction (HORIZONAL), or Vertical Prediction (VERTICAL). Then, the inter-frame and intra-frame prediction signals are derived using the standard intra-frame and inter-frame decoding process. Finally, a weighted average of the inter-frame and intra-frame prediction signals is performed to obtain the CIIP prediction.
[0113] Intra-frame prediction mode export.
[0114] In CIIP mode, up to four intra-prediction modes (DC, PLANA, HORIZONAL, and VERTICAL modes) can be used to predict the luma component. HORIZONAL mode is not allowed if the CU shape is very wide (i.e., width more than twice its height). VERTICAL mode is not allowed if the CU shape is very narrow (i.e., height more than twice its width). In these cases, only three intra-prediction modes are allowed.
[0115] CIIP mode uses three most probable modes (MPMs) for intra-frame prediction. The CIIP MPM candidate list is as follows:
[0116] - The left and top adjacent blocks are set to A and B respectively.
[0117] - The intra prediction modes of block A and block B, respectively represented as intramode A and intramode B, are derived as follows:
[0118] Let X be A or B.
[0119] If 1) block X is unavailable; or 2) block X is not predicted using CIIP mode or intra-frame mode; or 3) block B is outside the current CTU, then set intra-frame mode X to DC.
[0120] Otherwise, 1) if the intra-prediction mode of block X is DC or PLANA, then set intra-prediction mode X to DC or PLANA; or 2) if the intra-prediction mode of block X is a "vertical-like" angle mode (greater than 34), then set intra-prediction mode X to VERTICAL; or 3) if the intra-prediction mode of block X is a "horizontal-like" angle mode (less than or equal to 34), then set intra-prediction mode X to HORIZONAL.
[0121] -If intra-frame mode A and intra-frame mode B are the same:
[0122] If intra-frame mode A is PLANA or DC, then the three MPMs are set to {PLANAR, DC, VERTICAL} in the order {PLANAR, DC, VERTICAL}.
[0123] Otherwise, set the three MPMs to {Intra-Frame Mode A, PLANA, DC} in the order of {Intra-Frame Mode A, PLANA, DC}.
[0124] -Otherwise (intra-frame mode A and intra-frame mode B are different):
[0125] Set the first two MPMs to {Intra-frame Mode A, Intra-frame Mode B} in the order of {Intra-frame Mode A, Intra-frame Mode B}.
[0126] Compare the first two MPM candidate patterns, and check the uniqueness of PLANA, DC, and VERTICAL in the order of PLANA, DC, and VERTICAL; once a unique pattern is found, add it as the third MPM.
[0127] If the CU shape is very wide or very narrow (as defined above), the MPM flag is inferred to be 1, and no signaling is required. Otherwise, signaling sends the MPM flag to indicate whether the CIIP intra-frame prediction mode is one of the CIIP MPM candidate modes.
[0128] If the MPM flag is 1, further signaling is sent to the MPM index to indicate which MPM candidate mode to use in CIIP intra-frame prediction. Otherwise, if the MPM flag is 0, the intra-frame prediction mode is set to the "default" mode in the MPM candidate list. For example, if the PLANA mode is not in the MPM candidate list, then PLANA is the default mode, and the intra-frame prediction mode is set to PLANA. Since CIIP allows four possible intra-frame prediction modes, and the MPM candidate list only contains three, one of the four possible modes must be the default mode.
[0129] For the chromaticity component, the DM mode is always applied without additional signaling; that is, the chromaticity uses the same prediction mode as the luminance.
[0130] The intra-prediction mode of the CIIP-encoded CU will be saved and used in the intra-mode coding of future neighboring CUs.
[0131] Joint inter-frame and intra-frame prediction signals.
[0132] The inter-frame prediction signal P in CIIP mode is derived using the same inter-frame prediction process applied to the regular merging mode. interFurthermore, the CIIP intra-prediction mode is used to derive the intra-prediction signal P after the regular intra-prediction process. intra Then, a weighted average is used to combine the intra-frame and inter-frame prediction signals, where the weight values depend on the intra-frame prediction mode and where the sample is located in the coded block, as follows:
[0133] - If the intra-prediction mode is DC or PLANAR mode, or if the block width or height is less than 4, equal weights are applied to the intra-prediction and inter-prediction signals.
[0134] - Otherwise, the weights are determined based on the intra-prediction mode (in this case, either HORIZONAL or VERTICAL) and the sample positions within the block. Taking HORIZONAL prediction mode as an example (the weights for VERTICAL mode are derived in a similar way but in the orthogonal direction). Let W represent the width of the block and H represent the height of the block. First, the coded block is divided into four equal-area regions, each with a size of (W / 4) x H. Starting from the region closest to the intra-prediction reference sample and ending at the region furthest from the intra-prediction reference sample, the weights wt for each of the four regions are set to 6, 5, 3, and 2, respectively. The final CIIP prediction signal is derived using the following formula:
[0135] P CIIP =((8-wt)*P inter +wt*P intra )>>3.
[0136] Triangular partitioning for inter-frame prediction.
[0137] In VTM3, a new triangular partitioning mode was introduced for inter-frame prediction. The triangular partitioning mode is only applied to CUs that are 8x8 or larger and encoded in skip or merge mode. For CUs that meet these conditions and whose merge flag is enabled, signaling sends a CU-level flag to indicate whether the triangular partitioning mode is applied.
[0138] When using this mode, use diagonal partitioning or anti-diagonal partitioning. Figure 13A and Figure 13B (As described below), the CU is uniformly divided into two triangular partitions. Each triangular partition in the CU is predicted inter-frame using its own motion; only unidirectional prediction is allowed for each partition, that is, each partition has one motion vector and one reference index. Unidirectional prediction motion constraints are applied to ensure that, as with regular bidirectional prediction, only two motion-compensated predictions are required for each CU.
[0139] If the CU level identifier indicates that the current CU is encoded using the triangular partitioning pattern, further signaling is sent with an index in the range [0,39]. Using this triangular partitioning index, the orientation (diagonal or anti-diagonal) of the triangular partitions, and the motion of each partition, can be obtained via a lookup table. After predicting each triangular partition, a blending process with adaptive weights is used to adjust the sample values along the diagonal or anti-diagonal edges. This is the prediction signal for the entire CU, and as in other prediction patterns, the transformation and quantization processes are applied to the entire CU. Finally, the motion field of the CU predicted using the triangular partitioning pattern is stored in a 4x4 cell.
[0140] Figure 13A Inter-frame prediction based on triangular partitioning according to this disclosure is shown.
[0141] Figure 13B Inter-frame prediction based on triangular partitioning according to this disclosure is shown.
[0142] Context Adaptive Binary Arithmetic Coding (CABAC).
[0143] Context Adaptive Binary Arithmetic Coding (CABAC) is a form of entropy coding used in H.264 / MPEG-4 AVC, High Efficiency Video Coding (HEVC), and VVC standards. CABAC is based on arithmetic coding but incorporates several innovations and modifications to adapt it to the needs of video coding standards.
[0144] It encodes binary symbols, which maintains low complexity and allows for probabilistic modeling of more frequent bits in any symbol.
[0145] • Adaptively select the probability model based on local context, thus allowing for better probability modeling, because the coding patterns are usually well-correlated locally.
[0146] It uses multiplication-free range division by employing quantized probability ranges and probability states.
[0147] CABAC employs multiple probability models for different contexts. It first converts all non-binary symbols into binary. Then, for each binary bit (or digit), the encoder selects which probability model to use and optimizes the probability estimate using information from nearby elements. Finally, arithmetic coding is applied to compress the data.
[0148] Context modeling provides an estimate of the conditional probability of the encoded symbol. By using a suitable context model, redundancy between given symbols can be utilized by switching between different probability models based on the symbols already encoded in the neighborhood of the current symbol to be encoded.
[0149] The encoding of data symbols involves the following stages.
[0150] • Binarization: CABAC uses binary arithmetic coding, meaning that only binary decisions (1 or 0) are encoded. Non-binary value symbols (e.g., transform coefficients or motion vectors) are "binarized" or converted into binary code before arithmetic coding. This process is similar to converting data symbols into variable-length codes, but the binary code is further encoded (by the arithmetic encoder) before transmission.
[0151] • Repeat each stage for each binary bit (or “bit”) of the binary symbol.
[0152] • Context Model Selection: A "context model" is a probabilistic model of one or more bits of a binary symbol. This model can be selected from the available models based on statistics of recently encoded data symbols. The context model stores the probability that each bit is "1" or "0".
[0153] • Arithmetic Encoding: The arithmetic encoder encodes each binary bit according to the selected probability model. It's important to note that for each binary bit (corresponding to "0" and "1"), there are only two subranges.
[0154] • Probability update: Update the selected context model based on the actual encoded value (e.g., increment the frequency count of "1" if the binary bit value is "1").
[0155] Figure 14 A computing environment 1410 coupled to a user interface 1460 is shown. The computing environment 1410 may be part of a data processing server. The computing environment 1410 includes a processor 1420, memory 1440, and input / output interface 1450.
[0156] Processor 1420 typically controls the overall operation of computing environment 1410, such as operations associated with display, data acquisition, data communication, and image processing. Processor 1420 may include one or more processors to execute instructions to perform all or some of the steps of the methods described above. Furthermore, processor 1420 may include one or more modules that facilitate interaction between processor 1420 and other components. The processor may be a central processing unit (CPU), microprocessor, single-chip machine, GPU, etc.
[0157] Memory 1440 is configured to store various types of data to support the operation of computing environment 1410. Examples of such data include instructions for any application or method operating on computing environment 1410, MRI datasets, image data, etc. Memory 1440 can be implemented using any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0158] Input / output interface 1450 provides an interface between processor 1420 and peripheral interface modules (such as keyboards, click wheels, buttons, etc.). These buttons may include, but are not limited to, a home button, a start scan button, and a stop scan button. Input / output interface 1450 can be coupled to encoders and decoders.
[0159] In one embodiment, a non-transitory computer-readable storage medium is also provided, which includes a plurality of programs, such as programs included in memory 1440, which can be executed by processor 1420 in computing environment 1410 to perform the methods described above. For example, the non-transitory computer-readable storage medium may be ROM, RAM, CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0160] A non-transitory computer-readable storage medium stores a plurality of programs for execution by a computing device having one or more processors, wherein the plurality of programs, when executed by the one or more processors, cause the computing device to perform the methods described above for motion prediction.
[0161] In one embodiment, the computing environment 1410 may be implemented using one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0162] According to the method disclosed herein.
[0163] As described above, in VTM-3.0, merging modes are further categorized into five types, including regular merging, merge mode with MVD (MMVD), sub-block merging (including affine merging and sub-block-based temporal motion vector prediction), joint inter-frame intra-frame prediction (CIIP) merging, and triangular partition merging. The table below illustrates the semantics of merging mode signaling in the current VVC.
[0164] Table 3. Semantics of merge-related patterns in current VVC
[0165]
[0166]
[0167] In summary, in the current VVC, the semantics (associated identifiers) sent by signaling to indicate the corresponding merging mode are described below.
[0168] Table 4 Signaling of Merge-Related Modes in Current VVC
[0169] MMVD logo Sub-block identifier CIIP logo Triangular symbol MMVD 1 - - - sub-block 0 1 - - CIIP 0 0 1 - triangle 0 0 0 1 conventional 0 0 0 0 .
[0170] What was observed was that more than 50% of the merge patterns were regular merge patterns. However, in VTM-3.0, the codeword used for regular merge patterns was the longest among the five different merge patterns, which is not an efficient design in terms of semantic parsing. In the current VVC, skip patterns have a similar semantic design to merge patterns, except that there is no CIIP pattern for skipping. However, the same observation was made in skip patterns.
[0171] Semantics used for regular merging.
[0172] As mentioned above, among the various merge-related modes, including (regular merge, MMVD, sub-block merge, CIIP, and triangular merge), the regular merge mode scheme in the current VVC is the most frequently used. In embodiments of this disclosure, signaling is sent with an explicit identifier for the regular merge mode to indicate whether the regular merge mode is being used. As shown in the table below, a regular identifier (or regular merge identifier) is explicitly signaled into the bitstream, and all signaling associated with the identifier is modified accordingly. The regular merge identifier is context-encoded using CABAC. In one scheme, only one context is used to encode the regular merge identifier. In another scheme, multiple context models are used to encode the regular merge identifier, and the selection of the context model is based on information to be encoded, such as the regular merge identifiers of neighboring blocks or the current CU size.
[0173] Examples of signaling for the merge-related modes in the proposed scheme are shown in Table 5.
[0174] Standard signage MMVD logo Sub-block identifier CIIP logo conventional 1 - - - MMVD 0 1 - - sub-block 0 0 1 - CIIP 0 0 0 1 triangle 0 0 0 0 .
[0175] In the current VVC, the constraints for enabling merge-related modes are different, and therefore the signaling for the identifier of each merge-related mode is also different, as summarized below.
[0176] Table 6 Constraints for Enable / Signaling Sending Merge Related Modes
[0177]
[0178] Therefore, the signaling for the regular merge identifier should also consider the different constraints applied to each identifier signaling. For example, when the block size is 4x4, 8x4, or 4x8, only the regular merge mode and MMVD are valid. Under these conditions (block size 4x4, 8x4, or 4x8), only the regular merge identifier is sent; when the regular merge identifier equals 1, the regular merge mode is used; otherwise, when the regular merge identifier equals 0, MMVD is used. An example based on the semantics of the current VVC working draft is illustrated below.
[0179] Examples of semantics in the proposed schemes in Table 7
[0180]
[0181]
[0182] In this example, it's important to note that the regular merge identifier is explicitly signaled into the bitstream. However, the regular merge identifier can be signaled in any position, and it doesn't have to be the first position as described above. In yet another scenario, the regular merge identifier is signaled, but after the MMVD and sub-block merge identifiers.
[0183] Integrate the relevant merge patterns into the regular merge patterns.
[0184] In embodiments of this disclosure, MMVD, CIIP, and triangular merge are incorporated into a regular merge mode. In this scheme, all MMVD candidates, CIIP candidates, and triangular merge candidates are treated as regular merge candidates, and a regular merge index is used to indicate the candidate being used. Therefore, the size of the regular merge candidate list needs to be increased accordingly. In one example, a regular merge index equal to N (where N can be any positive integer and is less than the maximum size of the regular merge candidate list) signifies that the MMVD mode has been selected, and signaling transmission / reception provides further semantics to indicate which MMVD candidate is being used. The same scheme is also applied to CIIP and triangular merge modes.
[0185] In another example, CIIP and triangle merges are incorporated into the regular merge pattern. In this scheme, all CIIP candidates and triangle merge candidates are treated as regular merge candidates, and the regular merge index is used to indicate the candidate being used. Therefore, the size of the regular merge candidate list needs to be increased accordingly.
[0186] Align constraints.
[0187] As mentioned above, the constraints for enabling different merging-related modes are different. In the embodiments of this disclosure, the constraints for enabling different merging modes and signaling transmission-related identifiers are more aligned. In one example, the constraints are modified as illustrated in the table below.
[0188] Table 8 Modified Constraints for Enable / Signaling Sending Merge Related Modes
[0189]
[0190] In yet another example of this disclosure, the constraints have been modified as illustrated in the table below.
[0191] Table 9 Modified Constraints for Enable / Signaling Sending Merge Related Modes
[0192]
[0193] In yet another example of this disclosure, the constraints are modified as illustrated in the table below. In this scheme, it is important to note that when the block width = 128 or the block height = 128, the CIIP identifier is still signaled; however, when the block width = 128 or the block height = 128, the CIIP identifier is constrained to always be zero, because intra-frame prediction does not support these conditions.
[0194] Table 10 Modified Constraints for Enable / Signaling Sending Merge Related Modes
[0195] Constraints conventional Unconstrained conditions MMVD Unconstrained conditions sub-block Block width > 8 and block height > 8 CIIP Block width > 8 and block height > 8 triangle Block width > 8 and block height > 8 .
[0196] In yet another example of this disclosure, the constraints are modified as illustrated in the table below. In this scheme, it is important to note that when the block width = 128 or the block height = 128, the CIIP identifier is still signaled; however, when the block width = 128 or the block height = 128, the CIIP identifier is constrained to always be zero, because intra-frame prediction does not support these conditions.
[0197] Table 11 Modified Constraints for Enable / Signaling Sending Merge Related Modes
[0198] Constraints conventional Unconstrained conditions MMVD Unconstrained conditions sub-block Block width > 8 and block height > 8 CIIP (block width x block height) >= 64 triangle (block width x block height) >= 64 .
[0199] Switch the order of the CIIP identifier and the triangle merge identifier.
[0200] The signaling order of the CIIP identifier and the triangular merge identifier is switched because the triangular merge mode is observed to be used more frequently.
Claims
1. A method for video coding, comprising: signaling a regular merge flag for a coding unit (CU) coded in one of a regular merge mode and a merge-related mode, wherein the merge-related mode includes at least one of a merge mode with motion vector difference (MMVD), a subblock merge mode, a combined inter-intra prediction (CIIP) mode, and a triangle merge mode; when the regular merge flag is signaled as one, indicating that the regular merge mode or the merge mode with motion vector difference (MMVD) is used by the CU, constructing a single merge list for the CU, wherein the single merge list includes regular motion vector candidates and MMVD motion vector candidates, the regular motion vector candidates and MMVD motion vector candidates are selected by a regular merge index to indicate the used candidates, wherein the single merge list is constructed for both the regular merge mode and the MMVD; and when the regular merge flag is signaled as zero, indicating that the regular merge mode is not used by the CU, further signaling a mode flag, the mode flag indicates that an associated merge-related mode is used when a constraint condition of the mode flag is satisfied; wherein the method further comprises: when the regular merge flag is signaled as one, determining whether to signal a MMVD merge flag according to a value of a MMVD flag.
2. The method of claim 1, wherein, the mode flag is a flag of the merge mode with motion vector difference (MMVD), and the constraint condition of the MMVD flag includes: obtaining a coding block, wherein the coding block has a width and a height; specifying whether neither the width of the coding block nor the height of the coding block is equal to 4; specifying whether the width of the coding block is not equal to 8 or the height of the coding block is not equal to 4; specifying whether the width of the coding block is not equal to 4 or the height of the coding block is not equal to 8; and specifying that the regular merge flag is not set; or wherein the method further comprises: when the MMVD flag is equal to one, signaling the MMVD merge flag; or wherein the mode flag is a subblock flag, and the constraint condition of the subblock flag includes: obtaining a coding block, wherein the coding block has a width and a height; specifying whether a maximum number of subblock-based merge MVP candidates (MaxNumSubblockMergeCand) is greater than zero; specifying whether the width of the coding block is greater than or equal to 8; and specifying whether the height of the coding block is greater than or equal to 8; or wherein the mode flag is a combined inter-intra prediction (CIIP) flag, and the constraint condition of the CIIP flag includes: obtaining a coding block, wherein the coding block has a width and a height; specifying whether sps_mh_intra_enabled_flag is set; specifying whether cu_skip_flag is equal to zero; specifying whether the width of the coding block multiplied by the height of the coding block is greater than or equal to 64; specifying whether the width of the coding block is less than 128; and specifying whether the height of the coding block is less than 128; or wherein the mode flag is a triangle flag, and the constraint condition of the triangle flag includes: specifying whether a regular merge flag is not set; specifying whether a merge mode with motion vector difference (MMVD) flag is not set; specifying whether a merge sub-block flag is not set; and specifying whether a mode flag, a combined inter-intra prediction (CIIP) flag, is not set.
3. The method of claim 1, further comprising: signaling the regular merge flag before signaling the mode flag when a constraint condition of the mode flag is satisfied.
4. The method of claim 3, further comprising: signaling the regular merge flag before signaling a merge mode with motion vector difference (MMVD) flag when a constraint condition of the merge mode with motion vector difference (MMVD) flag is satisfied; or wherein the method further comprises: signaling the regular merge flag before signaling a combined inter-intra prediction (CIIP) flag when a constraint condition of the combined inter-intra prediction (CIIP) flag is satisfied.
5. The method of claim 1, further comprising: signaling the regular merge flag after signaling the mode flag when a constraint condition of the mode flag is satisfied.
6. The method of claim 5, further comprising: signaling the regular merge flag after signaling a sub-block merge mode flag when a constraint condition of the sub-block merge mode flag is satisfied.
7. The method of claim 1, further comprising: signaling the regular merge flag using context adaptive binary arithmetic coding (CABAC) with one context; or wherein the method further comprises: signaling the regular merge flag using context adaptive binary arithmetic coding (CABAC) with multiple context models, and a selection of the context models is based on coding information, wherein the coding information includes a size of a current CU.
8. A computing device comprising: one or more processors; a memory coupled to the one or more processors; and a plurality of programs stored in the memory that, when executed by the one or more processors, cause the computing device to perform the method of any of claims 1-7.
9. A non-transitory computer-readable storage medium storing a plurality of programs, wherein the plurality of programs, when executed by one or more processing units of a computing device, cause the computing device to perform the method of any of claims 1-7 to generate a bitstream and store the bitstream with the non-transitory computer-readable storage medium.
10. A computer program product comprising a plurality of program codes for execution by a computing device having one or more processing units, wherein the plurality of program codes, when executed by the one or more processing units, cause the computing device to perform the method of any of claims 1-7.
11. A method of storing a bitstream, wherein, the bitstream is generated according to the method of any of claims 1-7.
Citation Information
Patent Citations
Method for decoding inter predictive encoded motion pictures
CN103370940A
Intra block copy coding with temporal block vector prediction
CN107005708A