Motion vector prediction in video encoding and decoding
SbTMVP and separate CABAC coding models improve video encoding and decoding efficiency and complexity by optimizing motion vector prediction and transformation processes.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-12-02
- Publication Date
- 2026-03-04
AI Technical Summary
Existing video encoding and decoding systems face challenges in achieving improved compression efficiency and reduced complexity, particularly in handling motion vector prediction and transformation processes.
Implementing Sub-block Temporal Motion Vector Prediction (SbTMVP) and separate CABAC coding models for sub-block merge and inter-affine prediction modes to optimize motion vector derivation and encoding processes.
Enhances coding efficiency and reduces decoder complexity by optimizing motion vector prediction and transformation processes, leading to improved video compression and decoding performance.
Smart Images

Figure 2026035769000001_ABST
Abstract
Description
[Technical Field]
[0001] Technical Field TECHNICAL FIELD This disclosure relates to video encoding and decoding. [Background technology]
[0002] background To achieve high compression efficiency, image and video coding schemes typically use prediction and transformation to exploit spatial and temporal redundancy within the video content. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original picture block and the predicted picture block, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To recover the video, the compressed data is decoded by an inverse process corresponding to the prediction, transformation, quantization, and entropy coding. As described below, various modifications and embodiments are envisioned that can provide improvements to video encoding and / or decoding systems, including, but not limited to, improved compression or coding efficiency and / or reduced complexity. Summary of the Invention
[0003] overview In general, an example of an embodiment may include a method including: decoding a first flag included in a coded bitstream, where the first flag is coded using CABAC coding based on a first probability model; decoding a second flag included in the coded bitstream, where the second flag is coded using CABAC coding based on a second probability model; and decoding picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag, where the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode.
[0004] In general, another example of an embodiment may include a method including: determining a value of a first flag indicating a first prediction mode associated with encoding of picture information; determining a value of a second flag indicating a second prediction mode associated with encoding of the picture information; and encoding at least a portion of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model.
[0005] In general, another example of an embodiment may include an apparatus that includes one or more processors configured to: decode a first flag included in a coded bitstream, where the first flag was coded using CABAC coding based on a first probability model; decode a second flag included in the coded bitstream, where the second flag was coded using CABAC coding based on a second probability model; and decode coded picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag, wherein the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode.
[0006] In general, another example of an embodiment may include an apparatus including one or more processors configured to: determine a value of a first flag indicating a first prediction mode associated with encoding of picture information; determine a value of a second flag indicating a second prediction mode associated with encoding of the picture information; and encode at least a portion of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model.
[0007] In general, another example of an embodiment may include a bitstream formatted to include encoded picture information, where the encoded picture information has been encoded by processing the picture information according to any one or more of the example embodiments of the method according to the present disclosure.
[0008] In general, one or more other example embodiments may also provide a computer-readable storage medium, e.g., a non-volatile computer-readable storage medium, having stored thereon instructions for encoding or decoding picture information, such as video data, in accordance with the methods or apparatus described herein. One or more embodiments may also provide a computer-readable storage medium having stored thereon a bitstream generated in accordance with the methods or apparatus described herein. One or more embodiments may also provide methods and apparatus for transmitting or receiving a bitstream generated in accordance with the methods or apparatus described herein.
[0009] Various modifications and embodiments are envisioned as described below that may provide improvements to video encoding and / or decoding systems, including, but not limited to, one or more of improved compression efficiency, and / or improved coding efficiency, and / or improved processing efficiency, and / or reduced complexity.
[0010] The foregoing presents a brief summary of the subject matter to provide a basic understanding of some aspects of the present disclosure. This summary is not an extensive overview of the subject matter. It is not intended to identify key / critical elements of embodiments or to delineate the scope of the subject matter. Its sole purpose is to present some concepts of the subject matter in a simplified form as a prelude to the more detailed description provided below.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS The present disclosure can be better understood from the following detailed description considered in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0012] [Figure 1] 1 provides a block diagram illustrating an example of an embodiment of a video encoder. [Figure 2] 1 provides a block diagram illustrating an example of an embodiment of a video decoder. [Figure 3] 1 illustrates aspects of the present disclosure including a coding tree unit (CTU). [Figure 4] 1 illustrates aspects of the present disclosure including a CTU and a coding unit (CU). [Figure 5] For example, a flow diagram showing an example of an inter mode in VTM-5 is provided. [Figure 6] 1 provides a flow diagram illustrating an example of signaling inter prediction information, for example, according to HEVC. [Figure 7] We provide a block diagram showing an example of the location of spatial and temporal motion vector predictors used in e.g. merge mode in HEVC (left: spatial merge candidate, right: temporal merge candidate). [Figure 8] 1 provides a flow diagram illustrating an example of building a list of merge motion vector predictor candidates, for example in HEVC. [Figure 9] 1 provides another flow diagram illustrating another example of building a list of merge motion vector predictor candidates, e.g., in HEVC. [Figure 10] For example, an example of SbTMVP motion prediction for a CU in JEM is shown below. [Figure 11] 10 shows an example of inter-mode signaling using temporal merge lists. [Figure 12] An example of the location of the motion vector predictor is shown (left: spatial predictor, right: temporal predictor). [Figure 13] 1 illustrates an example of an embodiment of building a list of sub-block-based motion vector predictor candidates. [Figure 14] 1 provides a flow diagram illustrating an example of an embodiment of an inter-decoding process using temporal merge lists. [Figure 15] 1 provides a block diagram illustrating an example of an embodiment of an apparatus or system according to various aspects described herein. [Figure 16] 1 illustrates an example of an embodiment described herein. [Figure 17] 1 illustrates an example of an embodiment described herein. DETAILED DESCRIPTION OF THE INVENTION
[0013] It should be understood that the drawings are intended to illustrate examples of various aspects and embodiments and are not necessarily the only possible configuration. Like reference designators throughout the various drawings refer to the same or similar features.
[0014] Detailed Description Referring now to the drawings, Figure 1 illustrates an example video encoder 100, such as an HEVC encoder. HEVC is a compression standard developed by the Joint Collaborative Team on Video Coding (JCT-VC) (see, e.g., "ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265"). Figure 1 may also illustrate an encoder based on or improved upon the HEVC standard, such as an encoder based on the Joint Exploration Model (JEM) being developed by JVET, or an encoder that improves on the JEM, or an encoder that uses technology similar to HEVC.
[0015] In this application, the terms "reconstructed" and "decoded" may be used interchangeably, the terms "pixel" and "sample" may be used interchangeably, and the terms "picture" and "frame" may be used interchangeably.
[0016] The HEVC specification distinguishes between "blocks" and "units", where a "block" addresses a specific area within a sample array (e.g., luma, Y), and a "unit" contains an ordered block of all coded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data (e.g., motion vectors) associated with the block.
[0017] For coding, a picture is partitioned into square-shaped coding tree blocks (CTBs) with configurable sizes, and sets of consecutive coding tree blocks are grouped into slices. A coding tree unit (CTU) includes the CTB of a coded color component. The CTB is the root of a quadtree that partitions into coding blocks (CBs), which may be partitioned into one or more prediction blocks (PBs) and form the root of a quadtree that partitions into transform blocks (TBs). Corresponding to the coding blocks, prediction blocks, and transform blocks, a coding unit (CU) includes a prediction unit (PU) and a set of tree-structured transform units (TUs), where a PU includes prediction information for all color components, and a TU includes residual coding syntax structures for each color component. The sizes of the CBs, PBs, and TBs for the luma component correspond to the corresponding CUs, PUs, and TUs. In this application, the term "block" may be used to refer to any of the CTUs, CUs, PUs, TUs, CBs, PBs, and TBs. Additionally, "block" may also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally to refer to arrays of data of various sizes.
[0018] In the encoder 100 of FIG. 1, a picture is encoded by the following encoder elements: The picture to be encoded is processed in units of CUs. Each CU is encoded using intra mode or inter mode. If a CU is encoded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to encode the CU, and indicates the intra / inter decision with a prediction mode flag. A prediction residual is calculated by subtracting (110) the predicted block from the original image block.
[0019] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can also skip the transform and apply quantization directly to the untransformed residual signal on a 4x4 TU basis. The encoder can also skip both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process. In direct PCM coding, no prediction is applied and coded unit samples are coded directly into the bitstream.
[0020] The encoder decodes the coded block to provide a reference for further prediction. The quantized transform coefficients are dequantized (140) and inverse transformed (150) to decode the prediction residual. Combining (155) the decoded prediction residual and the predicted block reconstructs the image block. An in-loop filter (165) is applied to the reconstructed picture, e.g., to perform deblocking / Sample Adaptive Offset (SAO) filtering to reduce coding artifacts. The filtered image is stored in a reference picture buffer (180).
[0021] Figure 2 shows a block diagram of an example video decoder 200, such as an HEVC decoder. In the example decoder 200, a signal or bitstream is decoded by the following decoder elements: The video decoder 200 generally performs a decoding path that is the inverse of the encoding path as described in Figure 1, which performs video decoding as part of the encoding of the video data. Figure 2 may also show a decoder that has improvements to the HEVC standard, such as a decoder based on or an improvement to JEM, or a decoder that uses technology similar to HEVC.
[0022] Specifically, the decoder's input includes a video signal or bitstream that may be generated by a video encoder, such as video encoder 100 of FIG. 1. The signal or bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information. To decode the prediction residual, the transform coefficients are dequantized (240) and inverse transformed (250). An image block is reconstructed by combining (255) the decoded prediction residual and the predicted block. The predicted block may be obtained (270) from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275). Advanced Motion Vector Prediction (AMVP) and merge mode techniques may be used to derive motion vectors for motion compensation, which may use an interpolation filter to calculate interpolated values of sub-integer samples of a reference block. An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0023] The HEVC video compression standard uses motion-compensated temporal prediction to exploit redundancy that exists between successive pictures of a video. To do this, a motion vector is associated with each prediction unit (PU). Each coding tree unit (CTU) is represented in the compressed domain by a coding tree (CT), which is a quadtree decomposition of the CTU, with each leaf called a coding unit (CU), as shown in Figure 3.
[0024] Each CU is then given some intra- or inter-prediction parameters (prediction information). To that end, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. The intra- or inter-coding mode is assigned at the CU level, as shown in Figure 4, which shows an example of the division of a coding tree unit into coding units, prediction units, and transform units.
[0025] In HEVC, exactly one motion vector is assigned to each PU. This motion vector is used for motion-compensated temporal prediction of the considered PU. Therefore, in HEVC, the motion model linking a predicted block and its reference block includes translation.
[0026] To encode motion data, HEVC uses two modes, called Adaptive Motion Vector Prediction (AMVP) and Merge, respectively. AMVP involves signaling one or more reference pictures used to predict the current PU, a motion vector predictor index (obtained from a list of two predictors), and a motion vector difference. Generally, at least one embodiment described herein includes the Merge mode.
[0027] The merge mode involves signaling and decoding an index of some motion data collected in a list of motion data predictors. This list consists of five candidates and is constructed identically on the decoder side and the encoder side. Thus, the merge mode aims to derive some motion information obtained from the merge list. The merge list typically contains motion information related to several spatial and temporal surrounding blocks that are available in the decoding state when the current PU is being processed.
[0028] In the Joint Exploration Model (JEM) developed by the Joint Video Exploration Team (JVET) group, several additional temporal prediction tools with associated parameters determined on the decoder side include Sub-block Temporal Motion Vector Prediction (SbTMVP), sometimes called Alternative Temporal Motion Vector Prediction (ATMVP). The basic principle of SbTMVP is to derive certain motion information of the current CU from certain motion information contained in the reference picture of the current picture. To this end, a so-called temporal vector is first obtained as the motion vector of the first merge candidate of the current CU. This temporal vector with its associated reference picture index can be used to search for a motion vector from the reference picture under consideration. To this end, the current CU is divided into NxN sub-CUs, and the motion data indicated by the time from the center position of the considered sub-CU is regarded as the sub-CU MV SbTMVP predictor.
[0029] The SbTMVP motion prediction mode can be included as an additional candidate in the classical merge list. Another approach is to insert the SbTMVP motion prediction mode into the affine merge list, which is originally dedicated to building a set of merge candidates for predicting the affine motion model of the current CU when affine motion compensation is used for the current CU. The use of affine motion compensation for the current CU is signaled by a so-called affine flag at the CU level. It is believed that moving the SbTMVP motion predictor into the affine merge list may result in improved compression efficiency compared to other approaches, such as when the SbTMVP candidate is part of the classical merge candidate list.
[0030] Figure 5 shows an example of intermode coding, e.g. in VTM-5: First, the skip flag is decoded to indicate whether the CU contains residuals. Next, the regular_merge flag is decoded, which indicates the use of regular merging and includes TMVP (Temporal Motion Vector Predictor) candidates. Otherwise, the mmvd flag is decoded, indicating the use of the Merge Motion Vector Difference mode, which may also include candidates constructed from TMVP candidates. - If not MMDV, then the sub-block flags are decoded. Sub-block candidates can be either SbTMVP or affine candidates. -If it is not a sub-block, in case of merge (not skip), the CIIP flag (Combined Intra-Inter Prediction) is decoded. This mode may also include candidates built from TMVP candidates. Finally, if none of the above modes are selected, the triangle mode is inferred, which may also include candidates constructed from TMVP candidates.
[0031] In AMVP mode (Advanced Motion Vector Prediction), blocks can be either translational or affine. Note that the affine flag in AMVP is coded using the same CABAC context as the sub-block flags in skip or merge modes.
[0032] The merge mode in the HEVC standard involves deriving inter prediction information (hereinafter also referred to as motion information) of a given prediction unit from selected motion information predictor candidates. The motion information includes all inter prediction parameters of the PU, i.e., -One-way or two-way time prediction type - Reference picture index within each reference picture list one or more motion vectors Includes:
[0033] The coding and decoding of inter prediction information in HEVC is summarized in Figure 6, which shows the signaling of inter prediction information. As shown in the figure, the coding / decoding of motion information in merge mode occurs in two modes: skip mode and merge mode. In these two modes, one field is signaled to enable a decoder to retrieve the motion information of a PU: the so-called merge index. The merge index indicates which motion vector predictor in the list of merge motion information predictors is used to derive the motion information of the current PU. In the following, the list of motion information predictors is called a merge list or a merge candidate list. And, the candidate motion information predictors are called merge candidates.
[0034] In HEVC, a merge candidate list is systematically built from five merge candidates. The merge list is constructed on both the encoder side and the decoder side as follows: Figure 7 shows the locations of spatial and temporal motion vector predictors used in merge mode. Spatial merge candidates are shown on the left side of Figure 7. Temporal merge candidates are shown on the right side of Figure 7. As shown in Figure 7, up to five spatial locations can be considered to search for several potential candidates. They are in the following order: 1-Left (A1) 2-Top (B1) 3-Top right (B0) 4-Bottom left (A0) 5-Top left (B2) , where symbols A0, A1, B0, B1, and B2 denote the spatial positions shown on the left side of FIG. 7. Four distinct spatial candidates are selected. Next, a TMVP representing a temporal predictor is selected by considering the temporal motion information located at position H, and then the "center" is the candidate at position H in the unavailable considered reference picture. Next, a list of merge motion vector predictor candidates is constructed, as shown in FIG. 8, which shows a final pruning process performed to ensure that the set of selected spatial and temporal candidates does not contain redundant candidates.
[0035] Next, for B slices, if the merge list is not full, another type of candidate is sent to the merge list: a so-called combined candidate. This may involve generating a candidate consisting of motion information associated with one reference picture list (L0) from one candidate already in the list, and this motion associated with another reference picture list (L1) from another candidate already in the merge list. If the merge list is not yet full (five elements), zero motion vectors are sent to the end of the merge list until it is full. The overall process of merge list construction in HEVC is detailed in the diagram of Figure 9.
[0036] In the sub-block temporal motion vector prediction (SbTMVP) method (also called ATMVP), one or several temporal motion vector predictors for a current CU are retrieved from one or more reference pictures for the current CU. First, a so-called temporal motion vector and associated reference picture index are obtained as motion data associated with the first candidate in the current CU's regular merge list candidates. Next, the current CU is divided into N×N sub-CUs (N is generally equal to 4). This is indicated by "Error! Reference Not Found." For each N×N sub-block, one or more motion vectors and one or more reference picture indexes are identified in the reference picture associated with the temporal MV with the help of the temporal motion vector. The N×N sub-blocks in the reference picture indicated by the temporal MV from the current sub-CU position are considered. The motion data is considered as the SbTMVP motion data prediction for the current sub-CU. Then, it is converted into the motion vector and reference picture index of the current sub-CU by appropriate motion vector scaling.
[0037] In general, one aspect of at least one example of an embodiment includes organizing the number and content of various lists of motion data predictors in a manner that improves coding efficiency. In general, another aspect of at least one embodiment includes reducing complexity to improve worst-case performance on the decoder side, for example, by reducing the number of merge candidates to construct for a given mode. In general, another aspect of at least one embodiment includes making affine mode coding more consistent by not intermixing sub-block flag coding and affine flag coding.
[0038] At least one example of an embodiment provides a temporal merge list that includes one or more temporal predictors, such as SbTMVP, TMVP, MMVD based on TMVP, CIIP based on TMVP, and triangle based on TMVP, and these candidates are removed from the conventional merge list, thereby reducing the worst-case complexity at the decoder.
[0039] At least one example of an embodiment includes that temporal modes not included in the temporal merge list are maintained in their respective lists (eg, MMVD, CIIP, triangle).
[0040] At least one example of an embodiment includes that temporal modes not included in a temporal merge list are completely removed, leaving each list with no temporal predictors.
[0041] At least one example of an embodiment includes extracting SbTMVP from sub-blocks and providing consistent coding of affine flags between merge / skip mode and AMVP mode.
[0042] At least one example of an embodiment includes separating the sub-block and affine flag coding.
[0043] Figure 11 shows an example of a classical merge list construction (as opposed to sub-block). Classical here refers to a merge list used for translational motion compensated temporal prediction, where one motion vector per reference picture list is associated with one CU. This process involves: - Addition of spatial candidates A1, B1, B0 A0 B2 (some pruning is performed between them) (see Figure 13 for predictor locations) -Added time candidate (TMVP) (H or center) Addition of HMVP (History based Motion Vector Predictor) candidates, with at most some pruning for A1 and B1 for the first additional candidate. - generating a pair of (averaged) candidates if there are at least two candidates in the list; - stuffing the list with null motion vectors if necessary, Includes:
[0044] An example of constructing a list of sub-block-based motion predictor candidates is now provided. The affine merge list collects all merge MV predictors, including sub-block-based motion-compensated temporal prediction. Therefore, it includes the above-mentioned SbTMVP merge mode and affine merge mode. The affine merge candidate represents a merge candidate from which a 4x4 block-based affine motion field is derived for the current CU and used for temporal prediction of the current CU. Affine motion compensation will not be described in detail here. One aspect to note here is that, in at least one example, the SbTMVP candidate is placed at the beginning of the affine merge list.
[0045] Other examples of merge lists may include: The MMVD merge list consists of the first two classical merge candidates displaced by a given motion vector difference signaled in the bitstream. Therefore, this list may include candidates based on temporal candidates (displaced TMVP candidates). The CIIP merge list is the same as the classical merge list. Therefore, this list may include candidates based on temporal candidates. The final predictor is a mixture of motion-based predictor and intra-prediction. The triangle merge list consists of candidates generated from the candidates in the classical merge list. This list may therefore include candidates based on temporal candidates.
[0046] In general, at least one aspect of at least one example of an embodiment described herein includes generating a merge list that includes temporal candidates.
[0047] FIG. 14 illustrates an example of an embodiment of a modification process for decoding a temporal merge list that includes at least the following: - After decoding the mmvd mode flag, if the mmvd flag is false, the time mode flag is decoded. - If temporal mode is true, the temporal candidate index is decoded similarly to the classical merge candidate index - Otherwise, the affine mode flag is decoded. Note that the affine flag replaces the sub-block mode flag because the SbTMVP candidate has been removed from the list. Thanks to this division, the affine flag coding in merged mode and AMVP mode uses the same CABAC bins in a consistent manner (as opposed to the way the sub-block flag meant both SbTMVP or affine mode).
[0048] In general, at least one aspect of at least one example of an embodiment may include providing various arrangements of time merge lists. For example, the time merge list may be: -As the first merge list -Normally after the merge list -Affine merge list after -ciip merge list after -After the triangle merge list (in this case, if triangle mode is false, time mode is deduced to be true) It can also be inserted.
[0049] In general, the list of temporal candidates can be generated in the same way that a merge list is generated, by adding potential candidates (if any) in a fixed order until a maximum number of merge candidates is reached (typically 2 or 3). An index indicating the position in the list of temporal candidates to use is then transmitted to the decoder. An example list of temporal candidates is: -SbTMVP candidate -TMVP candidate -Time CIIP candidate is.
[0050] Variations may include one or more of the following: - TMVP candidates for MMVD and triangle mode may also be added to the list by further signaling, for which further information (typically mmvd flag and mmvd displacement) is required. - TMVP candidates for MMVD and / or triangle and / or CIIP are kept in their original list. - TMVP candidates related to MMVD and / or triangles and / or CIIP are completely removed, thereby leaving all other lists with no temporal candidates. Optionally, in this variant, the recovered affine candidates do not use temporal motion vectors, leaving the affine merge list with no temporal motion vectors. - Pure time-restored affine candidates are also added to the time list.
[0051] In general, one or more of the described embodiments may provide: - A reduction in the complexity of the decoder process where a given merge list (typically mmvd, temporal, affine, ciip, or triangular) has a reduction in worst-case complexity (e.g., one or more lengths of one or more lists are reduced). - The bandwidth required to build each candidate in the list is reduced because all spatial candidates are in a separate list, without the need to access the temporal memory buffer. Generally, in all lists except the temporal list, only the spatial information around the current block needs to be cached.
[0052] Table 1 attached to this disclosure provides an example of one embodiment of syntax for encoding the temporal merge list flags. The example shown in Table 1 includes providing syntax for encoding the temporal merge list flags after the MMVD flag. Note that the sub-block syntax is replaced by affine syntax as seen in the AMVP derivation (shaded in gray).
[0053] For SbTMVP and TMVP temporal merge lists only, the variable MaxNumTemporalMergeCand is at most equal to 2 if both candidates are available. Candidate availability is derived as before, - SbTMVP is available if the SPS level flag of SbTMVP is true and the current block size is 8 or more for both width and height. - TMVP is available if the time flag at the SPS level is true and width + height is greater than 12.
[0054] Another aspect of at least one example of an embodiment may include affine and sub-block flag coding. In certain systems, the coding of merge_subblock_flag (in merge / skip mode) and inter_affine_flag (in AMVP) may include sharing the same CABAC coding. In the case of separate temporal lists, such as in at least one embodiment described herein (i.e., SbTMVP is no longer in the sub-block list), merge_subblock_flag actually becomes merge_affine_flag, so that the CABAC coding is more consistent. Another variation that does not use separate temporal merge lists is to use two separate CABAC codings (one for merge_subblock_flag and one for inter_affine_flag).
[0055] Generally, another example of an embodiment is shown in Figure 16. In Figure 16, a bitstream including coded picture information also includes coded control information, such as a flag. At 1610, a first flag included in the coded bitstream is decoded using CABAC coding based on a first probability model. At 1620, a second flag included in the coded bitstream is decoded using CABAC coding based on a second probability model. That is, two separate CABAC codings are used to decode the first and second flags. Then, at 1630, coded picture information included in the bitstream is decoded based on a coding mode indicated by the first flag or the second flag to generate decoded picture information. For example, the first flag may indicate, correspond to, or be associated with a coding mode such as a sub-block merge prediction mode, and the second flag may indicate, correspond to, or be associated with a coding mode such as an inter-affine prediction mode.
[0056] Generally, another example of an embodiment is shown in Figure 17. In Figure 17, at 1710, a value of a first flag is determined, the first flag indicating a first prediction mode associated with encoding of picture information included in the input. At 1720, a value of a second flag is determined, the second flag indicating a second prediction mode associated with encoding of the picture information. At 1730, at least a portion of the picture information and the first and second flags are coded to generate a coded bitstream, where the first flag indicates, for example, a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates, for example, an inter-affine prediction mode and is coded using CABAC coding based on a second probability model. That is, two separate CABAC codings are used to code the first and second flags.
[0057] This specification describes various example embodiments, features, models, methods, and the like. Many such examples are described with specificity, often in a manner that may be considered limiting, at least to illustrate individual characteristics. However, this is for the purpose of clarity of description and not to limit its application or scope. Indeed, the various example embodiments, features, and the like described herein may be combined and interchanged in various ways to provide further example embodiments. Example embodiments according to the present disclosure include, but are not limited to, the following:
[0058] In general, an example of an embodiment may include a method including: decoding a first flag included in a coded bitstream, where the first flag is coded using CABAC coding based on a first probability model; decoding a second flag included in the coded bitstream, where the second flag is coded using CABAC coding based on a second probability model; and decoding picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag, where the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode.
[0059] In general, another example of an embodiment may include a method including: determining a value of a first flag indicating a first prediction mode associated with encoding of picture information; determining a value of a second flag indicating a second prediction mode associated with encoding of the picture information; and encoding at least a portion of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model.
[0060] In general, another example of an embodiment may include an apparatus that includes one or more processors configured to: decode a first flag included in a coded bitstream, where the first flag was coded using CABAC coding based on a first probability model; decode a second flag included in the coded bitstream, where the second flag was coded using CABAC coding based on a second probability model; and decode coded picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag, wherein the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode.
[0061] In general, another example of an embodiment may include an apparatus including one or more processors configured to: determine a value of a first flag indicating a first prediction mode associated with encoding of picture information; determine a value of a second flag indicating a second prediction mode associated with encoding of the picture information; and encode at least a portion of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model.
[0062] In general, another example of an embodiment may include a bitstream formatted to include encoded picture information, where the encoded picture information has been encoded by processing the picture information according to any one or more of the example embodiments of the method according to the present disclosure.
[0063] In general, one or more other example embodiments may also provide a computer-readable storage medium, e.g., a non-volatile computer-readable storage medium, having stored thereon instructions for encoding or decoding picture information, such as video data, in accordance with the methods or apparatus described herein.
[0064] In general, at least one example of an embodiment may include a computer program including instructions that, when executed by a computer, cause the computer to perform a method according to one or more example embodiments described herein.
[0065] In general, at least one example of an embodiment may include a non-transitory computer-readable medium storing executable program instructions that cause a computer executing the instructions to perform a method according to one or more example embodiments described herein.
[0066] In general, at least one example of an embodiment may include a signal including data generated according to any one or more example embodiments described herein.
[0067] In general, at least one example of an embodiment may include a bitstream formatted to include syntax elements and encoded image information generated in accordance with any one or more of the example embodiments described herein.
[0068] In general, at least one example of an embodiment may include a computer-readable storage medium having stored thereon a bitstream generated according to the methods or apparatus described herein.
[0069] In general, at least one example of an embodiment may involve transmitting or receiving a bitstream or signal generated in accordance with the methods or apparatus described herein.
[0070] In general, at least one example of an embodiment may include a device including an apparatus according to any one or more of the example embodiments described herein and at least one of: (i) an antenna configured to receive a signal, the signal including data representing image information; (ii) a band limiter configured to limit the received signal to a frequency band including the data representing the image information; and (iii) a display configured to display an image from the image information.
[0071] Generally, at least one example of an embodiment may include a device as described herein, where the device includes one of a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a mobile phone, a tablet, or other electronic device.
[0072] In general, another example of an embodiment may include an apparatus including one or more processors configured to: determine a value of a first flag indicating a first prediction mode associated with encoding of picture information; determine a value of a second flag indicating a second prediction mode associated with encoding of the picture information; and encode at least a portion of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model.
[0073] In general, example embodiments described and contemplated herein can be implemented in many different forms. While Figures 1 and 2 above and Figure 15 below provide some embodiments, other embodiments are contemplated, and the descriptions of Figures 1, 2, and 15 do not limit the breadth of implementations. At least one embodiment provides examples generally related to encoding and / or decoding video, and at least one other embodiment generally related to transmitting a generated or encoded bitstream or signal. These and other embodiments may be implemented as methods, apparatus, computer-readable storage media having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media having stored thereon a bitstream or signal generated according to any of the described methods.
[0074] The terms HDR (high dynamic range) and SDR (standard dynamic range) are used in this disclosure. These terms often convey specific values of dynamic range to those skilled in the art. However, further embodiments are contemplated in which reference to HDR is understood to mean "higher dynamic range" and reference to SDR is understood to mean "lower dynamic range." Such further embodiments are not constrained by any specific values of dynamic range that may often be associated with the terms "high dynamic range" and "standard dynamic range."
[0075] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the correct operation of the method, the order and / or use of specific steps and / or actions may be modified or combined.
[0076] Various methods and other aspects described herein can be used to modify modules, e.g., motion estimation and / or motion compensation modules, such as modules 170 and / or 175, of a video encoder, such as encoder 100 shown in FIG. 1, and / or motion compensation module 275 of video decoder 200 shown in FIG. 2. Additionally, the aspects are not limited to VVC or HEVC, but may be applied, for example, to other standards and recommendations, whether existing or later developed, and extensions of any such standards and recommendations (including VVC and HEVC). Unless indicated otherwise or technically precluded, aspects described herein may be used individually or in combination.
[0077] For example, various numerical values are used herein. The particular values are for illustrative purposes, and the aspects described are not limited to these particular values.
[0078] FIG. 15 shows a block diagram of an example system in which various aspects and embodiments may be implemented. System 1000 may be embodied as a device including the various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected consumer electronics, and servers. The elements of system 1000, singly or in combination, may be embodied in a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or separate components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described herein.
[0079] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 1010 may include embedded memory, input / output interfaces, and various other circuitry known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 1040 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0080] System 1000 includes an encoder / decoder module 1030 configured to process data to provide, for example, encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 1030 represents one or more modules that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Furthermore, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0081] Program code loaded into the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and later loaded into the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items while performing the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams or signals, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operational logic.
[0082] In some embodiments, memory internal to the processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory may be, for example, memory 1020 and / or storage device 1040, such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video coding and decoding operations, such as MPEG-2, HEVC, or Versatile Video Coding (VVC).
[0083] Input to the elements of system 1000 may be provided by various input devices as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals transmitted wirelessly, for example, by broadcast equipment, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0084] In various embodiments, the input devices of block 1130 have associated respective input processing elements known in the art. For example, the RF section may be associated with elements for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which may be referred to as a channel in certain embodiments (for example), (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a bandlimiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs some of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency near baseband) or to baseband. In some set-top box embodiments, the RF section and associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, downconverting, and filtering again to a desired frequency band. Various embodiments rearrange the order of the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements may include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0085] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via USB and / or HDMI connections. It is understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented in a separate input processing IC or within processor 1010. Similarly, aspects of USB or HDMI interface processing may be implemented in a separate interface IC or within processor 1010. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including processor 1010 and encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data stream for presentation on an output device.
[0086] The various elements of system 1000 may be provided within an integrated housing in which the various elements may be interconnected and data may be transmitted therebetween using a suitable connection arrangement 1140, such as an internal bus known in the art, including, for example, an I2C bus, wiring, and printed circuit boards.
[0087] System 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. Communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over communication channel 1060. Communication interface 1050 may also include, but is not limited to, a modem or a network card, and communication channel 1060 may be implemented, for example, in a wired and / or wireless medium.
[0088] In various embodiments, data is streamed to system 1000 using a Wi-Fi network, such as IEEE 802.11. The Wi-Fi signal in these embodiments is received over communication channel 1060 and communication interface 1050, which are adapted for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, enabling streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that delivers data over an HDMI connection in input block 1130. Still other embodiments provide streamed data to system 1000 using an RF connection in input block 1130.
[0089] System 1000 can provide output signals to various output devices, including display 1100, speakers 1110, and other peripheral devices 1120. Other peripheral devices 1120, in various example embodiments, include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, CEC, or other communication protocols that enable inter-device control with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections using respective interfaces 1070, 1080, and 1090. Alternatively, output devices may be connected to system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speakers 1110 may be integrated with other components of the system 1000 in a single unit, for example, in an electronic device such as a television. In various embodiments, the display interface 1070 includes a display driver, such as, for example, a timing controller (T Con) chip.
[0090] Display 1100 and speakers 1110 may alternatively be separate from one or more of the other components, for example, if the RF portion of input 1130 is part of a separate set-top box. In various embodiments in which display 1100 and speakers 1110 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0091] The embodiments may be implemented by computer software implemented by the processor 1010, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technical environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0092] Throughout this disclosure, various implementations include decoding. As used herein, "decoding" may encompass all or some of the processes performed on a received encoded sequence to generate a final output suitable for display, for example. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes additionally or alternatively include processes performed by decoders of various implementations described herein, such as extracting a picture from a non-overlapping (padded) picture, determining an upsample filter to use and then upsampling the picture, and flipping the picture back to its intended orientation.
[0093] As a further example, in one embodiment, "decoding" refers to entropy decoding only, in another embodiment, "decoding" refers to differential decoding only, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the term "decoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.
[0094] Various implementations also include encoding. Similar to the above description of "decoding," "encoding," as used herein, may encompass all or some of the processes performed on an input video sequence to generate, for example, an encoded bitstream or signal. In various embodiments, such processes include one or more of the processes typically performed by an encoder, such as, for example, partitioning, differential encoding, transforming, quantizing, and entropy coding. In various embodiments, such processes may additionally or alternatively include processes performed by the encoders of various implementations described herein.
[0095] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential and entropy encoding. Whether the expression "encoding process" is intended to refer specifically to a subset of operations or to refer generally to a broader encoding process will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.
[0096] It should be noted that the syntax elements used herein are descriptive terms, so they do not preclude the use of other syntax element names.
[0097] Where a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, where a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.
[0098] Various embodiments refer to rate-distortion optimization. Specifically, during the encoding process, a balance or trade-off between rate and distortion is typically considered, often given a computational complexity constraint. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist for solving the rate-distortion optimization problem. For example, these approaches may be based on extensive testing of all encoding options, including all considered modes or coding parameter values, with a thorough evaluation of the coding cost and associated distortion of the reconstructed signal after coding and decoding. To reduce coding complexity, faster approaches may also be used, specifically using approximate distortion calculations based on predicted or prediction residual signals rather than the reconstructed ones. A hybrid of these two approaches may also be used, such as by using approximate distortion for only some of the possible encoding options and full distortion for others. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches utilize any of a variety of techniques for performing optimization, but the optimization does not necessarily involve a thorough evaluation of both the coding cost and associated distortion.
[0099] The embodiments and aspects described herein may be implemented in, for example, a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of embodiment (e.g., described only as a method), the described feature implementation may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented in, for example, appropriate hardware, software, and firmware. A method may be implemented in, for example, a processor (which generally refers to processing devices including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device). Processors also include, for example, communication devices such as computers, mobile phones, personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0100] References to "one embodiment," or "an embodiment," or "one implementation," or "an implementation," and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with an embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment," or "in an embodiment," or "in one embodiment," or "in an embodiment," and other variations thereof, in various places throughout this specification are not necessarily all referring to the same embodiment.
[0101] Additionally, this specification may refer to "obtaining" various information. Obtaining information may include, for example, one or more of determining information, estimating information, calculating information, predicting information, or retrieving information from a memory.
[0102] Additionally, this specification may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0103] Additionally, this specification may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" generally includes in some way, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0104] For example, the use of either " / ", "and / or", and "at least one of" in the cases of "A / B", "A and / or B", and "at least one of A and B" shall be understood to be intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such language is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded for as many items as there are to be listed, as will be apparent to those skilled in this and related arts.
[0105] Also, in this specification, the term "signal" refers, among other things, to indicating something to a corresponding decoder. For example, in certain embodiments, an encoder signals a specific one of multiple parameters for improvement purposes. Thus, in some embodiments, the same parameter is used on both the encoder and decoder sides. Thus, for example, an encoder can transmit a specific parameter to a decoder so that the decoder can use the same specific parameter (explicit signaling). Conversely, if the decoder already has the specific parameter in addition to other parameters, signaling without transmission (implicit signaling) can be used to simply allow the decoder to recognize and select the specific parameter. By avoiding transmitting the actual function, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various ways. For example, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder in various embodiments. Although the above relates to the verb form of the term "signal," the term "signal" may also be used herein as a noun.
[0106] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be, for example, stored or transmitted. Information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream or signal of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0107] Various embodiments have been described. The embodiments may include any of the following features or entities, alone or in any combination, across a variety of different claim categories and types. In an encoder and / or decoder for processing video, providing at least one merge mode based on a merge list containing temporal candidates. In an encoder and / or decoder for processing video, providing at least one merge mode based on organizing the number and content of various lists of motion data predictors in a manner that improves coding efficiency. In an encoder and / or decoder for processing video, providing based on reducing complexity to improve the worst case on the decoder side, for example by reducing the number of merging candidates to construct for a given mode. To provide, in an encoder and / or decoder for processing video, improved consistency of coding in affine mode based on not mixing sub-block flag coding and affine flag coding. In an encoder and / or decoder for processing video, providing a temporal merge list based on the temporal merge list, the temporal merge list including one or more temporal predictors including one or more of SbTMVP, or TMVP, or TMVP-based MMVD, or TMVP-based CIIP, or TMVP-based triangle, where these candidates are removed from a conventional merge list, thereby reducing worst-case complexity at the decoder. In an encoder and / or decoder for processing video, providing based on at least one temporal mode that is not included in the temporal merge list and that is maintained in a respective list (e.g. MMVD, CIIP, triangle). In an encoder and / or decoder for processing video, providing based on at least one temporal mode that is not included in the temporal merge list and is completely removed so that each list has no temporal predictor. In an encoder and / or decoder for processing video, based on extracting SbTMVP from sub-blocks and providing consistent coding of affine flags between merge / skip mode and AMVP mode, Providing, in an encoder and / or decoder for processing video, sub-block coding and affine flag coding based on separation. In an encoder and / or decoder for processing video, After decoding the mmvd mode flag, if the mmvd flag is false, the time mode flag is decoded; If the temporal mode is true, the temporal candidate index is decoded; If temporal mode is not true, the affine mode flag is decoded, where the affine flag replaces the sub-block mode flag based on the SbTMVP candidate removed from the merge list, and the affine flag coding in merge mode and AMVP mode uses the same CABAC bins in a consistent manner; Based on the modified process for decoding the time merge list, the method includes: In an encoder and / or decoder for processing video, various arrangements of a temporal merge list are provided, wherein the temporal merge list comprises: As the first merge list, or Usually after the merge list, or After the affine merge list, or After the CIIP merge list, or ○ After the triangle merge list (in this case, if triangle mode is false, time mode is deduced to be true), It is also possible to insert, based on providing, providing. In an encoder and / or decoder for processing video, a list of temporal candidates generated by adding potential candidates in a fixed order until a maximum number of merge candidates is reached, an index indicating the position in the list of temporal candidates to be used is then transmitted to the decoder, and the potential candidates are ○SbTMVP candidate, or ○TMVP candidate, or ○Time CIIP candidate providing a time based on a list of candidate times including one or more of: In an encoder and / or decoder for processing video, ○ Signaling that a TMVP candidate for MMVD or a triangle mode candidate can be added to the merge list, or ○ TMVP candidates for MMVD and / or Triangle and / or CIIP are maintained in their respective original lists, or ○ TMVP candidates for MMVD and / or triangle and / or CIIP are completely removed, so that all other lists have no temporal candidates, or ○ TMVP candidates for MMVD and / or triangles and / or CIIP are completely removed, thereby leaving all other lists with no temporal candidates, and in this case the affine merge list with no temporal motion vectors, so that the recovered affine candidates do not use temporal motion vectors, or Pure time-restored affine candidates are added to the time list, Based on the generated merge list, In an encoder and / or decoder for processing video, providing a reduction in the complexity of a decoder process based on a given merge list having a worst-case complexity reduction. In an encoder and / or decoder for processing video, providing a reduction in the complexity of a decoder process based on a given merge list having a worst-case complexity reduction, including reducing the length of one or more merge lists. In an encoder and / or decoder for processing video, providing a reduction in the complexity of a decoder process, where a given merge list has a worst-case complexity reduction including reducing the length of one or more merge lists, and the given merge list includes at least one of normal, mmvd, temporal, affine, ciip, or triangular. In an encoder and / or decoder for processing video, the bandwidth associated with constructing each candidate in a merge list is reduced because all spatial candidates are in a separate list and no access to a temporal memory buffer is required. Inserting syntax elements into the signaling that allow the encoder to provide at least one merging mode based on the merging candidate list described herein. Selecting the merge mode to apply at the decoder based on these syntax elements. A bitstream or signal containing one or more of the described syntax elements or variations thereof. Inserting syntax elements into the signaling that allow the decoder to provide merge modes in a manner that corresponds to the manner used by the encoder. · Generating and / or transmitting and / or receiving and / or decoding bitstreams or signals that include one or more of the described syntax elements or variations thereof. A TV, set-top box, mobile phone, tablet, or other electronic device that encodes and / or decodes video according to any of the described embodiments and displays the resulting images (e.g., using a monitor, screen, or other type of display). A TV, set-top box, mobile phone, tablet, or other electronic device that tunes to a channel (e.g., using a tuner) to receive a signal containing encoded images and performs video encoding and / or decoding according to any of the described embodiments. A TV, set-top box, mobile phone, tablet, or other electronic device that receives a signal containing an encoded image wirelessly (e.g., using an antenna) and encodes and / or decodes the video according to any of the described embodiments. A computer program storing a program code which, when executed by a computer, performs the encoding and / or decoding of video according to any of the described embodiments. A non-transitory computer-readable medium containing executable program instructions that, when executed, cause a computer to perform video encoding and / or decoding according to any of the described embodiments.
[0108] Various other generalized and specialized embodiments are also supported and contemplated throughout this disclosure.
[0109] [Table 1]
Claims
1. decoding a first flag included in a coded bitstream, the first flag being coded using CABAC coding based on a first probability model; decoding a second flag included in the coded bitstream, the second flag being coded using CABAC coding based on a second probability model; and decoding coded picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag; Including, the first flag indicates a sub-block merge prediction mode; and the second flag indicates an inter-affine prediction mode; method.
2. determining a value of a first flag indicating a first prediction mode associated with encoding of the picture information; determining a value of a second flag indicating a second prediction mode associated with encoding of the picture information; encoding at least a portion of the picture information and the first flag and the second flag to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model; A method comprising:
3. decoding a first flag included in a coded bitstream, the first flag being coded using CABAC coding based on a first probability model; decoding a second flag included in the coded bitstream, the second flag being coded using CABAC coding based on a second probability model; and decoding coded picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag; one or more processors configured to: the first flag indicates a sub-block merge prediction mode; and the second flag indicates an inter-affine prediction mode; Device.
4. determining a value of a first flag indicating a first prediction mode associated with encoding of the picture information; determining a value of a second flag indicating a second prediction mode associated with encoding of the picture information; encoding at least a portion of the picture information and the first flag and the second flag to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model; 1. An apparatus comprising: one or more processors configured to:
5. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method of claim 1 or 2.
6. A non-transitory computer readable medium having stored thereon executable program instructions that cause a computer that executes said instructions to perform the method of claim 1 or 2.
7. A signal comprising data generated according to the method of claim 2.
8. A bitstream formatted to include syntax elements and image information encoded according to the method of claim 2.
9. An apparatus according to claim 3 or 4; at least one of: (i) an antenna configured to receive a signal, the signal including data representing the image information; (ii) a band limiter configured to limit the received signal to a frequency band including the data representing the image information; and (iii) a display configured to display an image from the image information; Devices containing:
10. The device of claim 10 , wherein the device comprises one of a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a mobile phone, a tablet, or other electronic device.
Citation Information
Patent Citations
Motion vector prediction in video encoding and decoding
WO2020061395A1