Motion vector prediction in video encoding and decoding
By using separate CABAC models for flags indicating sub-block merge and inter-affine prediction modes, the video encoding and decoding systems achieve improved coding efficiency and reduced complexity in motion vector prediction.
Patent Information
- Application Number
- JP2025047635
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-06-25
- Filing Date
- 2025-03-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing video encoding and decoding systems face challenges in achieving high compression efficiency and reduced complexity while effectively utilizing spatial and temporal redundancies in video content.
Implementing CABAC coding for flags indicating sub-block merge and inter-affine prediction modes, with separate probability models for each flag, to enhance the construction of motion vector prediction lists and improve coding efficiency.
This approach reduces decoder complexity and improves coding efficiency by optimizing the construction of merge candidate lists, leading to more efficient video encoding and decoding processes.
Smart Images

Figure 2025098127000001_ABST
Abstract
Description
Technical Field
[0001] Technical Field The present disclosure relates to video encoding and decoding.
Background Art
[0002] Background To achieve high compression efficiency, video coding systems typically use prediction and transformation to exploit spatial and temporal redundancies within video content. Generally, intra or inter prediction is used to exploit intra or inter-frame correlation, and then the difference between the original picture block and the predicted picture block, often called the prediction error or prediction residue, is transformed, quantized, and entropy-coded. To restore the video, the compressed data is decoded by inverse processes corresponding to prediction, transformation, quantization, and entropy coding. As will be described below, various modifications and embodiments are envisioned that can provide improvements to video encoding and / or decoding systems, including, but not limited to, one or both of improved compression or coding efficiency and reduced complexity.
Summary of the Invention
[0003] Summary Generally, an example of an embodiment includes decoding a first flag included in a coded bitstream, the first flag being coded using CABAC coding based on a first probability model, decoding a second flag included in the coded bitstream, the second flag being coded using CABAC coding based on a second probability model, and decoding picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag, where the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode.
[0004] Generally, another example of an embodiment may include determining a value of a first flag indicating a first prediction mode related to encoding of picture information, determining a value of a second flag indicating a second prediction mode related to encoding of picture information, and encoding at least a portion of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is encoded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is encoded using CABAC coding based on a second probability model.
[0005] Generally, another example of an embodiment may include one or more processors configured to decode a first flag included in an encoded bitstream, wherein the first flag is encoded using CABAC coding based on a first probability model, decode a second flag included in the encoded bitstream, wherein the second flag is encoded using CABAC coding based on a second probability model, and decode encoded picture information included in the encoded bitstream based on an encoding mode indicated by the first flag or the second flag, wherein the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode.
[0006] Generally, another example of an embodiment includes determining a value of a first flag indicating a first prediction mode related to encoding of picture information, determining a value of a second flag indicating a second prediction mode related to encoding of picture information, and encoding at least a part of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is encoded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is encoded using CABAC coding based on a second probability model, and an apparatus including one or more processors configured to perform the encoding.
[0007] Generally, another example of an embodiment may include a bitstream formatted to include encoded picture information, wherein the encoded picture information is encoded by processing picture information based on any one or more of the example embodiments of the method according to the present disclosure.
[0008] Generally, one or more other examples of embodiments may also provide a computer-readable storage medium storing instructions for encoding or decoding picture information, such as video data, according to the method or apparatus described herein, for example, a non-volatile computer-readable storage medium. One or more embodiments may also provide a computer-readable storage medium storing a bitstream generated according to the method or apparatus described herein. One or more embodiments may also provide a method and apparatus for transmitting or receiving a bitstream generated according to the method or apparatus described herein.
[0009] Various modifications and embodiments can be envisioned that provide improvements to video encoding and / or decoding systems, including one or more of (but not limited to) improved compression efficiency, and / or improved coding efficiency, and / or improved processing efficiency, and / or reduced complexity, as described below.
[0010] The above presents a brief overview of the subject matter to provide a basic understanding of some aspects of the present disclosure. This overview is not a broad summary of the subject matter. It is not intended to identify key / important elements of the embodiments or to elaborate on the scope of the subject matter. Its sole purpose is to present some concepts of the subject matter in a simplified form as a prelude to the more detailed description provided below.
[0011] Brief Description of the Drawings The present disclosure can be better understood by considering the following detailed description in conjunction with the accompanying drawings.
Brief Description of the Drawings
[0012]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Best Mode for Carrying Out the Invention
[0013] It should be understood that the drawings are for the purpose of showing examples of various aspects and embodiments and are not necessarily the only possible configurations. Throughout the various drawings, like reference indicators refer to the same or similar features.
[0014] Detailed Description Referring now to the drawings, FIG. 1 shows an example of a video encoder 100 such as an HEVC encoder. HEVC is a compression standard developed by JCT-VC (Joint Collaborative Team on Video Coding) (see, for example, "ITU-T H.265 TELECOMMUNICATION STANDARDIZATION SECTOR OF ITU (10 / 2014), SERIES H: AUDIOVISUAL AND MULTIMEDIA SYSTEMS, Infrastructure of audiovisual services - Coding of moving video, High efficiency video coding, Recommendation ITU-T H.265"). FIG. 1 can also show an encoder improved with respect to the HEVC standard, such as an encoder based on JEM (Joint Exploration Model) that JVET is developing, or an encoder that has improved JEM, or an encoder using a technology similar to HEVC.
[0015] In the present application, the terms "reconstructed" and "decoded" may be used synonymously, the terms "pixel" and "sample" may be used synonymously, and the terms "picture" and "frame" may be used synonymously.
[0016] The HEVC specification distinguishes between "block" and "unit", where "block" addresses a specific area (e.g., luma, Y) within the sample array, and "unit" includes the arrayed blocks of all encoded color components (Y, Cb, Cr, or monochrome), syntax elements, and prediction data (e.g., motion vectors) associated with the block.
[0017] For coding, the picture is partitioned into square-shaped coding tree blocks (CTBs) having a configurable size, and a set of consecutive coding tree blocks is grouped into slices. A coding tree unit (CTU) contains CTBs of coded color components. A CTB is the root of a quadtree partitioned into coding blocks (CBs), and a coding block may be partitioned into one or more prediction blocks (PBs) and forms the root of a quadtree partitioned into transform blocks (TBs). Corresponding to the coding block, prediction block, and transform block, a coding unit (CU) contains a set of prediction units (PUs) and tree-structured transform units (TUs), where a PU contains prediction information regarding all color components, and a TU contains a residual coding syntax structure regarding each color component. The sizes of the CB, PB, and TB of the luma component correspond to the corresponding CU, PU, and TU. In this application, the term "block" may be used to refer to any of a CTU, CU, PU, TU, CB, PB, and TB. Further, "block" may also be used to refer to macroblocks and partitions as specified in H.264 / AVC or other video coding standards, and more generally, to refer to an array of data of various sizes.
[0018] In the encoder 100 of FIG. 1, the picture is coded by encoder elements as follows. The picture to be coded is processed in units of CUs. Each CU is coded using either an intra mode or an inter mode. When a CU is coded in the intra mode, it performs intra prediction (160). In the inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) whether to use the intra mode or the inter mode for coding the CU, and indicates the intra / inter decision by a prediction mode flag. A prediction residual is calculated by subtracting (110) the block predicted from the original image block.
[0019] Next, the prediction residuals are transformed (125) and quantized (130). To output a bitstream, the quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145). The encoder can also skip the transformation and directly apply quantization to the untransformed residual signal on a 4×4 TU basis. The encoder can also bypass both transformation and quantization, i.e., the residuals are directly coded without applying the transformation or quantization process. In direct PCM coding, no prediction is applied and the coded unit samples are directly coded into the bitstream.
[0020] The encoder decodes the coded blocks to provide a reference for further prediction. To decode the prediction residuals, the quantized transform coefficients are inverse quantized (140) and inverse transformed (150). An image block is restored by combining the decoded prediction residuals and the predicted block (155). For example, an in-loop filter (165) is applied to the restored picture to perform deblocking / SAO (Sample Adaptive Offset) filtering to reduce coding artifacts. The filtered image is stored in the reference picture buffer (180).
[0021] FIG. 2 shows a block diagram of an example of a video decoder 200 such as an HEVC decoder. In the exemplary decoder 200, a signal or bitstream is decoded by decoder elements as follows. The video decoder 200 generally performs a decoding path that is inverse to the encoding path as described in FIG. 1, which performs video decoding as part of the encoding of video data. FIG. 2 can also show a decoder that has been improved with respect to the HEVC standard, such as a decoder based on JEM or a decoder that has improved JEM, or a decoder that uses a technology similar to HEVC.
[0022] Specifically, the input to the decoder includes a video signal or bitstream that can be generated by a video encoder such as the video encoder 100 of FIG. 1. The signal or bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information. The transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. By combining the decoded prediction residual and the predicted block (255), the image block is restored. The predicted block can be obtained from intra prediction (260) or motion compensation prediction (i.e., inter prediction) (275) (270). AMVP (Advanced Motion Vector Prediction) and merge mode techniques may be used to derive motion vectors for motion compensation, and an interpolation filter may be used to calculate the interpolation value of the sub-integer samples of the reference block. The in-loop filter (265) is applied to the restored image. The filtered image is stored in the reference picture buffer (280).
[0023] In the HEVC video compression standard, motion compensation temporal prediction is used to utilize the redundancy existing between consecutive pictures of a video. For this purpose, motion vectors are associated with each prediction unit (PU). Each coding tree unit (CTU) is represented by a coding tree (CT) in the compression region. This is a quadtree partitioning of the CTU, and each leaf is called a coding unit (CU), as shown in FIG. 3.
[0024] Next, each CU is given some intra or inter prediction parameters (prediction information). For this purpose, it is spatially partitioned into one or more prediction units (PUs), and certain prediction information is assigned to each PU. The intra or inter coding mode is assigned at the CU level, as shown in FIG. 4, which shows an example of the partitioning of the coding tree unit into coding units, prediction units, and transform units.
[0025] In HEVC, exactly one motion vector is assigned to each PU. This motion vector is used for motion-compensated temporal prediction of the PU under consideration. Thus, in HEVC, the motion model that links the predicted block and its reference block includes translational motion.
[0026] To encode motion data, HEVC uses two modes. These are called AMVP (Adaptive Motion Vector Prediction) and Merge, respectively. AMVP involves signaling one or more reference pictures used to predict the current PU, a motion vector predictor index (obtained from a list of two predictors), and a motion vector difference. Generally, at least one embodiment described herein includes the Merge mode.
[0027] The Merge mode involves signaling and decoding an index of some motion data collected in a list of motion data predictors. This list consists of five candidates and is constructed in the same way on the decoder side and the encoder side. Thus, the Merge mode aims to derive some motion information obtained from the merge list. The merge list generally includes motion information related to several spatial and temporal surrounding blocks that are available in the decoded state when the current PU is being processed.
[0028] In the JEM (Joint Exploration Model) developed by the JVET (Joint Video Exploration Team) group, some additional temporal prediction tools with related parameters determined on the decoder side include SbTMVP (Sub-block Temporal Motion Vector Prediction), sometimes also called ATMVP (Alternative Temporal Motion Vector Prediction). The basic principle of SbTMVP is to derive the motion information of the current CU from certain motion information included in the reference pictures of the current picture. For this purpose, first, a so-called temporal vector is obtained as the motion vector of the first merge candidate of the current CU. This temporal vector with the related reference picture index can be used to search for the motion vector from the considered reference picture. For this purpose, the current CU is divided into N×N sub-CUs, and the motion data indicated by time from the center position of the considered sub-CU is regarded as the sub-CU MV SbTMVP predictor.
[0029] The SbTMVP motion prediction mode can be included as an additional candidate in the classical merge list. Another approach is to insert the SbTMVP motion prediction mode into the affine merge list, which was originally dedicated to constructing a set of merge candidates for predicting the affine motion model of the current CU when affine motion compensation is used for the current CU. The use of affine motion compensation for the current CU is signaled by the so-called affine flag at the CU level. Moving the SbTMVP motion predictor into the affine merge list may result in an improvement in compression efficiency compared to other approaches, for example, when the SbTMVP candidate is part of the classical merge candidate list.
[0030] Figure 5 shows an example of inter-mode coding in VTM-5, for example: - First, a skip flag indicating whether the CU contains residuals is decoded - Next, the regular_merge flag including the TMVP (temporal Motion Vector Predictor) candidates is decoded, indicating the use of normal merge. - Otherwise, the mmvd flag indicating the use of the Merge Motion Vector Difference mode is decoded. This mode may also include candidates constructed from TMVP candidates. - If not MMDV, the sub-block flag is decoded. The sub-block candidates can be either SbTMVP or affine candidates. - If not sub-block, in the case of (not skip) merge, the CIIP flag (Combined Intra-Inter Prediction) is decoded. This mode may also include candidates constructed from TMVP candidates. - Finally, if none of the above modes are selected, the triangle mode is inferred. This mode may also include candidates constructed from TMVP candidates.
[0031] In the AMVP mode (Advanced Motion Vector Prediction), the block can be either translational or affine. Note that the affine flag in AMVP is coded using the same CABAC context as the sub-block flag in skip or merge mode.
[0032] The merge mode in the HEVC standard includes deriving the inter prediction information (hereinafter also referred to as motion information) of a given prediction unit from the selected motion information predictor candidates. The motion information includes all inter prediction parameters of the PU, namely, - Unidirectional or bidirectional temporal prediction type - Reference picture index within each reference picture list - One or more motion vectors are included.
[0033] The coding and decoding of inter-prediction information in HEVC are summarized in FIG. 6 showing the signaling of inter-prediction information. As shown in the figure, the coding / decoding of motion information by the merge mode occurs in two modes: the skip mode and the merge mode. In these two modes, one field is signaled to enable the decoder to retrieve the motion information of the PU: the so-called merge index. The merge index indicates which motion vector predictor in the list of merge motion information predictors is used to derive the motion information of the current PU. Hereinafter, the list of motion information predictors is referred to as the merge list or the merge candidate list. Also, the candidate motion information predictors are referred to as merge candidates.
[0034] In HEVC, the merge candidate list is systematically created from five merge candidates. The merge list is constructed on both the encoder side and the decoder side as follows. FIG. 7 shows the positions of the spatial and temporal motion vector predictors used in the merge mode. The spatial merge candidates are shown on the left side of FIG. 7. The temporal merge candidates are shown on the right side of FIG. 7. As shown in FIG. 7, up to five spatial positions can be considered to search for some potential candidates. They are in the following order: 1 - Left (A1) 2 - Top (B1) 3 - Top - Right (B0) 4 - Bottom - Left (A0) 5 - Top - Left (B2) are visited, and the symbols A0, A1, B0, B1, B2 indicate the spatial positions shown on the left side of FIG. 7. Four different spatial candidates are selected. Next, by considering the temporal motion information located at position H, the TMVP indicating the temporal predictor is selected, and then the "center" is the candidate at position H in the considered reference picture where it is not available. Next, as shown in FIG. 8 showing the final pruning process performed to ensure that the set of selected spatial and temporal candidates does not contain redundant candidates, a list of merge motion vector predictor candidates is constructed.
[0035] Next, for the B slice, if the merge list is not full, another type of candidate, so-called combined candidate, is sent to the merge list. This may include generating a candidate consisting of motion information related to one of the reference picture lists (L0) from one candidate already present in the list, where this motion is related to the other reference picture list (L1) from another candidate already present in the merge list. If the merge list is not yet full (5 elements), zero motion vectors are sent to the end of the merge list until it is full. The overall process of merge list construction in HEVC is detailed in the chart of Figure 9.
[0036] In the SbTMVP (sub-block temporal motion vector prediction) method (also called ATMVP), one or several temporal motion vector predictors for the current CU are retrieved from one or more reference pictures for the current CU. First, the so-called temporal motion vector and related reference picture index are obtained as motion data related to the first candidate in the normal merge list candidates of the current CU. Next, the current CU is divided into sub-CUs of N×N (N is generally equal to 4). This is shown in "Error! Reference source not found." For each N×N sub-block, one or more motion vectors and one or more reference picture indexes are identified with the help of the temporal motion vector in the reference picture related to the temporal MV. The N×N sub-block in the reference picture indicated by the temporal MV from the current sub-CU position is considered. Its motion data is regarded as the SbTMVP motion data prediction of the current sub-CU. Next, it is converted into the motion vector and reference picture index of the current sub-CU by appropriate motion vector scaling.
[0037] Generally, one aspect of at least one example of an embodiment includes organizing the number and content of various lists of motion data predictors in a way that improves coding efficiency. Generally, another aspect of at least one embodiment includes reducing complexity to improve the worst-case scenario on the decoder side, for example, by reducing the number of merge candidates constructed for a given mode. Generally, another aspect of at least one embodiment includes making affine mode coding more consistent by not mixing sub-block flag coding and affine flag coding.
[0038] At least one example of an embodiment provides a temporal merge list that includes one or more temporal predictors, such as SbTMVP, TMVP, MMVD based on TMVP, CIIP based on TMVP, triangles based on TMVP, and these candidates are removed from the conventional merge list, thereby reducing the worst-case complexity in the decoder.
[0039] At least one example of an embodiment includes that temporal modes not included in the temporal merge list are maintained in their respective lists (e.g., MMVD, CIIP, triangles).
[0040] At least one example of an embodiment includes that temporal modes not included in the temporal merge list are completely removed so that each list has no temporal predictor.
[0041] At least one example of an embodiment includes extracting SbTMVP from sub-blocks and providing consistent coding of affine flags between merge / skip modes and AMVP modes.
[0042] At least one example of an embodiment includes separating sub-block and affine flag coding.
[0043] FIG. 11 shows an example of the construction of classical merge list construction (in contrast to sub-blocks). Here, classical means a merge list used for translational motion compensation time prediction where one motion vector is associated with one CU for each reference picture list. This process involves - Addition of spatial candidates A1, B1, B0 A0 B2 (with some pruning among them) (see FIG. 13 for the location of predictors) - Addition of temporal candidates (TMVP) (H or center) - Addition of HMVP (History based Motion Vector Predictor) candidates with at most some pruning for the first added candidates with respect to A1 and B1 - Generating paired (averaged) candidates if at least two candidates are in the list - Stuffing the null motion vector into the list if necessary and includes.
[0044] An example of the construction of a list of sub-block based motion predictor candidates is provided hereinafter. The affine merge list collects all merge MV predictors including sub-block based motion compensation time prediction. Thus, this includes the SbTMVP merge mode and the affine merge mode described above. Affine merge candidates represent merge candidates that are derived from a 4×4 block-based affine motion field with respect to the current CU and are used for time prediction of the current CU. Affine motion compensation is not described in detail here. One aspect to note here is that in at least one example, the SbTMVP candidates are placed at the head of the affine merge list.
[0045] Other examples of merge lists may include: - The MMVD merge list consists of the first two classical merge candidates displaced by a given motion vector difference signaled in the bitstream. Thus, this list may include candidates based on temporal candidates (displaced TMVP candidates). - The CIIP merge list is the same as the classical merge list. Thus, this list may include candidates based on temporal candidates. The last predictor is a mixture of a motion-based predictor and an intra prediction. - The triangle merge list consists of candidates generated from candidates of the classical merge list. Thus, this list may include candidates based on temporal candidates.
[0046] Generally, at least one aspect of at least one example of an embodiment described herein includes generating a merge list that includes temporal candidates.
[0047] FIG. 14 shows an example of an embodiment of a correction process for decoding a temporal merge list including at least the following: - After decoding the mmvd mode flag, if the mmvd flag is false, the temporal mode flag is decoded. - If the temporal mode is true, a temporal candidate index is decoded similar to the classical merge candidate index - Otherwise, the affine mode flag is decoded. Note that since the SbTMVP candidates have been removed from the list, the affine flag replaces the sub-block mode flag. Thanks to this split, the affine flag coding in the merge mode and the AMVP mode uses the same CABAC bin in a consistent way (as opposed to a technique where the sub-block flag meant both SbTMVP or affine mode).
[0048] Generally, at least one aspect of at least one example of an embodiment may include providing various arrangements of the temporal merge list. For example, the temporal merge list may be - As a first merge list - After the normal merge list - After the affine merge list - After the ciip merge list - After the triangle merge list (in this case, if the triangle mode is false, the temporal mode is deduced to be true) It is also possible to be inserted.
[0049] Generally, the list of time candidates can be generated in the same way as the merge list is generated, by adding potential candidates (if any) in a certain order until the maximum number of merge candidates (usually 2 or 3) is reached. An index indicating the position in the list of time candidates to be used is then sent to the decoder. An example of the list of time candidates is - SbTMVP candidates - TMVP candidates - Temporal CIIP candidates is.
[0050] The variant form may include one or more of the following: - By further signaling, TMVP candidates regarding MMVD and triangular mode may also be added to the list. For that purpose, further information (generally, the mmvd flag and mmvd displacement). - TMVP candidates regarding MMVD and / or triangle and / or CIIP are maintained in their original lists. - TMVP candidates regarding MMVD and / or triangle and / or CIIP are completely removed, thereby making all other lists have no time candidates. Optionally, in this variant form, in order to make the affine merge list have no temporal motion vectors, the restored affine candidates do not use the temporal motion vectors. - Pure temporal restoration affine candidates are also added to the time list.
[0051] Generally, one or more of the described embodiments can provide the following: - Reduction in the complexity of the decoder process where a given merge list (usually mmvd, temporal, affine, ciip, or triangular) has a reduction in worst-case complexity (e.g., the length of one or more of one or more lists is reduced). - Since all spatial candidates are in a separate list without the need to access the time memory buffer, the bandwidth required to construct each candidate of the list is reduced. Generally, in all lists except the time list, only the spatial information around the current block needs to be cached.
[0052] Table 1 attached to this disclosure provides an example of an embodiment with syntax regarding the coding of time merge list flags. The example shown in Table 1 includes providing syntax regarding the coding of time merge list flags after the MMVD flag. Note that the sub-block syntax is replaced by the affine syntax as seen in (shaded in gray) AMVP derivation.
[0053] In the case of only the time merge list of SbTMVP and TMVP, the variable MaxNumTemporalMergeCand is equal to at most 2 when both candidates are available. Candidate availability is derived as before, - SbTMVP is available when the SPS level flag of SbTMVP is true and the current block size is 8 or more in both width and height - TMVP is available when the temporal flag of the SPS level is true and the width + height exceeds 12.
[0054] Another aspect of at least one example of an embodiment may include affine and sub-block flag coding. In a particular system, the coding of merge_subblock_flag (in merge / skip mode) and inter_affine_flag (in AMVP) may include sharing the same CABAC coding. In the case of a separate temporal list in at least one embodiment described herein (i.e., SbTMVP is no longer in the sub-block list), the merge_subblock_flag actually becomes the merge_affine_flag so that the CABAC coding is more consistent. Another variant that does not use a separate temporal merge list is to use two separate CABAC codings (one for merge_subblock_flag and one for inter_affine_flag).
[0055] Generally, another example of an embodiment is shown in FIG. 16. In FIG. 16, a bitstream including encoded picture information also includes encoded control information such as flags. At 1610, a first flag included in the encoded bitstream is decoded using CABAC coding based on a first probability model. At 1620, a second flag included in the encoded bitstream is decoded using CABAC coding based on a second probability model. That is, two separate CABAC codings are used to decode the first and second flags. Next, at 1630, the encoded picture information included in the bitstream is decoded based on the coding mode indicated by the first flag or the second flag to generate decoded picture information. For example, the first flag may indicate, correspond to, or be associated with a coding mode such as a sub-block merge prediction mode, and the second flag may indicate, correspond to, or be associated with a coding mode such as an inter-affine prediction mode.
[0056] Generally, another example of an embodiment is shown in FIG. 17. In FIG. 17, at 1710, the value of a first flag is determined, and this first flag indicates a first prediction mode related to the coding of picture information included in the input. At 1720, the value of a second flag is determined, and this second flag indicates a second prediction mode related to the coding of picture information. At 1730, at least a part of the picture information, as well as the first and second flags, are coded to generate a coded bitstream. The first flag indicates, for example, a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model. The second flag indicates, for example, an inter-affine prediction mode and is coded using CABAC coding based on a second probability model. That is, two separate CABAC codings are used to code the first and second flags.
[0057] This specification describes various examples of embodiments, features, models, techniques, etc. Many such examples are described with specificity and are often described in a manner that may seem limiting in order to show at least individual characteristics. However, this is for the purpose of clarity of explanation and does not limit its application or scope. In fact, the various examples of embodiments, features, etc. described in this specification may be combined and interchanged in various forms to provide further examples of embodiments. Examples of embodiments according to the present disclosure include (but are not limited to) the following.
[0058] Generally, an example of an embodiment includes decoding a first flag included in a coded bitstream, where the first flag is coded using CABAC coding based on a first probability model, decoding a second flag included in the coded bitstream, where the second flag is coded using CABAC coding based on a second probability model, and decoding picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag, and the method may include the first flag indicating a sub-block merge prediction mode and the second flag indicating an inter-affine prediction mode.
[0059] Generally, another example of an embodiment includes determining a value of a first flag indicating a first prediction mode related to encoding of picture information, determining a value of a second flag indicating a second prediction mode related to encoding of picture information, and encoding at least a part of the picture information and the first and second flags to generate a coded bitstream, where the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model. The method may include the encoding.
[0060] Generally, another example of an embodiment is to decode a first flag included in an encoded bitstream, where the first flag is encoded using CABAC encoding based on a first probability model, to decode, and to decode a second flag included in the encoded bitstream, where the second flag is encoded using CABAC encoding based on a second probability model, and to decode the encoded picture information included in the encoded bitstream based on the encoding mode indicated by the first flag or the second flag, and includes one or more processors configured to perform the above, where the first flag indicates a sub-block merge prediction mode and the second flag indicates an inter-affine prediction mode, and the apparatus may include.
[0061] Generally, another example of an embodiment is to determine the value of a first flag indicating a first prediction mode related to the encoding of picture information, to determine the value of a second flag indicating a second prediction mode related to the encoding of picture information, and to encode at least a part of the picture information and the first and second flags to generate an encoded bitstream, where the first flag indicates a sub-block merge prediction mode and is encoded using CABAC encoding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is encoded using CABAC encoding based on a second probability model, and the apparatus may include one or more processors configured to perform the above.
[0062] Generally, another example of an embodiment is a bitstream formatted to include encoded picture information, where the encoded picture information is encoded by processing picture information based on any one or more of the example embodiments of the method according to the present disclosure, and the bitstream may be included.
[0063] In general, one or more other examples of embodiments may also provide a computer-readable storage medium, such as a non-volatile computer-readable storage medium, that stores instructions for encoding or decoding picture information, such as video data, according to the methods or apparatuses described herein.
[0064] In general, at least one example of an embodiment may include a computer program that, when executed by a computer, includes instructions that cause the computer to perform a method according to one or more examples of the embodiments described herein.
[0065] In general, at least one example of an embodiment may include a non-transitory computer-readable medium that stores the above instructions for causing a computer that executes executable program instructions to perform a method according to one or more examples of the embodiments described herein.
[0066] In general, at least one example of an embodiment may include a signal that includes data generated according to any one or more examples of the embodiments described herein.
[0067] In general, at least one example of an embodiment may include a bitstream that includes syntactic elements and encoded picture information generated according to any one or more of the examples of the embodiments described herein.
[0068] In general, at least one example of an embodiment may include a computer-readable storage medium that stores a bitstream generated according to the methods or apparatuses described herein.
[0069] In general, at least one example of an embodiment may include transmitting or receiving a bitstream or signal generated according to the methods or apparatuses described herein.
[0070] Generally, at least one example of an embodiment may include a device according to any one or more of the examples of the embodiments described herein, and (i) an antenna configured to receive a signal, the signal including data representing image information, (ii) a band limiter configured to limit the signal received in a frequency band including data representing image information, and (iii) at least one of a display configured to display an image from the image information.
[0071] Generally, at least one example of an embodiment may include a device as described herein, the device including one of a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a mobile phone, a tablet, or other electronic device.
[0072] Generally, another example of an embodiment may include determining a value of a first flag indicating a first prediction mode related to encoding of picture information, determining a value of a second flag indicating a second prediction mode related to encoding of picture information, and encoding at least a part of the picture information and the first and second flags to generate an encoded bitstream, wherein the first flag indicates a sub-block merge prediction mode and is encoded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is encoded using CABAC coding based on a second probability model, and may include a device including one or more processors configured to perform the encoding.
[0073] In general, examples of the embodiments described and contemplated herein can be implemented in many different forms. The above FIGS. 1 and 2, as well as FIG. 15 below, provide some embodiments, but other embodiments are contemplated, and the descriptions of FIGS. 1, 2, and 15 do not limit the scope of the embodiments. At least one embodiment generally provides examples related to video encoding and / or decoding, and at least one other embodiment generally relates to transmitting a generated or encoded bitstream or signal. These embodiments and other embodiments may be implemented as a method, an apparatus, a computer-readable storage medium storing instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium storing a bitstream or signal generated according to any of the described methods.
[0074] The terms HDR (High Dynamic Range) and SDR (Standard Dynamic Range) are used in this disclosure. These terms often convey to those skilled in the art specific values of the dynamic range. However, further embodiments are also contemplated where a reference to HDR is understood to mean "higher dynamic range" and a reference to SDR is understood to mean "lower dynamic range." Such further embodiments are not restricted by any particular values of the dynamic range that are often associated with the terms "high dynamic range" and "standard dynamic range."
[0075] Various methods are described herein, and each of these methods includes one or more steps or actions for implementing the described method. The order and / or use of specific steps and / or actions may be modified or combined, provided that a particular order of steps or actions is not required for the correct operation of the method.
[0076] Using the various methods and other aspects described herein, modules, such as module 170 and / or 175 of a video encoder, such as encoder 100 shown in FIG. 1, of a motion estimation module and / or a motion compensation module, and / or the motion compensation module 275 of the video decoder 200 shown in FIG. 2, can be modified. Also, this aspect is not limited to VVC or HEVC, and can be applied to other standards and recommendations, as well as extended versions of any such standards and recommendations (including VVC and HEVC), whether existing or developed in the future. Unless otherwise indicated or technically excluded, the aspects described herein can be used individually or in combination.
[0077] For example, various numerical values are used in this specification. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0078] FIG. 15 shows a block diagram of an example of a system in which various aspects and embodiments can be implemented. System 1000 can be embodied as a device that includes the various components described below and is configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 1000 can be embodied in a single integrated circuit, multiple ICs, and / or discrete components, either alone or in combination. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more of the aspects described herein.
[0079] System 1000 includes at least one processor 1010 configured to execute instructions loaded therein to implement various aspects described herein, for example. The processor 1010 may include an embedded memory, an input / output interface, and various other circuitry known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). System 1000 includes a storage device 1040, and the storage device 1040 may include non-volatile memory and / or volatile memory including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drive, and / or optical disk drive. The storage device 1040 may include, as non-limiting examples, an internal storage device, an additional storage device, and / or a network-accessible storage device.
[0080] System 1000 includes an encoder / decoder module 1030 configured to process data to provide encoded video or decoded video, for example. The encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents one or more modules that may be included within a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Further, the encoder / decoder module 1030 may be implemented as a separate element of System 1000 or may be incorporated within the processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0081] To perform the various aspects described herein, the program code loaded into processor 1010 or encoder / decoder 1030 can be stored within storage device 1040 and later loaded into memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 can store one or more of various items while performing the processes described herein. Such stored items can include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams or signals, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logics.
[0082] In some embodiments, the internal memory of processor 1010 and / or encoder / decoder module 1030 is used to store instructions and to provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either processor 1010 or encoder / decoder module 1030) is used for one or more of these functions. The external memory can be, for example, memory 1020 and / or storage device 1040 such as dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory such as RAM is used as working memory for video encoding and decoding operations such as MPEG-2, HEVC, or VVC (Versatile Video Coding).
[0083] Inputs to the elements of system 1000 can be provided by various input devices as shown in block 1130. Such input devices include, but are not limited to, (i) an RF portion that receives RF signals wirelessly transmitted, for example, by a broadcast device, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0084] In various embodiments, the input device of block 1130 has respective input processing elements known in the art. For example, the RF portion can be associated with elements for (i) selecting a desired frequency (also referred to as selecting a signal or bandlimiting a signal to a certain frequency band), (ii) downconverting the selected signal, (iii) bandlimiting again to a narrower frequency band to select a signal frequency band, which may be called a channel in certain embodiments, (iv) demodulating the downconverted and bandlimited signal, (v) performing error correction, and (vi) demultiplexing to select a desired stream of data packets. The RF portion of various embodiments includes one or more elements for performing these functions, such as a frequency selector, signal selector, bandlimiter, channel selector, filter, downconverter, demodulator, error corrector, and demultiplexer. The RF portion can include a tuner that performs some of these functions, including, for example, downconverting a received signal to a lower frequency (e.g., an intermediate frequency or a frequency close to baseband) or to baseband. In certain set-top box embodiments, the RF portion and its associated input processing elements perform frequency selection by receiving an RF signal transmitted over a wired (e.g., cable) medium and filtering, downconverting, and filtering again to a desired frequency band. Various embodiments can reorder the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF portion includes an antenna.
[0085] Furthermore, the USB and / or HDMI terminals may each include an interface processor for connecting the system 1000 to other electronic devices via a USB and / or HDMI connection. For example, it is to be understood that various aspects of input processing, such as Reed-Solomon error correction, may be implemented within another input processing IC or within the processor 1010. Similarly, aspects of USB or HDMI interface processing may be implemented within a separate interface IC or within the processor 1010. The demodulated, error-corrected, and de-multiplexed stream is provided to various processing elements, including the processor 1010 and the encoder / decoder 1030, which operate with memory and storage elements to process the data stream, for example, for presentation on an output device.
[0086] The various elements of the system 1000 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and may transmit data between them using a suitable connection configuration 1140, such as an internal bus known in the art, including, for example, an I2C bus, wiring, and a printed circuit board.
[0087] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include a transceiver (but is not limited to) configured to transmit and receive data on the communication channel 1060. The communication interface 1050 may include a modem or a network card (but is not limited to), and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.
[0088] In various embodiments, data is streamed to system 1000 using a Wi-Fi network such as IEEE 802.11. The Wi-Fi signals of these embodiments are received on communication channel 1060 and communication interface 1050 adapted for Wi-Fi communication. The communication channel 1060 of these embodiments is generally connected to an access point or router that provides access to an external network, including the Internet, that enables streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that delivers data over the HDMI connection of input block 1130. Still other embodiments provide streamed data to system 1000 using the RF connection of input block 1130.
[0089] System 1000 can provide output signals to various output devices including a display 1100, a speaker 1110, and other peripheral devices 1120. The other peripheral devices 1120 include, in various examples of embodiments, one or more of a stand-alone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functions based on the output of system 1000. In various embodiments, control signals are transmitted between system 1000 and the display 1100, the speaker 1110, or the other peripheral devices 1120 using signaling such as AV.Link, CEC, or other communication protocols that enable device - to - device control with or without user intervention. The output devices can be communicatively coupled to system 1000 via dedicated connections using respective interfaces 1070, 1080, and 1090. Alternatively, the output devices may be connected to system 1000 using a communication channel 1060 via a communication interface 1050. The display 1100 and the speaker 1110 may be integrated in a single unit with other components of system 1000, for example, within an electronic device such as a television. In various embodiments, the display interface 1070 includes a display driver such as, for example, a timing controller (T Con) chip.
[0090] Alternatively, the display 1100 and the speaker 1110 may be separate from one or more of the other components, for example, if the RF portion of the input 1130 is part of a separate set - top box. In various embodiments where the display 1100 and the speaker 1110 are external components, the output signal may be provided via a dedicated output connection including, for example, an HDMI port, a USB port, or a COMP output.
[0091] The embodiments may be implemented by computer software implemented by the processor 1010, or by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type suitable for the technical environment and may be implemented using any suitable data storage technology, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type suitable for the technical environment and may include, as non-limiting examples, one or more of a microprocessor, a general-purpose computer, a dedicated computer, and a processor based on a multi-core architecture.
[0092] Throughout the present disclosure, various embodiments include decoding. As used in this application, "decoding" may include, for example, all or part of the process performed on a received encoded sequence to generate a final output suitable for display. In various embodiments, such a process may include, for example, one or more of the processes generally performed by a decoder, such as entropy decoding, inverse quantization, inverse transformation, and differential decoding. In various embodiments, such a process may additionally or alternatively include, for example, extracting a picture from a non-overlapping (packed) picture, determining an upsampling filter to be used and then performing upsampling of the picture, and inverting the picture to return it to the intended orientation, among other processes performed by the decoders of the various embodiments described in this application.
[0093] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the expression "decoding process" is intended to specifically refer to a subset of operations or to generally refer to a broader decoding process will be apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.
[0094] Also, various embodiments include encoding. Similar to the above description regarding "decoding", "encoding" as used in this application may include all or part of the process performed on an input video sequence to generate, for example, an encoded bitstream or signal. In various embodiments, such a process may include one or more of the processes generally performed by an encoder such as, for example, partitioning, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such a process may additionally or alternatively include the processes performed by the encoders of the various embodiments described in this application.
[0095] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the expression "encoding process" is intended to specifically refer to a subset of operations or to generally refer to a broader encoding process will be apparent based on the context of the specific description and is considered to be fully understood by those skilled in the art.
[0096] Note that the syntactic elements used in this specification are descriptive terms. Therefore, they do not prevent the use of other syntactic element names.
[0097] If the figure is presented as a flow chart, it should be understood that a block diagram of the corresponding apparatus is also provided. Similarly, if the figure is presented as a block diagram, it should be understood that a flow chart of the corresponding method / process is also provided.
[0098] Various embodiments refer to rate distortion optimization. Specifically, during the encoding process, often given the constraint of computational complexity, the balance or trade-off between rate and distortion is usually considered. Rate distortion optimization is typically formulated as minimizing the rate distortion function, which is a weighted sum of rate and distortion. Different techniques exist for solving the rate distortion optimization problem. For example, these techniques can be based on tests that span a wide range of all encoding options, including all considered modes or encoding parameter values, involving a complete evaluation of the encoding cost and the associated distortion of the restored signal after encoding and decoding. To reduce the encoding complexity, more specifically, faster techniques can be used that employ the calculation of approximate distortion based on predicted or prediction residual signals rather than the restored ones. A mixture of these two techniques can also be used, such as using approximate distortion for only some of the possible encoding options and complete distortion for other encoding options. Other techniques evaluate only a subset of the possible encoding options. More generally, many techniques utilize any of various techniques for optimization, but the optimization is not necessarily a complete evaluation of both the encoding cost and the associated distortion.
[0099] The embodiments and aspects described herein can be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of embodiment (e.g., described only as a method), the embodiments of the described features can also be implemented in other forms (e.g., an apparatus or a program). An apparatus can be implemented, for example, in appropriate hardware, software, and firmware. A method can be implemented, for example, in a processor (which generally refers to a processing device including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device). The processor also includes, for example, a communication device such as a computer, a mobile phone, a personal digital assistant (PDA), and other devices that facilitate the communication of information between end users.
[0100] References to "one embodiment", or "an embodiment", or "one implementation", or "an implementation", and other variations thereof, mean that a particular feature, structure, characteristic, etc. described in connection with the embodiment is included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment", or "in an embodiment", or "in one implementation", or "in an implementation", and other variations thereof, which occur in various places throughout this specification, do not necessarily all refer to the same embodiment.
[0101] In addition, this specification may refer to "obtaining" various information. Obtaining information can include, for example, one or more of determining information, estimating information, calculating information, predicting information, or retrieving information from memory.
[0102] Furthermore, this specification may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0103] In addition, this specification may refer to "receiving" various information. Receiving is intended to be a broad term, similar to "accessing". Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from memory). Further, "receiving" generally includes, in some form, during operations such as storing information, processing information, transmitting information, moving information, copying information, deleting information, calculating information, determining information, predicting information, or estimating information.
[0104] For example, in the cases of "A / B", "A and / or B", and "at least one of A and B", the use of any of " / ", "and / or", and "at least one of ~" is understood to be intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such expressions are intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or the selection of all three options (A and B and C). As will be apparent to those skilled in the art in the relevant art and related fields, this can be extended to the same number of items as the enumeration.
[0105] Also, in this specification, the term "signal" refers, among other things, to indicating something to the corresponding decoder. For example, in certain embodiments, the encoder signals a particular one of a plurality of parameters for improvement purposes. Thus, in one embodiment, the same parameter is used on both the encoder side and the decoder side. Accordingly, for example, the encoder can send a particular parameter to the decoder (explicit signaling) so that the decoder can use the same particular parameter. Conversely, if the decoder already has the above particular parameter in addition to other parameters, signaling without transmission (implicit signaling) can be used simply to enable the decoder to recognize and select the above particular parameter. By avoiding the transmission of actual functions, bit savings are achieved in various embodiments. It should be understood that signaling can be accomplished in various ways. For example, one or more syntax elements, flags, etc. are used in various embodiments to signal information to the corresponding decoder. The above relates to the verb form of the term "signal", but the term "signal" can also be used as a noun in this specification.
[0106] As will be apparent to those skilled in the art, embodiments can generate various signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described embodiments. For example, the signal can be formatted as a bitstream of the described embodiment or to carry a signal. Such a signal can be formatted, for example, as an electromagnetic wave (using, for example, the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier wave with the encoded data stream. The information carried by the signal can be analog information or digital information, for example. The signal can be transmitted over various different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0107] Various embodiments have been described. The embodiments can include any of the following features or entities, alone or in any combination, across various different claim categories and types. · Providing at least one merge mode in an encoder and / or decoder for processing video, based on a merge list including temporal candidates. · Providing at least one merge mode in an encoder and / or decoder for processing video, based on arranging the number and content of various lists of motion data predictors in a way that improves coding efficiency. · Providing in an encoder and / or decoder for processing video, based on reducing complexity to improve the worst case on the decoder side, for example, by reducing the number of merge candidates to be constructed for a given mode. · Providing in an encoder and / or decoder for processing video, based on not mixing sub-block flag coding and affine flag coding, with improved consistency of affine mode coding. · In an encoder and / or decoder for processing video, a temporal merge list including one or more temporal predictors including one or more of SbTMVP, or TMVP, or MMVD based on TMVP, or CIIP based on TMVP, or triangles based on TMVP, wherein these candidates are removed from a conventional merge list, thereby reducing the worst-case complexity in the decoder, and providing based on the temporal merge list. · In an encoder and / or decoder for processing video, providing based on at least one temporal mode that is not included in the temporal merge list and is maintained in each list (e.g., MMVD, CIIP, triangles). · In an encoder and / or decoder for processing video, providing based on at least one temporal mode that is not included in the temporal merge list and is completely removed to make each list have no temporal predictor. · In an encoder and / or decoder for processing video, providing based on extracting SbTMVP from a sub-block and making the coding of the affine flag consistent between the merge / skip mode and the AMVP mode. · In an encoder and / or decoder for processing video, providing based on separating sub-block coding and affine flag coding. · In an encoder and / or decoder for processing video, ○ After decoding the mmvd mode flag, if the mmvd flag is false, the temporal mode flag is decoded; ○ If the temporal mode is true, the temporal candidate index is decoded; ○ If the temporal mode is not true, the affine mode flag is decoded, and based on the SbTMVP candidates removed from the merge list, the affine flag replaces the sub-block mode flag, and the affine flag coding in the merge mode and the AMVP mode uses the same CABAC bin in a consistent manner and is decoded. Provided based on a modified process for decoding a temporal merge list including · In an encoder and / or decoder for processing video, providing various arrangements of a temporal merge list, wherein the temporal merge list is ○ Insertable as a first merge list, or ○ After a normal merge list, or ○ After an affine merge list, or ○ After a CIIP merge list, or ○ After a triangle merge list (in which case, if the triangle mode is false, the temporal mode is deduced to be true), Provided based on the above-provided ability to be inserted. · In an encoder and / or decoder for processing video, a list of temporal candidates generated by adding potential candidates in a certain order until the maximum number of merge candidates is reached, wherein an index indicating the position in the list of temporal candidates to be used is then sent to the decoder, and the potential candidates are ○ SbTMVP candidates, or ○ TMVP candidates, or ○ Temporal CIIP candidates Provided based on the list of temporal candidates including one or more of the above. · In an encoder and / or decoder for processing video, ○ Signaling indicating that TMVP candidates or triangle mode candidates for MMVD can be added to the merge list, or ○ TMVP candidates for MMVD and / or triangle and / or CIIP are maintained in their respective original lists, or ○ TMVP candidates for MMVD and / or triangle and / or CIIP are completely removed, thereby causing all other lists to have no temporal candidates, or ○ TMVP candidates related to MMVD and / or triangle and / or CIIP are completely removed, so that all other lists have no temporal candidates. In this case, in order for the affine merge list to have no temporal motion vectors, the restored affine candidates do not use temporal motion vectors, or ○ Adding pure temporal restoration affine candidates to the temporal list, Providing based on the generated merge list including at least one of the above. · In an encoder and / or decoder for processing video, reducing the complexity of the decoder process, and providing based on reducing the worst-case complexity where a given merge list has a reduction in worst-case complexity. · In an encoder and / or decoder for processing video, reducing the complexity of the decoder process, and providing based on reducing the worst-case complexity where a given merge list has a reduction in worst-case complexity including reducing the length of one or more merge lists. · In an encoder and / or decoder for processing video, reducing the complexity of the decoder process, and providing based on reducing the worst-case complexity where a given merge list has a reduction in worst-case complexity including reducing the length of one or more merge lists, and a given merge list typically includes at least one of mmvd, temporal, affine, ciip, or triangle. · In an encoder and / or decoder for processing video, reducing the bandwidth associated with constructing each candidate in the merge list, and providing based on reducing where all spatial candidates are in another list and there is no need to access the temporal memory buffer. · Inserting syntax elements signaling into the signaling that enable the encoder to provide at least one merge mode based on the merge candidate list described herein. · Selecting the merge mode to be applied by the decoder based on these syntax elements. · A bitstream or signal that includes one or more of the described syntax elements or variations thereof. · Inserting into signaling a syntax element that enables a decoder to provide a merge mode in a manner corresponding to the manner used by an encoder. · Generating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof. · A TV, set-top box, mobile phone, tablet, or other electronic device that encodes and / or decodes video according to any of the described embodiments and displays the resulting image (e.g., using a monitor, screen, or other type of display). · A TV, set-top box, mobile phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal that includes an encoded image and encodes and / or decodes video according to any of the described embodiments. · A TV, set-top box, mobile phone, tablet, or other electronic device that wirelessly receives a signal that includes an encoded image (e.g., using an antenna) and encodes and / or decodes video according to any of the described embodiments. · A computer program that stores program code which, when executed by a computer, performs video encoding and / or decoding according to any of the described embodiments. · A non-transitory computer-readable medium that includes the above instructions for causing a computer that executes executable program instructions to perform video encoding and / or decoding according to any of the described embodiments.
[0108] A variety of other generalized and specialized embodiments are also supported and contemplated throughout the present disclosure.
[0109]
Table 1
Claims
1. decoding a first flag included in a coded bitstream, the first flag being coded using CABAC coding based on a first probability model; decoding a second flag included in the coded bitstream, the second flag being coded using CABAC coding based on a second probability model; decoding coded picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag; Including, the first flag indicates a sub-block merge prediction mode; and the second flag indicates an inter-affine prediction mode; method.
2. determining a value of a first flag indicating a first prediction mode associated with encoding of the picture information; determining a value of a second flag indicating a second prediction mode associated with encoding of the picture information; encoding at least a portion of the picture information and the first flag and the second flag to generate an encoded bitstream, where the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model; The method includes:
3. decoding a first flag included in a coded bitstream, the first flag being coded using CABAC coding based on a first probability model; decoding a second flag included in the coded bitstream, the second flag being coded using CABAC coding based on a second probability model; decoding encoded picture information included in the coded bitstream based on a coding mode indicated by the first flag or the second flag; [0023] In one embodiment, the method comprises: the first flag indicates a sub-block merge prediction mode; and the second flag indicates an inter-affine prediction mode; Device.
4. determining a value of a first flag indicating a first prediction mode associated with encoding of the picture information; determining a value of a second flag indicating a second prediction mode associated with encoding of the picture information; encoding at least a portion of the picture information and the first flag and the second flag to generate an encoded bitstream, where the first flag indicates a sub-block merge prediction mode and is coded using CABAC coding based on a first probability model, and the second flag indicates an inter-affine prediction mode and is coded using CABAC coding based on a second probability model; 16. An apparatus comprising: one or more processors configured to:
5. A computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to claim 1 or 2.
6. A non-transitory computer readable medium having stored thereon executable program instructions that cause a computer executing said instructions to perform the method of claim 1 or 2.
7. A signal comprising data generated according to the method of claim 2.
8. A bitstream formatted to include syntax elements and image information encoded according to the method of claim 2.
9. An apparatus according to claim 3 or 4, at least one of: (i) an antenna configured to receive a signal, the signal including data representative of the image information; (ii) a band limiter configured to limit the received signal to a frequency band including the data representative of the image information; and (iii) a display configured to display an image from the image information; The device that contains
10. The device of claim 10 , wherein the device comprises one of a television, a television signal receiver, a set-top box, a gateway device, a mobile device, a mobile phone, a tablet, or other electronic device.