Affine mode signaling in video encoding and decoding
Independent CABAC contexts and probability models for affine mode signaling in video encoding and decoding improve compression efficiency by addressing statistical inefficiencies in existing systems.
Patent Information
- Application Number
- JP2025183702
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-10-10
- Filing Date
- 2025-10-30
- Publication Date
- 2026-01-27
Smart Images

Figure 2026012915000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD This disclosure relates to video encoding and decoding. [Background technology]
[0002] To achieve high compression efficiency, image and video coding schemes, such as those defined by the High Efficiency Video Coding (HEVC) standard, typically utilize prediction and transform coding to exploit spatial and temporal redundancy within video content. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original and predicted block, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to prediction, transform, quantization, and entropy coding. Recent additions to video compression technology include various versions of the Joint Exploration Model (JEM) reference software and / or documentation being developed by the Joint Video Exploration Team (JVET). The goal of efforts such as JEM is to provide further improvements over existing standards, such as HEVC. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] H.265 / HEVC High Efficiency Video Coding, ITU-T H.265 Telecommunication Standardization Sector of ITU, “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services- Coding of moving video,High efficiency video coding” Summary of the Invention [Problem to be solved by the invention]
[0004] A novel affine mode signaling in video encoding and decoding is desired. [Means for solving the problem]
[0005] Generally, at least one example embodiment includes a method for encoding video data, the method comprising: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and encoding the video data based on the first CABAC context and the first CABAC probability model during the affine mode, and based on the second CABAC context and the second CABAC probability model during the second mode.
[0006] Generally, at least one example embodiment includes a method for decoding video data, the method comprising: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and decoding the video data based on the first CABAC context and the first CABAC probability model to decode the encoded video data during the affine mode, and based on the second CABAC context and the second CABAC probability model to decode the encoded video data during the second mode.
[0007] Generally, at least one example embodiment includes an apparatus for encoding video data having at least one or more processors, the one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of an affine mode; determine a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and encode the video data based on the first CABAC context and the first CABAC probability model during the affine mode, and based on the second CABAC context and the second CABAC probability model during the second mode.
[0008] Generally, at least one example embodiment includes an apparatus for decoding video data, comprising one or more processors, the one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of an affine mode; determine a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and decode the video data based on the first CABAC context and the first CABAC probability model during the affine mode to decode the encoded video data, and based on the second CABAC context and the second CABAC probability model during the second mode to decode the encoded video data.
[0009] Generally, at least one example of an embodiment includes a device, such as, but not limited to, a television, set-top box, cell phone, tablet, or other electronic device, that performs any of the described embodiments and / or displays the resulting image (e.g., using a monitor, screen, or other type of display), and / or tunes a channel (e.g., using a tuner) to receive a signal containing the encoded image and / or receives the signal containing the encoded image wirelessly (e.g., using an antenna), and performs any of the described embodiments.
[0010] In general, at least one example embodiment includes a bitstream formatted to include encoded video data, the encoded video data being encoded by at least one method described herein.
[0011] In general, at least one example embodiment provides a computer-readable storage medium having stored thereon instructions for encoding or decoding video data in accordance with the methods or apparatus described herein.
[0012] In general, at least one example embodiment provides a computer-readable storage medium having stored thereon a bitstream generated according to the methods or apparatus described herein.
[0013] Generally, various example embodiments provide methods and / or apparatus for transmitting or receiving bitstreams generated in accordance with the methods or apparatus described herein.
[0014] The foregoing presents a simplified summary of the present subject matter in order to provide a basic understanding of some aspects of the disclosure. This summary is not an extensive overview of the present subject matter. It is not intended to identify key / critical elements of embodiments or to delineate the scope of the present subject matter. Its sole purpose is to present some concepts of the present subject matter in a simplified form as a prelude to the more detailed description that is provided below.
[0015] The present disclosure may be better understood by consideration of the following detailed description taken in conjunction with the accompanying drawings. [Effects of the Invention]
[0016] A novel affine mode signaling in video encoding and decoding is provided. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a diagram illustrating video information partitioning for video encoding and decoding including coding tree units (CTUs) in HEVC. [Figure 2]FIG. 1 illustrates a partitioning of video information for video encoding and decoding, including CTUs and coding units (CUs). [Figure 3] FIG. 1 illustrates aspects of an affine motion model. [Figure 4] FIG. 1 illustrates an embodiment including an affine motion model. [Figure 5] FIG. 1 illustrates control point motion vector prediction (CPMVP) associated with affine inter mode. [Figure 6] FIG. 10 illustrates candidate positions associated with an affine merge mode. [Figure 7] FIG. 10 illustrates the motion vectors involved in determining the control point motion vector (CPMV) during affine merge mode. [Figure 8] FIG. 10 illustrates an embodiment of determining a context associated with a flag, such as a CABAC context associated with an affine flag. [Figure 9] FIG. 1 illustrates an example of determining a probability model for encoding and / or decoding, such as a CABAC probability model. [Figure 10] FIG. 10 illustrates another example of determining a probability model for encoding and / or decoding, such as a CABAC probability model. [Figure 11] FIG. 10 illustrates an embodiment of determining a context associated with a flag, such as a CABAC context associated with an affine flag. [Figure 12] FIG. 10 illustrates an embodiment of determining a context associated with a flag, such as a CABAC context associated with an affine flag. [Figure 13] 1A-1C illustrate various example embodiments of determining a context associated with a flag, such as a CABAC context associated with an affine flag. [Figure 14] 1 illustrates an embodiment of an encoder suitable for encoding video data according to one or more of the embodiments herein. [Figure 15]FIG. 2 illustrates an embodiment of a decoder suitable for decoding video data according to one or more of the embodiments herein. [Figure 16] 1 illustrates an embodiment of a system suitable for encoding and / or decoding video data according to one or more of the embodiments herein. DETAILED DESCRIPTION OF THE INVENTION
[0018] It should be understood that the drawings are for purposes of illustrating examples of various aspects, embodiments and features, and are not necessarily the only possible configuration. Like reference designators throughout the various views refer to the same or similar features.
[0019] To achieve high compression efficiency, image and video coding schemes typically utilize prediction and transformation, including motion vector prediction, to exploit spatial and temporal redundancy within video content. Typically, intra- or inter-prediction is used to exploit intra- or inter-frame correlation, and then the difference between the original and predicted image, often referred to as the prediction error or prediction residual, is transformed, quantized, and entropy coded. To reconstruct the video, the compressed data is decoded by an inverse process corresponding to entropy coding, quantization, transformation, and prediction. Entropy coding / decoding typically involves context-adaptive binary arithmetic coding (CABAC).
[0020] A recent addition to high compression techniques involves the use of motion models based on affine modeling. Affine modeling is used for motion compensation for encoding and decoding video pictures. In general, affine modeling is a model that uses at least two parameters, such as two control point motion vectors (CPMVs), to represent motion at each corner of a block of a picture, and allows for deriving a motion field for the entire block of a picture, making it possible to simulate, for example, rotations and similarities (zooms).
[0021] The general aspects described herein are in the field of video compression and are aimed at improving compression efficiency compared to existing video compression systems.
[0022] In the HEVC video compression standard (see Non-Patent Document 1), motion compensated temporal prediction is used to exploit the redundancy that exists between successive pictures of a video.
[0023] To do so, a motion vector is associated with each prediction unit (PU), which we will now introduce. Each CTU (coding tree unit) is represented in the compressed domain by a coding tree, which is a quadtree division of the CTU, where each leaf is called a coding unit (CU), see Figure 1.
[0024] Each CU is then given some intra- or inter-prediction parameters (prediction information). To do so, it is spatially partitioned into one or more prediction units (PUs), and each PU is assigned some prediction information. The intra- or inter-coding mode is assigned at the CU level, see Figure 2 for this.
[0025] In HEVC, exactly one motion vector is assigned to each PU. This motion vector is used for motion-compensated temporal prediction of the considered PU. Therefore, in HEVC, the motion model linking a predicted block with its reference block simply involves a transformation.
[0026] In the Joint Exploration Model (JEM), developed by the JVET (Joint Video Exploration Team) group, and the subsequent Versatile Video Coding (VVC) Test Model (VTM), several richer motion models are supported to improve temporal prediction. To do so, a PU can be spatially divided into sub-PUs, and a richer model can be used to assign a dedicated motion vector to each sub-PU.
[0027] CUs are no longer divided into PUs or TUs, and some motion data is directly assigned to each CU. In this new codec design, CUs can be divided into sub-CUs, and motion vectors can be calculated for each sub-CU.
[0028] One of the new motion models introduced in JEM is the affine model, which basically involves using an affine model to represent the motion vectors in a CU.
[0029] The motion model used is illustrated by Figure 3. The affine motion field contains, for each position (x,y) inside the considered block, the following motion vector component values:
[0030]
number
[0031] Equation 1: The affine model used to generate the motion field inside the CU to predict
[0032] Coordinate(v 0x, v 0y ) and (v 1x ,v1y ) are the so-called control point motion vectors used to generate the affine motion field. 0x ,v 0y ) is the upper left corner control point of the motion vector, and (v 1x ,v 1y ) is the motion vector upper right corner control point.
[0033] In practice, to keep complexity reasonable, we calculate motion vectors for each 4x4 sub-block (sub-CU) of the considered CU, as illustrated in Figure 4. Affine motion vectors are calculated from control point motion vectors at the center position of each sub-block. The obtained MVs are expressed with 1 / 16 pixel accuracy.
[0034] As a result, the temporal coding of a coding unit in affine mode involves motion compensated prediction of each sub-block using its own motion vector.
[0035] Note that a model using three control points is also possible.
[0036] In VTM, affine motion compensation can be used in two ways: affine inter (AF_INTER), affine merge, and affine template, which are introduced below.
[0037] - Affine Inter (AF_INTER): CUs in AMVP mode with a size larger than 8x8 can be predicted in affine inter mode. This is signaled through a flag in the bitstream. The generation of an affine motion field for that inter CU involves determining a control point motion vector (CPMV), which is obtained by the decoder through the addition of a motion vector difference and a control point motion vector prediction (CPMVP). A CPMVP is a pair of motion vector candidates obtained from lists (A, B, C) and (D, E), respectively, as illustrated in Figure 5.
[0038]
number
[0039] First, the CPMVP is checked for validity using Equation 2 for a block of height H and width W.
[0040]
number
[0041] Equation 2: Validity test for each CPMVP
[0042] Then the valid CPMVP is the third motion vector (obtained from position F or G)
[0043]
number
[0044] to the vectors given by the affine motion model for , which is better: CPMVP.
[0045] For a block of height H and width W, the cost of each CPMVP is calculated using Equation 3. In the following equation, X and Y are the horizontal and vertical components of the motion vector, respectively.
[0046]
number
[0047] Equation 3: Cost calculated for each CPMVP
[0048] - Affine merge: In affine merge mode, a CU-level flag indicates whether the merged CU uses affine motion compensation. If so, the first available neighboring CU coded in affine mode is selected from the ordered set of candidate positions (A, B, C, D, E) in Figure 6.
[0049] Once the first neighboring CU in affine mode is obtained, the top left corner of the neighboring CU,
[0050]
number
[0051] (See FIG. 7.) Based on these three vectors, two CPMVs for the top-left and top-right corners of the current CU are derived as follows:
[0052]
number
[0053] Equation 4: Deriving the CPMV of the current CU based on the three corner motion vectors of neighboring CUs
[0054]
number
[0055] Then, the motion field inside the current CU is calculated based on the 4x4 sub-CUs.
[0056] For the affine merge mode, more candidates can be added, and the best candidate is selected from up to seven candidates, and the index of the best candidate is coded into the bitstream.
[0057] Another type of candidate is called temporal affine. Similar to TMVP (Temporal Motion Vector Predictor) candidates, affine CUs are searched for in the reference image and added to the candidate list.
[0058] The process can create "virtual" affine candidates that are added. Such a process can be useful for creating affine candidates when no affine CUs are available around the current CU. To do so, an affine model is created by taking the motion of individual sub-blocks at the corners and creating an "affine" model.
[0059] There are two constraints that must be considered during shortlisting for complexity reasons.
[0060] - Total number of potential candidates: Increases the total amount of computation required.
[0061] Final list size: Increases the delay in the decoder by increasing the number of comparisons required for each successive candidate.
[0062] Generally, at least one embodiment involves video coding using affine flag coding for motion compensation. Affine flags are known in the context of video coding systems such as VVC. The affine flag signals, at a coding unit (CU) level, whether affine motion compensation is used for temporal prediction of the current CU. In the following description, a coding unit is referred to as a coding unit, CU, or block.
[0063] In inter mode, the affine flag signals the use of an affine motion field to predict a block, as described above. Affine motion models typically use four or six parameters, as described above. This mode is used in both AMVP, where mvd (motion vector difference with motion predictor) is coded, and merge, where mvd is estimated to be zero. In general, at least one example embodiment described herein can also be applied to other modes, such as mmvd (merged motion vector difference, also known as UMVE), where mvd can be signaled in merge mode, or DMVR (decoder-side motion refinement), where the motion predictor is refined on the decoder side. In general, at least one embodiment provides an improved coding of the affine flag.
[0064] Currently, the affine flag is conveyed using a CABAC context. To encode using CABAC, non-binary syntax element values are mapped to a binary sequence, called a bin string, through a binarization process. A context model is selected for each bin. A "context model" is a probability model for one or more bins, selected from available model choices depending on the statistics of recently coded symbols. The context model for each bin is identified by a context model index (also used as a "context index"), with different context indices corresponding to different context models. The context model stores the probability that each bin is "1" or "0" and can be adaptive or static. A static model triggers the coding engine when a bin has equal probabilities for "0" and "1." In an adaptive coding engine, the context model is updated based on the actual coded value of the bin. The operating modes corresponding to the adaptive and static models are called normal and bypass modes, respectively. Based on the context, the binary arithmetic coding engine encodes or decodes the bin according to the corresponding probability model.
[0065] As mentioned above, currently, the affine flag is propagated using a CABAC context, which is a function of the values of the affine flags associated with neighboring blocks. An example is illustrated in FIG. 8, which shows an example of affine flag context derivation for both AMVP and merged affine. In FIG. 8, in 3010, for a current CU located at x, y and having dimensions width, height, an affine context Ctx is initialized, e.g., to 0. In 3020, the left-neighboring CU is checked to see if it is affine. If 3020 determines that the left-neighboring CU is affine, Ctx is incremented in 3030. If 3020 determines that the left-neighboring CU is not affine, in 3040, the upper-neighboring CU is checked to see if it is affine. If the check at 3040 determines that the upper neighbor CU is affine, Ctx is incremented at 3050, and then the process ends at 3060 with the Ctx value obtained by the described operations. If the check at 3040 determines that the upper neighbor CU is not affine, the process ends at 3060 with the Ctx value obtained after 3040. The affine flag context derived by the exemplary embodiment illustrated in Figure 8 can have three different values: Ctx = 0 (neither the left nor the above neighbor is affine), or Ctx = 1 (only one of the left or above neighbor is affine), or Ctx = 2 (both the left and above neighbors are affine).
[0066] The same CABAC context is shared by both AMVP and merge modes. As a result, the same CABAC bins or probability models are indicated or selected for both modes. Therefore, encoding or decoding uses the same model for both modes. One problem is that systems such as VVC also support other affine candidates (called virtual candidates or constructed candidates) that are constructed without using the affine flag values of neighboring CUs. These other affine candidates use affine motion models from individual motion vectors. Using one single context to encode affine flags for merge, AMVP, and other affine candidate modes is inefficient in capturing the various statistical behaviors of affine flags.
[0067] Generally, at least one embodiment described herein takes into account the different statistical occurrences of affine mode usage between two inter-prediction modes, e.g., AMVP and merge, by independently encoding affine flags in these two modes. An example is illustrated in FIG. 9. In FIG. 9, at 910, an affine context Ctx is determined or obtained for a current CU. This determination of Ctx may be performed according to an example embodiment, such as that shown in FIG. 8. Following the determination of Ctx at 910, at 920, a CABAC bin or probability model for a first inter-prediction mode, e.g., affine mode, is determined or obtained, for example, based on a particular formula or relationship, or a table or list of associations or correspondences between Ctx values and CABAC bins or models. At 930, CABAC bins or probability models for a second inter-prediction mode, e.g., merge mode, are determined or obtained, e.g., based on a specific formula or relationship, or a table or list of associations or correspondences between Ctx values and CABAC bins or models. Thus, CABAC bins or models are determined independently for each mode, i.e., at 920 for one mode and at 930 for another mode.
[0068] In the example of Figure 9, the total number of CABAC contexts for the affine flag is doubled, i.e., three for each of the first inter-prediction mode A and the second inter-prediction mode B. That is, an example embodiment for determining affine flag contexts as in Figure 9 can use six contexts instead of three. In other words, the CABAC context values can be derived in the same way for each of the two inter-prediction modes, but the probability models associated with the CABAC contexts can be determined independently and can be different for each of two different inter-prediction modes that include affine flags, e.g., AMVP and merge.
[0069] An example of independently encoding affine flags for two inter-prediction modes, i.e., affine and merge, is illustrated by the following code: The following example illustrates the determination of a CABAC context associated with each of the two modes, affine and merge, indicated by the respective mode flags AffineFlag and SubblockMergeFlag. As described above, the derivation of the CABAC context can be the same and can be based on an exemplary embodiment, such as that of FIG. 8. Mode 1 - Affine (AffineFlag): ... unsigned ctxId=DeriveCtx::CtxAffineFlag(cu); / / Derive context based on the upper and left neighbors. Context can be 0, 1, or 2. cu.affine=m_BinDecoder.decodeBin(Ctx::AffineFlag(ctxId)); / / Use the AffineFlag associated with the probabilistic model. ... Mode 2 - Merge (SubblockMergeFlag) ... unsignedctxId=DeriveCtx::CtxAffineFlag(cu); / / Note that the context derivation is the same as in affine. cu.subblock=m_BinDecoder.decodeBin(Ctx::SubblockMergeFlag(ctxId)); / / Note that the model (bin) derivation or determination is different from the affine one. ... Another example is shown in FIG. 10. In FIG. 10, 1011 and 1012 determine first and second CABAC contexts associated with respective first and second flags, e.g., an affine flag and a sub-block merge flag. The affine flag indicates that the affine mode is used. The sub-block merge flag indicates that either the affine mode or a second mode different from the affine mode is used. The context is associated with the mode to be used. For example, the second mode can be a merge mode such as SbTMVP. In 1013 and 1014, first and second CABAC probability models associated with the first and second contexts, and thus the modes, are independently determined.
[0070] Other example embodiments for improving context modeling include at least the following. One example includes completely removing CABAC context modeling based on spatial affine neighborhoods. In this case, only one context is used. Another example includes reducing modeling complexity by using only whether spatial neighborhoods are available. For this example, two contexts are possible: no affine neighborhoods are available (context 0), or at least one is available (context 1). An example in which two contexts are available is illustrated in FIG. 11. In FIG. 11, at 1111, an affine context Ctx is initialized, e.g., set equal to 0. At 1112, if either the left-neighboring CU or the top-neighboring CU is affine, Ctx is incremented at 1113, after which the Ctx value from 1113, e.g., 1, is returned, and the process ends at 1114. If neither is affine, then after 1112, the returned Ctx value is either 0 or 1, and the process ends at 1114.
[0071] In general, at least one embodiment may include the possibility of constructing virtual affine candidates as an aspect of context modeling, as shown in FIG. 12, which illustrates an example of context modeling for affine that takes virtual affine candidates into account. In FIG. 12, a context Ctx is initialized, e.g., set equal to zero, at 1211. Virtual affine candidates may be constructed from neighboring CUs that are coded in inter mode but not in affine mode. At 1212, the neighboring CUs are checked to determine whether they are inter-coded. If not, after 1212, the process ends at 1216 with the value of Ctx at an initialized value, e.g., 0. If the check at 1212 determines that the neighboring CU is inter-coded, then 1212 is followed by 1213, where Ctx is incremented (initial value + 1). After 1213, a check is made at 1214 to determine whether the neighboring CU is affine. If it is affine, then in 1215 Ctx is incremented ((initial value + 1) + 1) and then the Ctx value is returned, and the process ends in 1216; if it is not affine, then after the check in 1214, the process ends in 1216. To summarize the results of the arrangement in FIG. ○ 0 if no inter-neighborhood is available, ○ If inter-neighborhood is available but affine neighborhood is not available, then 1, ○2 if affine neighborhood is available is.
[0072] In a variant, at least one embodiment includes considering the presence of reference pictures that are inter-pictures, as it allows for the creation of virtual temporal candidates, as illustrated by the example in Figure 13. That is, Figure 13 illustrates an example of context modeling for affine, taking into account virtual temporal affine candidates. Virtual temporal candidates can be constructed from temporally co-located CUs coded in inter mode. This example in Figure 13 ○ An inter neighbor is available or an inter reference picture with reference picture index 0 is available Thus, this involves modifying the above-described exemplary embodiment of context modeling in Figure 12 by changing 1212 in Figure 12 to 1312 in Figure 13. Other features of the embodiment of Figure 13 correspond to similar features of Figure 12 described above and will not be described again with respect to Figure 13.
[0073] This document describes various example embodiments, features, models, techniques, etc. Many such examples are described with specificity, often in a manner that may appear limiting, at least to illustrate individual features. However, this is for clarity of description and does not limit applicability or scope. Indeed, the various example embodiments, features, etc. described herein can be combined and interchanged in various ways to provide further example embodiments.
[0074] Examples of embodiments according to the present disclosure include, but are not limited to, the following.
[0075] In general, at least one example embodiment may include a method for encoding video data, the method including the steps of: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and encoding the video data based on the first CABAC context and the first CABAC probability model during the affine mode, and based on the second CABAC context and the second CABAC probability model during the second mode.
[0076] In general, at least one example embodiment may include an apparatus for encoding video data comprising one or more processors, the one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of an affine mode; determine a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and encode the video data based on the first CABAC context and the first CABAC probability model during the affine mode, and based on the second CABAC context and the second CABAC probability model during the second mode.
[0077] In general, at least one example embodiment may include a method for decoding video data, the method including the steps of: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; and decoding the encoded video data based on the first CABAC context and the first CABAC probability model during the affine mode, and decoding the encoded video data based on the second CABAC context and the second CABAC probability model during the second mode.
[0078] In general, at least one example embodiment may include an apparatus for decoding video data, comprising one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of an affine mode; determine a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, the first CABAC context corresponding to a first CABAC probability model and the second CABAC context corresponding to a second CABAC probability model different from the first CABAC probability model; decode the encoded video data based on the first CABAC context and the first CABAC probability model during the affine mode; and decode the encoded video data based on the second CABAC context and the second CABAC probability model during the second mode.
[0079] In general, at least one example embodiment may include a method for encoding video data, the method including: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode; and encoding the video data to generate encoded video data, wherein the video data generated based on the affine mode is encoded based on a first CABAC probability model associated with the first CABAC context, and the video data generated based on the second mode is encoded based on a second CABAC probability model associated with the second CABAC context, the second CABAC probability model being different from the first CABAC probability model.
[0080] Generally, at least one example embodiment may include an apparatus for encoding video data, comprising one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of an affine mode; determine a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode; and encode the video data to generate encoded video data, wherein the video data generated based on the affine mode is encoded based on a first CABAC probability model associated with the first CABAC context, and the video data generated based on the second mode is encoded based on a second CABAC probability model associated with the second CABAC context, different from the first CABAC probability model.
[0081] In general, at least one example embodiment may include a method for decoding video data, the method including: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode; and decoding the video data to generate decoded video data, wherein the video data encoded based on the affine mode is decoded based on a first CABAC probability model associated with the first CABAC context and the video data encoded based on the second mode is decoded based on a second CABAC probability model associated with the second CABAC context, the second CABAC probability model being different from the first CABAC probability model.
[0082] In general, at least one example embodiment may include an apparatus for decoding video data, comprising one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of an affine mode; determine a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode; and decode the video data to generate decoded video data, wherein the video data encoded based on the affine mode is decoded based on a first CABAC probability model associated with the first CABAC context, and the video data encoded based on the second mode is decoded based on a second CABAC probability model associated with the second CABAC context, the second CABAC probability model being different from the first CABAC probability model.
[0083] In general, at least one example embodiment may include a method for encoding video data, the method including: determining a first CABAC context associated with a first flag indicating use of an affine mode; determining a second CABAC context associated with a second flag indicating use of either the affine mode or a second mode different from the affine mode, wherein determining the second CABAC context occurs independently from determining the first CABAC context; and encoding the video data to generate encoded video data based on the first and second flags and the first and second CABAC contexts during the first and second modes, respectively.
[0084] In general, at least one example embodiment may include an apparatus for encoding video data comprising one or more processors, the one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of affine mode; determine a second CABAC context associated with a second flag indicating use of either affine mode or a second mode different from the affine mode, where determining the second CABAC context occurs independently from determining the first CABAC context; and, during the first and second modes, encode the video data to generate encoded video data based on the first and second flags and the first and second CABAC contexts, respectively.
[0085] In general, at least one example embodiment may include a method for decoding video data, the method including: determining a first CABAC context associated with a first flag indicating use of affine mode; determining a second CABAC context associated with a second flag indicating use of either affine mode or a second mode different from the affine mode, where determining the second CABAC context occurs independently from determining the first CABAC context; and decoding the encoded video data based on the first and second CABAC contexts during the first and second modes, respectively, to generate decoded video data.
[0086] Generally, at least one example embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors configured to: determine a first CABAC context associated with a first flag indicating use of affine mode; determine a second CABAC context associated with a second flag indicating use of either affine mode or a second mode different from the affine mode, where determining the second CABAC context occurs independently from determining the first CABAC context; and decode the encoded video data based on the first and second CABAC contexts during the first and second modes, respectively, to generate decoded video data.
[0087] In general, at least one example embodiment may include a method for encoding video data, the method including: determining a CABAC context associated with a sub-block merging mode flag indicating use of a mode including an affine mode or a second mode different from the affine mode, the CABAC context corresponding to a first CABAC probability model during the affine mode and a second CABAC probability model different from the first CABAC probability model during the second mode; and encoding the video data based on the CABAC context, the mode, and the CABAC probability model corresponding to the mode.
[0088] Generally, at least one example embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors configured to: determine a CABAC context associated with a sub-block merging mode flag indicating use of a mode including an affine mode or a second mode different from the affine mode, the CABAC context corresponding to a first CABAC probability model during the affine mode and a second CABAC probability model different from the first CABAC probability model during the second mode; and encode the video data based on the CABAC context, the mode, and the CABAC probability model corresponding to the mode.
[0089] In general, at least one example embodiment may include a method for decoding video data, the method including: determining a CABAC context associated with a sub-block merging mode flag indicating use of a mode including an affine mode or a second mode different from the affine mode, the CABAC context corresponding to a first CABAC probability model during the affine mode and a second CABAC probability model different from the first CABAC probability model during the second mode; and decoding the video data encoded based on the affine mode based on the first CABAC probability model and decoding the video data encoded based on the second mode based on the second CABAC context.
[0090] Generally, at least one example embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors configured to: determine a CABAC context associated with a sub-block merging mode flag indicating use of a mode including an affine mode or a second mode different from the affine mode, the CABAC context corresponding to a first CABAC probability model during the affine mode and a second CABAC probability model different from the first CABAC probability model during the second mode; decode video data encoded based on the affine mode based on the first CABAC probability model; and decode video data encoded based on the second mode based on the second CABAC context.
[0091] In general, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, wherein the affine mode includes an AMVP mode and the second mode includes a merge mode.
[0092] In general, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, wherein the affine mode includes an AMVP mode and the second mode includes one of merge, SbTMVP, mmvd, or DMVR.
[0093] In general, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, wherein determining a CABAC context, such as a first or second CABAC context, does not consider spatial affine neighborhoods.
[0094] Generally, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, wherein a CABAC context, such as a first or second CABAC context, has only one context.
[0095] In general, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, wherein determining a CABAC context, such as a first or second CABAC context, is based solely on whether spatial neighborhoods are available.
[0096] In general, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, including constructing virtual affine candidates to be considered when determining a CABAC context, such as a first or second CABAC context.
[0097] In general, at least one example embodiment may include a method or apparatus according to any embodiment described herein, including a virtual affine candidate, where constructing the virtual affine candidate is based on the neighboring CU being coded in inter mode and not affine mode.
[0098] In general, at least one example embodiment may include a method or apparatus according to any of the embodiments described herein, including constructing virtual affine candidates based on neighboring CUs coded in inter mode, wherein the CABAC context includes one of: 0 if no inter neighbors are available, 1 if inter neighbors are available but no affine neighbors are available, and 2 if affine neighbors are available.
[0099] In general, at least one example of an embodiment may include a method or apparatus according to any embodiment described herein, further including considering the presence of reference pictures that are interpictures to enable creation of virtual temporal candidates.
[0100] In general, at least one example embodiment may include a method or apparatus according to any embodiment described herein, including constructing a virtual time candidate, where the virtual time candidate may be constructed based on temporally co-located CUs coded in inter mode.
[0101] In general, at least one example embodiment may include a method or apparatus according to any embodiment described herein, including constructing a virtual temporal candidate based on a co-located CU coded in inter mode, and determining the context includes that an inter neighbor is available or that an inter reference picture with reference picture index 0 is available.
[0102] In general, at least one example of an embodiment may include a computer program product including computing instructions for performing a method according to any embodiment described herein when executed by one or more processors.
[0103] In general, at least one example of an embodiment may include a non-transitory computer-readable medium that stores executable program instructions that cause a computer executing the instructions to perform a method according to any of the embodiments described herein.
[0104] In general, at least one example embodiment may include a bitstream formatted to include encoded video data, the encoded video data including an indicator associated with obtaining a CABAC context according to any of the methods described herein and picture data encoded based on the CABAC context.
[0105] In general, at least one example embodiment may include a device comprising an apparatus according to any embodiment described herein and at least one of: (i) an antenna configured to receive a signal, the signal including data representing video data; (ii) a band limiter configured to limit the received signal including the data representing the video data to a band of frequencies; and (iii) a display configured to display an image from the video data.
[0106] In general, at least one example embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors configured to: determine, during a first inter-prediction mode, a first CABAC context associated with a flag indicating use of an affine motion model; determine, during a second inter-prediction mode different from the first inter-prediction mode, a second CABAC context associated with a flag indicating use of an affine motion model different from the first CABAC context; and encode the video data to generate encoded video data, wherein, during the first inter-prediction mode, the video data generated based on the affine motion model is encoded based on the first CABAC probability model associated with the first CABAC context, and during the second inter-prediction mode, the video data generated based on the affine motion model is encoded based on the second CABAC probability model associated with the second CABAC context.
[0107] In general, at least one other example of an embodiment may include a method for encoding video data, the method including: determining, during a first inter-prediction mode, a first CABAC context associated with a flag indicating use of an affine motion model; determining, during a second inter-prediction mode different from the first inter-prediction mode, a second CABAC context associated with a flag indicating use of an affine motion model different from the first CABAC context; and encoding the video data to generate encoded video data, wherein, during the first inter-prediction mode, the video data generated based on the affine motion model is encoded based on a first CABAC probability model associated with the first CABAC context, and during the second inter-prediction mode, the video data generated based on the affine motion model is encoded based on a second CABAC probability model associated with the second CABAC context.
[0108] In general, at least one other example of an embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors configured to: determine, during a first inter-prediction mode, a first CABAC context associated with a flag indicating use of an affine motion model; determine, during a second inter-prediction mode different from the first inter-prediction mode, a second CABAC context associated with a flag indicating use of an affine motion model different from the first CABAC context; and decode the video data to generate decoded video data, wherein, during the first inter-prediction mode, the video data generated based on the affine motion model is decoded based on a first CABAC probability model associated with the first CABAC context, and during the second inter-prediction mode, the video data generated based on the affine motion model is decoded based on a second CABAC probability model associated with the second CABAC context.
[0109] In general, at least one other example of an embodiment may include a method for decoding video data, the method including: determining, during a first inter-prediction mode, a first CABAC context associated with a flag indicating use of an affine motion model; determining, during a second inter-prediction mode different from the first inter-prediction mode, a second CABAC context associated with a flag indicating use of an affine motion model different from the first CABAC context; and decoding the video data to generate decoded video data, wherein, during the first inter-prediction mode, the video data generated based on the affine motion model is decoded based on a first CABAC probability model associated with the first CABAC context, and during the second inter-prediction mode, the video data generated based on the affine motion model is decoded based on a second CABAC probability model associated with the second CABAC context.
[0110] In general, at least one other example of an embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors configured to: determine a first CABAC context associated with an affine motion compensation flag during a first inter prediction mode; determine a second CABAC context associated with an affine motion compensation flag during a second inter prediction mode different from the first inter prediction mode, wherein determining the second CABAC context occurs independently from determining the first CABAC context; and encode the video data to generate encoded video data based on the affine motion compensation flag and the first and second CABAC contexts during the first and second inter prediction modes, respectively.
[0111] In general, at least one other example of an embodiment may include a method for encoding video data, the method including: determining, during a first inter prediction mode, a first CABAC context associated with an affine motion compensation flag; and determining, during a second inter prediction mode different from the first inter prediction mode, a second CABAC context associated with the affine motion compensation flag, wherein the determining of the second CABAC context occurs independently from the determining of the first CABAC context; and encoding the video data to generate encoded video data based on the affine motion compensation flag and the first and second CABAC contexts during the first and second inter prediction modes, respectively.
[0112] Generally, at least one other example of an embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors configured to: determine a first CABAC context associated with an affine motion compensation flag during a first inter prediction mode; determine a second CABAC context associated with an affine motion compensation flag during a second inter prediction mode different from the first inter prediction mode, wherein determining the second CABAC context occurs independently from determining the first CABAC context; and decode the video data to generate decoded video data based on the affine motion compensation flag and the first and second CABAC contexts during the first and second inter prediction modes, respectively.
[0113] Generally, at least one other example of an embodiment may include a method for decoding video data, the method including: determining, during a first inter prediction mode, a first CABAC context associated with an affine motion compensation flag; and determining, during a second inter prediction mode different from the first inter prediction mode, a second CABAC context associated with the affine motion compensation flag, wherein the determining of the second CABAC context occurs independently from the determining of the first CABAC context; and decoding the video data to generate decoded video data based on the affine motion compensation flag and the first and second CABAC contexts during the first and second inter prediction modes, respectively.
[0114] In general, at least one other example of an embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors being configured to: determine an inter-prediction mode; acquire a first CABAC context associated with an affine motion compensation flag based on the inter-prediction mode being a first mode; acquire a second CABAC context associated with an affine motion compensation flag based on the inter-prediction mode being a second mode different from the first mode, wherein acquiring the second CABAC context occurs independently from acquiring the first CABAC context; and, during the first and second modes, encode the video data to generate encoded video data based on the affine motion compensation flag and the first and second CABAC contexts, respectively.
[0115] In general, at least one other example of an embodiment may include a method for encoding video data, the method including: determining an inter prediction mode; acquiring a first CABAC context associated with an affine motion compensation flag based on the inter prediction mode being a first mode; acquiring a second CABAC context associated with an affine motion compensation flag based on the inter prediction mode being a second mode different from the first mode, wherein acquiring the second CABAC context occurs independently from acquiring the first CABAC context; and encoding the video data to generate encoded video data based on the affine motion compensation flag and the first and second CABAC contexts during the first and second modes, respectively.
[0116] In general, at least one other example of an embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors being configured to: determine an inter-prediction mode; acquire a first CABAC context associated with an affine motion compensation flag based on the inter-prediction mode being a first mode; acquire a second CABAC context associated with an affine motion compensation flag based on the inter-prediction mode being a second mode different from the first mode, wherein acquiring the second CABAC context occurs independently from acquiring the first CABAC context; and, during the first and second modes, decode the video data to generate decoded video data based on the affine motion compensation flag and the first and second CABAC contexts, respectively.
[0117] In general, at least one other example of an embodiment may include a method for decoding video data, the method including: determining an inter prediction mode; acquiring a first CABAC context associated with an affine motion compensation flag based on the inter prediction mode being a first mode; acquiring a second CABAC context associated with an affine motion compensation flag based on the inter prediction mode being a second mode different from the first mode, the acquiring the second CABAC context occurring independently from the acquiring the first CABAC context; and decoding the video data to generate decoded video data based on the affine motion compensation flag and the first and second CABAC contexts during the first and second modes, respectively.
[0118] In general, at least one other example of an embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors being configured to: determine, during an inter-prediction mode, a CABAC context associated with a flag indicating the use of an affine motion model, where if the inter-prediction mode is a first mode, the CABAC context corresponds to a first CABAC probability model, and if the inter-prediction mode is a second mode different from the first mode, the CABAC context corresponds to a second CABAC probability model different from the first CABAC probability model; and encode the video data based on the flag, the CABAC context, the inter-prediction mode, and one of the first and second CABAC probability models corresponding to the inter-prediction mode.
[0119] Generally, at least one other example of an embodiment may include a method for encoding video data, the method including: determining a CABAC context associated with a flag indicating the use of an affine motion model during an inter-prediction mode, wherein if the inter-prediction mode is a first mode, the CABAC context corresponds to a first CABAC probability model, and if the inter-prediction mode is a second mode different from the first mode, the CABAC context corresponds to a second CABAC probability model different from the first CABAC probability model; and encoding the video data based on the flag, the CABAC context, the inter-prediction mode, and one of the first and second CABAC probability models corresponding to the inter-prediction mode.
[0120] In general, at least one other example of an embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors being configured to: determine, during an inter-prediction mode, a CABAC context associated with a flag indicating the use of an affine motion model, where if the inter-prediction mode is a first mode, the CABAC context corresponds to a first CABAC probability model, and if the inter-prediction mode is a second mode different from the first mode, the CABAC context corresponds to a second CABAC probability model different from the first CABAC probability model; and decode the video data based on the flag, the CABAC context, the inter-prediction mode, and one of the first and second CABAC probability models corresponding to the inter-prediction mode.
[0121] Generally, at least one other example of an embodiment may include a method for decoding video data, the method including: determining a CABAC context associated with a flag indicating the use of an affine motion model during an inter-prediction mode, wherein if the inter-prediction mode is a first mode, the CABAC context corresponds to a first CABAC probability model, and if the inter-prediction mode is a second mode different from the first mode, the CABAC context corresponds to a second CABAC probability model different from the first CABAC probability model; and decoding the video data based on the flag, the CABAC context, the inter-prediction mode, and one of the first and second CABAC probability models corresponding to the inter-prediction mode.
[0122] In general, at least one other example of an embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors configured to: acquire a flag indicating use of an affine motion model; determine a CABAC context associated with the flag, where the CABAC context is the only CABAC context associated with the flag and is determined without CABAC context modeling based on a spatial affine neighborhood of a current coding unit; and encode the video data to generate encoded video data based on the affine flag and the CABAC context.
[0123] In general, at least one other example embodiment may include a method for encoding video data, the method including the steps of: acquiring a flag indicating use of an affine motion model; determining a CABAC context associated with the flag, where the CABAC context is the only CABAC context associated with the flag and is determined without CABAC context modeling based on a spatial affine neighborhood of the current coding unit; and encoding the video data to generate encoded video data based on the affine flag and the CABAC context.
[0124] In general, at least one other example of an embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors configured to: acquire a flag indicating use of an affine motion model; determine a CABAC context associated with the flag, where the CABAC context is the only CABAC context associated with the flag and is determined without CABAC context modeling based on spatial affine neighborhoods of the current coding unit; and decode the video data to generate decoded video data based on the affine flag and the CABAC context.
[0125] In general, at least one other example embodiment may include a method for decoding video data, the method including the steps of: acquiring a flag indicating use of an affine motion model; determining a CABAC context associated with the flag, where the CABAC context is the only CABAC context associated with the flag and is determined without CABAC context modeling based on spatial affine neighborhoods of the current coding unit; and decoding the video data to generate decoded video data based on the affine flag and the CABAC context.
[0126] In general, at least one other example of an embodiment may include an apparatus for encoding video data, comprising one or more processors, the one or more processors being configured to: determine a CABAC context associated with a flag indicating the use of an affine motion model during an inter-prediction mode based solely on the availability of affine spatial neighbors of a current coding unit, wherein the CABAC context is one of a first context corresponding to availability indicating that no spatial affine neighbors are available or a second context corresponding to availability indicating that at least one spatial affine neighbor is available; obtain a CABAC probability model based on the CABAC context; and encode the video data based on the flag, the CABAC context, and the CABAC probability model corresponding to the CABAC context.
[0127] In general, at least one other example of an embodiment may include a method for encoding video data, the method including: determining a CABAC context associated with a flag indicating the use of an affine motion model during an inter-prediction mode based solely on the availability of affine spatial neighbors of a current coding unit, the CABAC context being one of a first context corresponding to an availability indicating that no spatial affine neighbors are available or a second context corresponding to an availability indicating that at least one spatial affine neighbor is available; obtaining a CABAC probability model based on the CABAC context; and encoding the video data based on the flag, the CABAC context, and the CABAC probability model corresponding to the CABAC context.
[0128] In general, at least one other example of an embodiment may include an apparatus for decoding video data, comprising one or more processors, the one or more processors being configured to: determine, based solely on the availability of affine spatial neighbors of a current coding unit, a CABAC context associated with a flag indicating the use of an affine motion model during an inter-prediction mode, wherein the CABAC context is one of a first context corresponding to an availability indicating that no spatial affine neighbors are available or a second context corresponding to an availability indicating that at least one spatial affine neighbor is available; obtain a CABAC probability model based on the CABAC context; and decode the video data based on the flag, the CABAC context, and the CABAC probability model corresponding to the CABAC context.
[0129] In general, at least one other example of an embodiment may include a method for decoding video data, the method including: determining a CABAC context associated with a flag indicating the use of an affine motion model during an inter-prediction mode based solely on the availability of affine spatial neighbors of a current coding unit, the CABAC context being one of a first context corresponding to an availability indicating that no spatial affine neighbors are available or a second context corresponding to an availability indicating that at least one spatial affine neighbor is available; obtaining a CABAC probability model based on the CABAC context; and decoding the video data based on the flag, the CABAC context, and the CABAC probability model corresponding to the CABAC context.
[0130] Various example embodiments described and contemplated in this document can be implemented in many different forms. Figures 14, 15, and 16 provide examples of some embodiments, as described below, although other embodiments are contemplated, and the descriptions of Figures 14, 15, and 16 do not limit the breadth of implementations. At least one of the aspects relates generally to video encoding and decoding, and at least one other aspect relates generally to transmitting a generated or encoded bitstream. These and other embodiments, features, aspects, etc. can be implemented as a method, an apparatus, a computer-readable storage medium having stored thereon instructions for encoding or decoding video data according to any of the described methods, and / or a computer-readable storage medium having stored thereon a bitstream generated according to any of the described methods.
[0131] In this application, the terms "reconstructed" and "decoded" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image," "picture," and "frame" can be used interchangeably. Typically, but not necessarily, the term "reconstructed" is used on the encoder side, while "decoded" is used on the decoder side.
[0132] Various methods are described herein, each of which includes one or more steps or actions for achieving the described method. Unless a specific order of steps or actions is required for the proper operation of the method, the order and / or use of specific steps and / or actions can be varied or combined.
[0133] Various methods and other aspects described in this document can be used to modify modules, e.g., entropy encoding module 145 and / or entropy decoding module 230 of video encoder 100 and decoder 200, respectively, as shown in Figures 14 and 15. Furthermore, the aspects are not limited to VVC or HEVC, but can be applied, for example, to other standards and recommendations, whether existing or developed in the future, and to extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically precluded, the aspects described in this document can be used individually or in combination.
[0134] Various numerical values are used in this document, for example, {{1,0}, {3,1}, {1,1}}, etc. The specific values are for illustrative purposes, and the described aspects are not limited to these specific values.
[0135] 14 illustrates an example of an encoder 100. Although variations of this encoder 100 are contemplated, the encoder 100 is described below for clarity without describing all possible variations.
[0136] Before being encoded, the video sequence may pass through a pre-encoding process (101) that, for example, applies a color transformation to the input color picture (e.g., converting from RGB 4:4:4 to YCbCr 4:2:0) or performs a remapping of the input picture components to obtain a signal distribution that is more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata may be associated with the pre-processing and attached to the bitstream.
[0137] In encoder 100, a picture is encoded by encoder elements as described below. The encoded picture is partitioned (102) and processed, for example, in units of CUs. Each unit is encoded, for example, using either intra mode or inter mode. When a unit is encoded in intra mode, it performs intra prediction (160). In inter mode, motion estimation (175) and compensation (170) are performed. The encoder decides (105) whether to use intra mode or inter mode to encode the unit and indicates the intra / inter decision, for example, by a prediction mode flag. A prediction residual is calculated, for example, by subtracting (110) the predicted block from the original image block.
[0138] The prediction residual is then transformed (125) and quantized (130). The quantized transform coefficients, as well as motion vectors and other syntax elements, are entropy coded (145) to output a bitstream. The encoder can skip the transform and apply quantization directly to the untransformed residual signal. The encoder can bypass both the transform and quantization, i.e., the residual is coded directly without applying a transform or quantization process.
[0139] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual is combined with the predicted block (155) to reconstruct an image block. An in-loop filter (165) is applied to the reconstructed picture to reduce encoding artifacts, for example, to perform deblocking / SAO (sample adaptive offset) filtering. The filtered image is stored in a reference picture buffer (180).
[0140] Figure 15 illustrates a block diagram of an example video decoder 200. In the decoder 200, the bitstream is decoded by decoder elements as described below. The video decoder 200 generally performs a decoding pass that is the reciprocal of the encoding pass as described with respect to Figure 14. The encoder 100 also generally performs video decoding as part of encoding the video data.
[0141] The decoder's input includes a video bitstream, which may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other coded information. Picture partition information indicates how the picture is partitioned. Thus, the decoder can divide (235) the picture according to the decoded picture's partitioning information. The transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. An image block is reconstructed by combining (255) the decoded prediction residual with a predicted block. The predicted block can be obtained (270) from intra prediction (260) or motion-compensated prediction (i.e., inter prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in a reference picture buffer (280).
[0142] The decoded picture may further pass through a post-decoding process (285), such as an inverse color conversion (e.g., YCbCr 4:2:0 to RGB 4:4:4) or an inverse remapping that performs the inverse of the remapping process performed in the pre-encoding process (101). The post-decoding process may use metadata derived in the pre-encoding process and conveyed in the bitstream.
[0143] FIG. 16 illustrates a block diagram of an example system in which various aspects and embodiments may be implemented. System 1000 may be embodied as a device including various components described below and configured to perform one or more of the aspects described herein. Examples of such devices include, but are not limited to, various electronic devices, such as personal computers, laptop computers, smartphones, tablet computers, digital multimedia set-top boxes, digital television sets, personal video recording systems, connected home appliances, and servers. The elements of system 1000 may be embodied, singly or in combination, in a single integrated circuit, multiple ICs, and / or separate components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or separate components. In various embodiments, system 1000 is communicatively coupled to other similar systems or other electronic devices, for example, via a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to perform one or more of the aspects described herein.
[0144] The system 1000 includes at least one processor 1010 configured to execute instructions loaded therein, for example, to implement various aspects described herein. The processor 1010 may include embedded memory, input / output interfaces, and various other circuits known in the art. The system 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non-volatile memory device). The system 1000 includes a storage device 1040, which may include non-volatile memory and / or volatile memory, including, but not limited to, EEPROM, ROM, PROM, RAM, DRAM, SRAM, flash, magnetic disk drives, and / or optical disk drives. The storage device 1040 may include, by way of non-limiting example, an internal storage device, an attached storage device, and / or a network-accessible storage device.
[0145] System 1000 includes, for example, an encoder / decoder module 1030 configured to process data to provide encoded or decoded video, which may include its own processor and memory. Encoder / decoder module 1030 represents a module that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding module and a decoding module. Additionally, encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software, as is known to those skilled in the art.
[0146] Program code loaded onto the processor 1010 or the encoder / decoder 1030 to perform various aspects described herein may be stored in the storage device 1040 and then loaded onto the memory 1020 for execution by the processor 1010. According to various embodiments, one or more of the processor 1010, the memory 1020, the storage device 1040, and the encoder / decoder module 1030 may store one or more of various items during execution of the processes described herein. Such stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and arithmetic logic.
[0147] In some embodiments, memory internal to the processor 1010 and / or the encoder / decoder module 1030 is used to store instructions and provide working memory for processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device can be either the processor 1010 or the encoder / decoder module 1030) is used for one or more of these functions. The external memory can be the memory 1020 and / or the storage device 1040, e.g., dynamic volatile memory and / or non-volatile flash memory. In some embodiments, the external non-volatile flash memory is used to store the television's operating system. In at least one embodiment, high-speed external dynamic volatile memory, such as RAM, is used as working memory for video encoding and decoding operations, such as MPEG-2, HEVC, or VVC (Versatile Video Coding).
[0148] Input to the elements of system 1000 can be provided through various input devices, as shown in block 1130. Such input devices include, but are not limited to, (i) an RF section that receives RF signals, e.g., transmitted over the air by a broadcast station, (ii) a composite input terminal, (iii) a USB input terminal, and / or (iv) an HDMI input terminal.
[0149] In various embodiments, the input devices of block 1130 have associated respective input processing elements as known in the art. For example, the RF section may be associated with elements necessary to (i) select a desired frequency (also referred to as selecting a signal or band-limiting a signal to a band of frequencies), (ii) downconvert the selected signal, (iii) band-limit again to a narrower band of frequencies to select (for example) a signal frequency band, which in some embodiments may be referred to as a channel, (iv) demodulate the downconverted and band-limited signal, (v) perform error correction, and (vi) demultiplex to select a desired stream of data packets. The RF section of various embodiments includes one or more elements for performing these functions, such as a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a downconverter, a demodulator, an error corrector, and a demultiplexer. The RF section may include, for example, a tuner that performs a variety of these functions, including downconverting a received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or to baseband. In one embodiment of a set-top box, the RF section and its associated input processing elements perform frequency selection by receiving RF signals transmitted over a wired (e.g., cable) medium, filtering to a desired frequency band, downconverting, and filtering again. Various embodiments rearrange the order of the above-described (and other) elements, delete some of these elements, and / or add other elements that perform similar or different functions. Adding elements can include inserting elements between existing elements, such as inserting amplifiers and analog-to-digital converters. In various embodiments, the RF section includes an antenna.
[0150] Additionally, the USB and / or HDMI terminals may include respective interface processors for connecting system 1000 to other electronic devices via USB and / or HDMI connections. It should be understood that various aspects of input processing, e.g., Reed-Solomon error correction, may be implemented, for example, in a separate input processing IC or within processor 1010, as desired. Similarly, aspects of USB or HDMI interface processing may be implemented, for example, in a separate interface IC or within processor 1010, as desired. The demodulated, error corrected, and demultiplexed stream is provided to various processing elements, including, for example, processor 1010 and encoder / decoder 1030, operating in combination with memory and storage elements to process the data stream, as desired, for presentation on an output device.
[0151] The various elements of system 1000 may be provided within an integrated housing in which the various elements are interconnected and can transmit data between them using a suitable connection arrangement 1140, e.g., an internal bus known in the art, including an I2C bus, wiring, and a printed circuit board.
[0152] The system 1000 includes a communication interface 1050 that enables communication with other devices over a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data over the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented within a wired and / or wireless medium, for example.
[0153] In various embodiments, data is streamed to system 1000 using a wireless network, such as IEEE 802.11. The wireless signal in these embodiments is received over communication channel 1060 and communication interface 1050, adapted for example for Wi-Fi communication. Communication channel 1060 in these embodiments is typically connected to an access point or router that provides access to external networks, including the Internet, to enable streaming applications and other over-the-top communications. Other embodiments provide streamed data to system 1000 using a set-top box that delivers data over the HDMI connection of input block 1130. Still other embodiments provide streamed data to system 1000 using the RF connection of input block 1130.
[0154] System 1000 can provide output signals to various output devices, including a display 1100, speakers 1110, and other peripheral devices 1120. In various example embodiments, other peripheral devices 1120 include one or more of a standalone DVR, a disc player, a stereo system, a lighting system, and other devices that provide functionality based on the output of system 1000. In various embodiments, control signals are communicated between system 1000 and display 1100, speakers 1110, or other peripheral devices 1120 using signaling such as AV.Link, CEC, or other communication protocols that allow inter-device control with or without user intervention. Output devices can be communicatively coupled to system 1000 via dedicated connections through respective interfaces 1070, 1080, and 1090. Alternatively, output devices can be connected to system 1000 using communication channel 1060 via communication interface 1050. The display 1100 and speakers 1110 can be integrated into a single unit along with other components of the system 1000 in an electronic device, such as a television. In various embodiments, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip.
[0155] Alternatively, the display 1100 and speakers 1110 can be separate from one or more of the other components, for example, if the RF section of the input 1130 is part of a separate set-top box. In various embodiments in which the display 1100 and speakers 1110 are external components, the output signal can be provided via a dedicated output connection, including, for example, an HDMI port, a USB port, or a COMP output.
[0156] The embodiments may be implemented by computer software executed by the processor 1010, by hardware, or by a combination of hardware and software. As a non-limiting example, the embodiments may be implemented by one or more integrated circuits. The memory 1020 may be of any type appropriate to the technology environment and may be implemented using any suitable data storage technology, such as, by way of non-limiting examples, optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 may be of any type appropriate to the technology environment and may include, by way of non-limiting examples, one or more of a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.
[0157] Various implementations include decoding. As used herein, "decoding" can encompass all or part of the processes performed, for example, on a received encoded sequence to generate a final output suitable for display. In various embodiments, such processes include processes typically performed by a decoder, such as one or more of entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also include, or alternatively include, processes performed by decoders of various implementations described herein, such as extracting indexes of weights used for various intra-prediction reference arrays.
[0158] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to refer specifically to a subset of operations or to the broader decoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.
[0159] Various implementations include encoding. Similar to the above description of "decoding," as used herein, "encoding" can encompass all or part of the processes performed, for example, on an input video sequence to generate an encoded bitstream. In various embodiments, such processes include one or more of processes typically performed by an encoder, such as partitioning, differential encoding, transforming, quantization, and entropy encoding. In various embodiments, such processes also include, or alternatively include, processes performed by encoders of various implementations described herein, such as weighting of intra-prediction reference arrays.
[0160] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to refer specifically to a subset of operations or to the broader encoding process generally will be clear based on the context of the specific description and is believed to be well understood by one of ordinary skill in the art.
[0161] Note that syntax elements, as used herein, are descriptive terms, so they do not preclude the use of other syntax element names.
[0162] When a figure is presented as a flow diagram, it should be understood that it also provides a block diagram of the corresponding apparatus. Similarly, when a figure is presented as a block diagram, it should be understood that it also provides a flow diagram of the corresponding method / process.
[0163] Various embodiments refer to rate-distortion calculation or rate-distortion optimization. During the encoding process, a balance or trade-off between rate and distortion is typically considered, often subject to computational complexity constraints. Rate-distortion optimization is typically formulated as minimizing a rate-distortion function, which is a weighted sum of rate and distortion. Different approaches exist for solving the rate-distortion optimization problem. For example, an approach can be based on extensive testing of all encoding options, including all considered modes or coding parameter values, with a thorough evaluation of their coding costs and associated distortion of the reconstructed signal after encoding and decoding. To reduce encoding complexity, faster approaches can also be used, particularly those that use approximated distortion calculations based on predicted or prediction residual signals rather than reconstructed ones. A hybrid of these two approaches can also be used, such as by using approximated distortion for only some of the possible encoding options and full distortion for other encoding options. Other approaches evaluate only a subset of the possible encoding options. More generally, many approaches utilize any of a variety of techniques to perform optimization, but the optimization is not necessarily a complete assessment of both the coding cost and the associated distortion.
[0164] The implementations and aspects described herein may be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if described only in the context of a single form of implementation (e.g., described only as a method), the implementation of the described features may also be implemented in other forms (e.g., an apparatus or a program). An apparatus may be implemented, for example, in appropriate hardware, software, and firmware. A method may be implemented, for example, in a processor, which generally refers to a processing device and includes, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. Processors also include communication devices, such as, for example, computers, cell phones, portable / personal digital assistants ("PDAs"), and other devices that facilitate communication of information between end users.
[0165] References to "one embodiment" or "embodiment," or "one implementation" or "implementation," and other variations thereof, mean that particular features, structures, characteristics, etc. described in connection with an embodiment are included in at least one embodiment. Thus, appearances of the phrases "in one embodiment" or "in an embodiment," or "in one implementation" or "in an implementation," and any other variations thereof, appearing in various places throughout this document are not necessarily all referring to the same embodiment.
[0166] Additionally, this document may refer to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from memory.
[0167] Additionally, this document may refer to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.
[0168] Additionally, this document may refer to "receiving" various information. Receiving, like "accessing," is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is typically included in various ways during operations such as, for example, storing information, processing information, transmitting information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.
[0169] For example, it should be understood that the use of any of the following " / ," "and / or," and "at least one of" in the cases of "A / B," "A and / or B," and "at least one of A and B" is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of both alternatives (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C," such language is intended to encompass the selection of only the first listed alternative (A), or the selection of only the second listed alternative (B), or the selection of only the third listed alternative (C), or the selection of only the first and second listed alternatives (A and B), or the selection of only the first and third listed alternatives (A and C), or the selection of only the second and third listed alternatives (B and C), or the selection of all three alternatives (A, B, and C). This can be expanded by the number of items listed, as will be apparent to those skilled in the art and related fields.
[0170] Also, as used herein, the term "signal" refers, among other things, to indicating something to a corresponding decoder. For example, in one embodiment, an encoder signals a specific one of multiple weights to be used for an intra-prediction reference array. Thus, in an embodiment, the same parameters are used on both the encoder and decoder sides. Thus, for example, the encoder can transmit specific parameters to the decoder so that the decoder can use the same specific parameters (explicit signaling). Conversely, if the decoder already has certain parameters and others, it can use signaling without transmission to simply enable the decoder to know and select certain parameters (implicit signaling). By avoiding the transmission of any actual function, bit savings are realized in various embodiments. It should be understood that signaling can be achieved in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to convey information to a corresponding decoder. Although the above relates to the verb form of the word "signal," the word "signal" can also be used as a noun in this specification.
[0171] As will be apparent to those skilled in the art, implementations can generate a variety of signals formatted to carry information that can be stored or transmitted, for example. The information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted, for example, as an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or as a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog information or digital information. The signal can be transmitted over a variety of different wired or wireless links, as is known. The signal can be stored on a processor-readable medium.
[0172] Embodiments may include one or more of the following features or entities, alone or in combination, across a variety of different claim categories and types.
[0173] Affine mode encoding and decoding, taking into account the different statistical occurrences of affine mode usage during AMVP and merging.
[0174] Affine mode encoding and decoding takes into account the different statistical occurrence of affine mode usage between AMVP and merge by encoding the affine flag independently in AMVP and merge modes.
[0175] Affine mode encoding and decoding that takes into account the different statistical occurrences of affine mode usage between AMVP and merge by encoding the affine flag independently in AMVP and merge mode, whereby the total number of CABAC contexts for the affine flag is doubled.
[0176] Affine mode encoding and decoding that takes into account the different statistical occurrences of affine mode usage between AMVP and merge by encoding the affine flag independently in AMVP and merge mode, where the total number of CABAC contexts for affine flags is doubled and the actual affine flag context modeling uses six contexts instead of three.
[0177] • Removing CABAC context modeling based on spatial affine neighborhoods.
[0178] • Eliminating CABAC context modeling based on spatial affine neighborhoods using only one context.
[0179] • Modeling context using only the availability of spatial neighborhoods.
[0180] ● Modeling the context using only whether spatial neighborhoods are available, where two contexts are possible: no affine neighborhoods are available (context 0) or at least one affine neighborhood is available (context 1).
[0181] • In context modeling, constructing virtual affine candidates to be considered, constructed from neighboring CUs that are coded in inter mode but not in affine mode.
[0182] In context modeling, constructing virtual affine candidates to be considered, constructed from neighboring CUs coded in inter mode but not in affine mode, where the context is ○ 0 if no inter-neighborhood is available, ○ If inter-neighborhood is available but affine neighborhood is not available, then 1, ○2 if affine neighborhood is available That is, to build.
[0183] Considering the existence of reference pictures that are interpictures to enable the creation of virtual temporal candidates.
[0184] ●Consider the presence of reference pictures that are interpictures to enable the creation of virtual time candidates, where the virtual time candidates can be constructed from temporally co-located CUs coded in inter mode.
[0185] Considering the presence of reference pictures that are inter pictures to enable the creation of virtual temporal candidates, which can be constructed from temporally co-located CUs coded in inter mode, and context modeling ○ An inter neighbor is available or an inter reference picture with reference picture index 0 is available To consider, including:
[0186] • A bitstream or signal containing one or more of the described syntax elements or variations thereof.
[0187] Inserting syntax elements into the signaling that allow the decoder to provide affine mode processing in a manner that corresponds to that used by the encoder.
[0188] - Selecting the affine mode processing to apply at the decoder based on these syntax elements.
[0189] • Creating and / or transmitting and / or receiving and / or decoding a bitstream or signal that includes one or more of the described syntax elements or variations thereof.
[0190] ● A television, set-top box, cell phone, tablet, or other electronic device that performs any of the described embodiments.
[0191] ● A television, set-top box, cell phone, tablet, or other electronic device that performs any of the described embodiments and displays the resulting images (e.g., using a monitor, screen, or other type of display).
[0192] ●A television, set-top box, cell phone, tablet, or other electronic device that tunes a channel (e.g., using a tuner) to receive a signal containing an encoded image and performs any of the described embodiments.
[0193] ● A television, set-top box, cell phone, tablet, or other electronic device that receives a signal containing an encoded image wirelessly (e.g., using an antenna) and performs any of the described embodiments.
[0194] A computer program product storing program code which, when executed by a computer, implements any of the described embodiments.
[0195] • A non-transitory computer-readable medium containing executable program instructions that cause a computer executing the instructions to perform any of the described embodiments.
[0196] Various other generalized and specialized embodiments are also supported and contemplated throughout this disclosure. [Industrial Applicability]
[0197] The present invention can be used in video signal processing.
Claims
1. obtaining a CABAC context for encoding a flag, the flag indicating use of an affine mode for predicting a block of video in a first inter mode or in a second inter mode; encoding the flag using the CABAC context, wherein encoding the flag takes into account the occurrence of use of the affine mode between the first inter mode and the second inter mode; A method for providing
2. obtaining a CABAC context for decoding a flag, the flag indicating use of an affine mode for predicting a block of video in a first inter mode or in a second inter mode; decoding the flag using the CABAC context, wherein decoding the flag takes into account the occurrence of use of the affine mode between the first inter mode and the second inter mode; A method for providing
3. The method of claim 1 or 2, wherein the CABAC context is shared between the first inter mode and the second inter mode.
4. The method of claim 1 or 2, wherein the CABAC context is obtained as a function of values of affine flags associated with neighboring blocks of the block.
5. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled in affine mode; 1 if the left neighboring block is signaled in the affine mode and the upper neighboring block is not signaled in the affine mode; 2 if both the left neighbor block and the upper neighbor block are signaled in the affine mode The method according to claim 4, wherein the method is any one of the following:
6. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled in affine mode; 1 if at least one of the left neighboring block or the upper neighboring block is signaled in the affine mode The method according to claim 1 or 2, wherein the method is any one of the following:
7. The method of claim 1 or 2, wherein the step of obtaining the CABAC context considers whether virtual affine candidates can be constructed from neighboring blocks.
8. 8. The method of claim 7, wherein virtual affine candidates are constructed from neighboring blocks if the neighboring blocks are coded in inter mode but without using an affine motion model.
9. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled in inter mode; is 1 if at least one of the left neighbor block or the above neighbor block is signaled in the inter mode and neither the left neighbor block nor the above neighbor block is signaled in the affine mode; 2 if at least one of the left neighboring block or the above neighboring block is signaled in the inter mode and at least one of the left neighboring block or the above neighboring block is signaled in the affine mode. The method according to claim 8, wherein the method is any one of the following:
10. 8. The method of claim 7, wherein the step of obtaining the CABAC context further includes considering whether a hypothetical temporal affine candidate can be constructed from a temporally co-located block in a reference picture coded in inter mode.
11. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled as inter mode and no inter reference picture with reference picture index 0 is available; 1 if at least one of the left neighboring block or the above neighboring block is signaled in the inter mode and neither the left neighboring block nor the above neighboring block is signaled in the affine mode, or an inter reference picture with reference picture index 0 is available; 2 if at least one of the left neighboring block or the above neighboring block is signaled in the inter mode and at least one of the left neighboring block or the above neighboring block is signaled in the affine mode. The method according to claim 10, wherein the method is any one of the following:
12. The method of claim 1 or 2, wherein the CABAC context is obtained using the same derivation in both the first inter mode and the second inter mode.
13. The step of encoding the flag comprises: if the block is in the first inter mode, obtaining a first CABAC probability model associated with the CABAC context and encoding the flag using the first CABAC probability model; and if the block is in the second inter mode, obtaining a second CABAC probability model associated with the CABAC context, the second CABAC probability model being different from the first CABAC probability model, and encoding the flag using the second CABAC probability model.
10. The method of claim 1, comprising:
14. The step of decoding the flag comprises: if the block is in the first inter mode, obtaining a first CABAC probability model associated with the CABAC context and decoding the flag using the first CABAC probability model; and if the block is in the second inter mode, obtaining a second CABAC probability model associated with the CABAC context, the second CABAC probability model being different from the first CABAC probability model, and decoding the flag using the second CABAC probability model.
3. The method of claim 2, comprising:
15. 3. The method of claim 1, wherein the use of the affine mode is signaled using three CABAC contexts, and each of the three CABAC contexts is associated with two CABAC probability models, one of the two CABAC probability models considering the occurrence of the use of the affine mode in the first inter mode, and another of the two CABAC probability models considering the occurrence of the use of the affine mode in the second inter mode.
16. The method according to claim 1 or 2, wherein the second inter mode is a merge mode.
17. 3. The method according to claim 1, wherein the first inter mode is an AMVP mode.
18. obtaining a CABAC context for encoding a flag, the flag indicating use of an affine mode for predicting a block of video in a first inter mode or in a second inter mode; The CABAC context is used to encode the flag, and encoding the flag takes into account the occurrence of use of the affine mode between the first inter mode and the second inter mode. One or more processors configured to A device comprising:
19. obtaining a CABAC context for decoding a flag, the flag indicating use of an affine mode for predicting a block of video in a first inter mode or in a second inter mode; Decode the flag using the CABAC context, and the decoding of the flag takes into account the occurrence of use of the affine mode between the first inter mode and the second inter mode. One or more processors configured to A device comprising:
20. 20. The apparatus of claim 18 or 19, wherein the CABAC context is shared between the first inter mode and the second inter mode.
21. Decoding the flag comprises: if the block is in the first inter mode, obtaining a first CABAC probability model associated with the CABAC context and decoding the flag using the first CABAC probability model; and if the block is in the second inter mode, obtaining a second CABAC probability model associated with the CABAC context, the second CABAC probability model being different from the first CABAC probability model, and decoding the flag using the second CABAC probability model.
20. The apparatus of claim 19, comprising:
22. 20. The apparatus of claim 18 or 19, wherein the use of the affine mode is signaled using three CABAC contexts, each of the three CABAC contexts being associated with two CABAC probability models, one of the two CABAC probability models considering the occurrence of use of the affine mode in the first inter mode and another of the two CABAC probability models considering the occurrence of use of the affine mode in the second inter mode.
23. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled in affine mode; 1 if the left neighboring block is signaled in the affine mode and the upper neighboring block is not signaled in the affine mode; 2 if both the left neighbor block and the upper neighbor block are signaled in the affine mode 20. The device according to claim 18 or 19, wherein
24. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled in affine mode; 1 if at least one of the left neighboring block or the upper neighboring block is signaled in the affine mode 20. The device according to claim 18 or 19, wherein
25. 20. The apparatus of claim 18 or 19, wherein obtaining the CABAC context considers whether virtual affine candidates can be constructed from neighboring blocks.
26. 20. The apparatus of claim 18 or 19, wherein virtual affine candidates are constructed from neighboring blocks if the neighboring blocks are coded in inter mode but without using an affine motion model.
27. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled in inter mode; is 1 if at least one of the left neighbor block or the above neighbor block is signaled in the inter mode and neither the left neighbor block nor the above neighbor block is signaled in the affine mode; 2 if at least one of the left neighboring block or the above neighboring block is signaled in the inter mode and at least one of the left neighboring block or the above neighboring block is signaled in the affine mode.
20. The device according to claim 18 or 19, wherein
28. 28. The apparatus of claim 27, wherein obtaining the CABAC context further considers whether a hypothetical temporal affine candidate can be constructed from a temporally co-located block in a reference picture coded in inter mode.
29. The value of the CABAC context is 0 if neither the left neighbor nor the top neighbor is signaled as inter mode and no inter reference picture with reference picture index 0 is available; 1 if at least one of the left neighboring block or the above neighboring block is signaled in the inter mode and neither the left neighboring block nor the above neighboring block is signaled in the affine mode, or an inter reference picture with reference picture index 0 is available; 2 if at least one of the left neighboring block or the above neighboring block is signaled in the inter mode and at least one of the left neighboring block or the above neighboring block is signaled in the affine mode.
20. The device according to claim 18 or 19, wherein
30. A computer-readable storage medium comprising computing instructions for performing the method of claim 1 or 2 when executed by one or more processors.
31. 20. The apparatus of claim 19; (i) an antenna configured to receive a signal, the signal including data representing video data; (ii) a band limiter configured to limit the received signal including the data representing the video data to a band of frequencies; and (iii) a display configured to display an image from the video data. A device with.