Apparatus and method for refined probability estimation of transform coefficients for video coding
By refining probability estimation through state-driven trellis-coded quantization and adaptive context modeling, the encoding of transform coefficients in video coding is optimized for efficiency and reduced complexity, addressing the challenges in VVC.
Patent Information
- Application Number
- PCT/EP2025/055225
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-29
- Filing Date
- 2025-02-26
- Publication Date
- 2025-09-04
AI Technical Summary
Existing video coding technologies face challenges in achieving efficient encoding of transform coefficients while balancing coding efficiency and implementation complexity, particularly in scenarios like Versatile Video Coding (VVC), where dependent quantization increases decoder complexity.
The proposed solution involves refining probability estimation of transform coefficients by employing a state-driven trellis-coded quantization (TCQ) with two scalar quantizers and adaptive context modeling, limiting context-coded bins to enhance throughput and reduce complexity.
This approach enhances coding efficiency and reduces decoder complexity by optimizing the transition between quantization sets and employing context models, resulting in improved encoding performance.
Smart Images

Figure EP2025055225_04092025_PF_FP_ABST
Abstract
Description
[0001] Apparatus and Method for Refined Probability Estimation of Transform Coefficients for Video Coding
[0002] Description
[0003] The present invention relates to video coding, in particular, to probability estimation of transform coefficients for video coding, and, more particularly, to an apparatus and a method for refined probability estimation of transform coefficients for video coding.
[0004] At first, level coding in state-of-the-art hybrid video compression is considered. In state-of- the-art hybrid video coding, such as Versatile Video Coding (VVC), prediction errors, known as residuals, undergo a transformation, and the resultant transform coefficients are quantized. These quantized levels are referred to as transform coefficient levels and are encoded into the bitstream using entropy coding. The process of encoding transform coefficient levels is referred to as level coding.
[0005] The description particularly emphasizes the design aspects specified in the Versatile Video Coding standard (VVC). It should, however, be noted that several variations of this particular design are possible, and the principles herein are applicable in other non-VVC scenarios and in non-VVC applications.
[0006] In particular, in block-based hybrid video coding, the pictures are in general partitioned into rectangular blocks. Each block is then predicted using intra- or inter-block prediction. Resulting error blocks are coded employing transform coding, which usually comprises an orthogonal transform, a quantization of transform coefficients, and an entropy coding of the resulting quantization indices.
[0007] In more detail, lossy coding of a block of residual samples is conducted. The residual samples represent the difference between the original block of samples and the samples of a prediction signal (the prediction signal can be obtained by intra-picture prediction or inter-picture prediction, or by a combination of intra- and inter-picture prediction, or by any other means; in a special case the prediction signal could be set equal to zero).
[0008] The residual block of samples is transformed using a signal transform. Typically, a linear and separable transform is used (linear means that the transforms are linear, but may incorporate an additional rounding of transform coefficients). Often, an integer approximation of the DCT-II or an integer approximation of other transforms of the DCT / DST family is used. Different transforms may be used in horizontal or vertical direction. The transform is not restricted to linear and separable transforms. Any other transform (linear and non-separable, or non-linear) may be used. In particular, a second order non-separable transform may be applied to the transform coefficients or a part of the transform coefficients obtained after a first separable transform. As result of the signal transform, a block of transform coefficients is obtained that represents the original block of residual samples in a different signal space. In a special case, the transform can be equal to the identify transform (i.e. , the block of transform coefficients can be equal to the block of residual samples). The block of transform coefficients is coded using lossy coding. At the decoder side, the block of reconstructed transform coefficients is inverse transformed so that a reconstructed block of residual samples is obtained. And finally, by adding the prediction signal, a reconstructed block of image samples is obtained.
[0009] Dependent quantization refers to a quantization method with the following properties: At the encoder side, the block of transform coefficients is mapped to a block of transform coefficient levels (i.e., quantization indexes), which represent the transform coefficients with reduced fidelity. At the decoder side, the quantization indexes are mapped to reconstructed transform coefficients (which differ from the original transform coefficients due to quantization). In contrast to conventional scalar quantization, the transform coefficients are not independently quantized. Instead, the admissible set of reconstruction levels for a certain transform coefficient depends on the chosen quantization indexes for other transform coefficients.
[0010] There are multiple variants of dependent quantization, which mainly deviate in the number of possible so-called quantization states. Increasing the number of quantization states typically improves the coding efficiency, but also increases the encoder complexity. Hence, it is reasonable to support multiple variants of dependent quantization, in order to give an encoder the freedom to choose a suitable trade-off between coding efficiency and implementation complexity. But for a decoder, the implementation complexity increases with the number of supported variants for quantization.
[0011] Dependent quantization of transform coefficients refers to a concept in which the set of available reconstruction levels for a transform coefficient depends on the chosen quantization indexes for preceding transform coefficients in reconstruction order (inside the same transform block). Multiple sets of reconstruction levels are pre-defined and, based on the quantization indexes for preceding transform coefficients (in coding order), one of the predefined sets is selected for reconstructing the current transform coefficient. The set of admissible reconstruction levels for a current transform coefficient is selected (based on the quantization indexes for preceding transform coefficients in coding order) among a collection (two or more sets) of pre-defined sets of reconstruction levels. The values of the reconstruction levels of the sets of reconstruction levels are parameterized by a block-based quantization parameter. The block-based quantization parameter (QP) determines a quantization step size A and all reconstruction levels (in all sets of reconstruction levels) represent integer multiples of the quantization step size A. The quantization step size Akfor a particular transform coefficient tk(with k indicating the reconstruction order) may not be solely determined by the block quantization parameter QP, but it is also possible that the quantization step size Akfor a particular transform coefficient tkis determined by a quantization weighting matrix and the block quantization parameter. Typically, the quantization step size Akfor a transform coefficient tkis given by the product of the weighting factor wkfor the transform coefficient tk(specified by the quantization weighting matrix) and the block quantization step size Ablock(specified by the block quantization parameter), fc = wk• Ablock.
[0012] In addition to the traditional scalar quantizer, VVC incorporates trellis-coded quantization (TCQ), which employs two scalar quantizers (e.g., Qo, Q ) with different alignments but the same step sizes. TCQ operates in a state-driven manner, meaning that the current state determines which quantizer out of the two quantizers is used for each scanning position. Each block begins with the initial state (SO). In VVC, the TCQ design comprises four states, where the first two states represent the utilization of the first quantizer, while the third and fourth states indicate the usage of the second quantizer.
[0013] The current state is determined based on the previous state (e.g., sk) and the parity of the previously decoded level. Since the level coding does not fully encode the level for each scanning position, the parity is encoded as a dedicated flag for context modeling purposes. A notable characteristic of TCQ is that the state is utilized to select among three different context model sets for sig. Specifically, the first two states point to the same context model set (a first set), while the third and fourth states indicate each a dedicated context model set (a second set and a third set).
[0014] In more detail, in VVC, the dependent scalar quantization for transform coefficients uses exactly two different sets of reconstruction levels. All reconstruction levels of the two sets for a transform coefficient tkrepresent integer multiples of the quantization step size Akfor this transform coefficient (which is, at least partly, determined by a block-based quantization parameter). Note that the quantization step size Akjust represents a scaling factor for the admissible reconstruction values in both sets. Except of a possible individual quantization step size Akfor the different transform coefficients tkinside a transform block (and, thus, an individual scaling factor), the same two sets of reconstruction levels are used for all transform coefficients.
[0015] Fig. 1 shows two sets of reconstruction levels. The hollow and filled circles indicate two different subsets inside the sets of reconstruction levels; the subsets can be used for determining the set of reconstruction levels for the next transform coefficient in reconstruction order.
[0016] The two sets of reconstruction levels are shown in Fig. 1:
[0017] • The reconstruction levels that are contained in the first quantization set (labeled as set O in fig. 1) represent the even integer multiples of the quantization step size.
[0018] • The second quantization set (labeled as set 1 in fig. 1) contains all odd integer multiples of the quantization step size and additionally the reconstruction level equal to zero.
[0019] Note that both reconstruction sets are symmetric about zero. The reconstruction level equal to zero is contained in both reconstruction sets, otherwise the reconstruction sets are disjoint. The union of both reconstruction sets contains all integer multiples of the quantization step size.
[0020] The reconstruction levels that the encoder selects among the admissible reconstruction levels have to be indicated inside the bitstream. As in conventional independent scalar quantization, this can be achieved using so-called quantization indexes, which are also referred to as transform coefficient levels. Quantization indexes (or transform coefficient levels) are integer numbers that uniquely identify the available reconstruction levels inside a quantization set (i.e. , inside a set of reconstruction levels). The quantization indexes are sent to the decoder as part of the bitstream (using any entropy coding technique). At the decoder side, the reconstructed transform coefficients can be uniquely calculated based on a current set of reconstruction levels (which is determined by the preceding quantization indexes in coding / reconstruction order) and the transmitted quantization index for the current transform coefficient.
[0021] The reconstruction levels in Fig. 1 are labeled with an associated quantization index (the quantization indexes are given by the numbers below the circles that represent the reconstruction levels). The quantization index equal to 0 is assigned to the reconstruction level equal to 0. The quantization index equal to 1 is assigned to the smallest reconstruction level greater than 0, the quantization index equal to 2 is assigned to the next reconstruction level greater than 0 (i.e., the second smallest reconstruction level greater than 0), etc. Or, in other words, the reconstruction levels greater than 0 are labeled with integer numbers greater than 0 (i.e., with 1, 2, 3, etc.) in increasing order of their values. Similarly, the quantization index -1 is assigned to the largest reconstruction level smaller than 0, the quantization index -2 is assigned to the next (i.e., the second largest) reconstruction level smaller than 0, etc. Or, in other words, the reconstruction levels smaller than 0 are labeled with integer numbers less than 0 (i.e., -1, -2, -3, etc.) in decreasing order of their values.
[0022] Before the scaling process, used when TOO is disabled, e.g., when operating with a single scalar quantizer, an intermediate level x' is derived from the decoded level x according to: x' = sgn(x)(2 ■ |x| - [S / 2J)-
[0023] The intermediate levels may, for example, then used as the input for the scaling process instead of the decoded level when TOO is enabled.
[0024] The usage of reconstruction levels that represent integer multiples of a quantization step sizes allow computationally low complex algorithms for the reconstruction of transform coefficients at the decoder side. This is illustrated based on the preferred example of Fig. 1 in the following. The first quantization set includes all even integer multiples of the quantization step size and the second quantization set includes all odd integer multiples of the quantization step size plus the reconstruction level equal to 0 (which is contained in both quantization sets). The reconstruction process for a transform coefficient can be implemented similar to the algorithm specified in the pseudo-code of Fig. 2.
[0025] Fig. 2 shows a Pseudo-code illustrating the reconstruction process for transform coefficients, k represents an index that specifies the reconstruction order of the current transform coefficient, the quantization index for the current transform coefficient is denoted by level[k], the quantization step size 1kthat applies to the current transform coefficient is denoted by quant_step_size[k], and trec[k] represents the value of the reconstructed transform coefficient tk. The variable setld[k] specifies the set of reconstruction levels that applies to the current transform coefficient. It is determined based on the preceding transform coefficients in reconstruction order; the possible values of setld[k] are 0 and 1. The variable n specifies the integer factor of the quantization step size; it is given by the chosen set of reconstruction levels (i.e., the value of setld[k]) and the transmitted quantization index level[k].
[0026] In the pseudo-code of Fig. 2, level[k] denotes the quantization index that is transmitted for a transform coefficient tkand setld[k] (being equal to 0 or 1) specifies the identifier of the current set of reconstruction levels (it is determined based on preceding quantization indexes in reconstruction order as will be described in more detail below). The variable n represents the integer multiple of the quantization step size given by the quantization index level[k] and the set identifier setld[k]. If the transform coefficient is coded using the first set of reconstruction levels (setld[k] == 0), which contains the even integer multiples of the quantization step size Afe, the variable n is two times the transmitted quantization index. If the transform coefficient is coded using the second set of reconstruction levels (setld[k] == 1), we have the following three cases: (a) if level[k] is equal to 0, n is also equal to 0; (b) if level[k] is greater than 0, n is equal to two times the quantization index level[k] minus 1; and (c) if level[k] is less than 0, n is equal to two times the quantization index level[k] plus 1. This can be specified using the sign function
[0027] M : x > 0 sign(x) = 0 : x = 0 .
[0028] (-1 : x < 0
[0029] Then, if the second quantization set is used, the variable n is equal to two times the quantization index level[k] minus the sign function sign(level[k]) of the quantization index.
[0030] Once the variable n (specifying the integer factor of the quantization step size) is determined, the reconstructed transform coefficient tkis obtained by multiplying n with the quantization step size Afe.
[0031] Fig. 3 shows a Pseudo-code illustrating an alternative implementation of the pseudo-code in Fig. 2. The main change is that the multiplication with the quantization step is represented using an integer implementation using a scale and a shift parameter. Typically, the shift parameter (represented by shift) is constant for a transform block and only the scale parameter (given by scale[k]) may depend on the location of the transform coefficient. The variable add represents a rounding offset, it is typically set equal to add = (1«(shift-1)). With Akbeing the nominal quantization step for the transform coefficient, the parameters shift and scale[k] are chosen in a way that we have Ak« scale [k] • 2~shl^t. As mentioned above, instead of an exact multiplication with the quantization step size Ak, the reconstructed transform coefficient can be obtained by an integer approximation. This is illustrated in the pseudo-code in Fig. 3. Here, the variable shift represents a bit shift to the right. Its value typically depends only on the quantization parameter for the block (but it is also possible that the shift parameter can be changed for different transform coefficients inside a block). The variable scale[k] represents a scaling factor for the transform coefficient tk; in addition to the block quantization parameter, it can, for example, depend on the corresponding entry of the quantization weighting matrix. The variable add specifies a rounding offset, it is typically set equal to add = (1 «(shift-1)). It should be noted that the integer arithmetic specified in the pseudo-code of Fig. 3 (last line) is, with exception of the rounding, equivalent to a multiplication with a quantization step size Ak, given by
[0032] Ak= scale [Zc] • 2-shlft.
[0033] Another (purely cosmetic) change in Fig. 3 relative to Fig. 2 is that the switch between the two sets of reconstruction levels is implemented using the ternary if-then-else operator ( a ? b : c ), which is known from programming languages such as the C programming language.
[0034] In the following, dependent Reconstruction of Transform Coefficients is considered. In particular, another important design aspect of dependent scalar quantization is the algorithm used for switching between the defined quantization sets (sets of reconstruction levels). The used algorithm determines the “packing density” that can be achieved in the / V-dimensional space of transform coefficients (and, thus, also in the / V-dimensional space of reconstructed samples). A higher packing density eventually results in an increased coding efficiency.
[0035] In VTM-7, the transition between the quantization sets (set 0 and set 1) is determined by a state variable (or quantization state). For the first transform coefficient in reconstruction order, the state variable is set equal to a pre-defined value. Typically, the pre-defined value is equal to 0. The state variable for the following transform coefficients in coding order are determined by an update process. The state for a particular transform coefficient only depends on the state for the previous transform coefficient in reconstruction order and the value of the previous transform coefficient.
[0036] The state variable has four possible values (0, 1 , 2, 3). On the one hand, the state variable specifies the quantization set that is used for the current transform coefficient. The quantization set 0 is used if and only if the state variable is equal to 0 or 1 , and the quantization set 1 is used if and only if the state variable is equal to 2 or 3. On the other hand, the state variable also specifies the possible transitions between the quantization sets.
[0037] The state for a particular transform coefficient only depends on the state for the previous transform coefficient in reconstruction order and a binary function of the value of the previous transform coefficient. The binary function is referred to as path in the following. In VTM-7, the following state transition table is used, where “path” refers to the said binary function of the previous transform coefficient level in reconstruction order.
[0038] Table 1: State transition table in VTM-6.
[0039] In VTM-7, the path is given by the parity of the quantization index. With level[k] being the transform coefficient level, it can be determined according to path = ( level [ k ] & 1 ), where the operator & represents a bit-wise “and” in two-complement integer arithmetic.
[0040] As an alternative, the path could also represent other binary functions of level[ k]. As an example, it can specify whether a transform coefficient level is equal or not equal to 0:
[0041] . fO level[
[0042] Pat(1 : level[
[0043] The concept of state transition for the dependent scalar quantization allows low- complexity implementations for the reconstruction of transform coefficients in a decoder. A preferred example for the reconstruction process of transform coefficients of a single transform block is shown in Fig. 4 using C-style pseudo-code. Fig. 4 shows a Pseudo-code illustrating the reconstruction process of transform coefficients for a transform block. The array level represents the transmitted transform coefficient levels (quantization indexes) for the transform block and the array tree represent the corresponding reconstructed transform coefficients. The 2d table state_trans_table specifies the state transition table and the table setld specifies the quantization set that is associated with the states. The function path() specifies a binary function of the transform coefficient level.
[0044] In the pseudo-code of Fig. 4, the index k specifies the reconstruction order of transform coefficients. It should be noted that, in the example code, the index k decreases in reconstruction order. The last transform coefficient has the index equal to k = 0. The first index kstart specifies the reconstruction index (or, more accurately, the inverse reconstruction index) of the first reconstructed transform coefficient. The variable kstart may be set equal to the number of transform coefficients in the transform block minus 1, or it may be set equal to the index of the first non-zero quantization index (for example, if the location of the first non-zero quantization index is transmitted in the applied entropy coding method) in coding / reconstruction order. In the latter case, all preceding transform coefficients (with indexes k > kstart) are inferred to be equal to 0. The reconstruction process for each single transform coefficient is the same as in the example of Fig. 3. As for the example in Fig. 3, the quantization indexes are represented by level[k] and the associated reconstructed transform are represented by trec[k]. The state variable is represented by state. The 1d table setld[] specifies the quantization sets that are associated with the different values of the state variable and the 2d table state_trans_table[][] specifies the state transition given the current state (first argument) and the path (second argument). As an example, the path could be given by the parity of the quantization index (using the bit-wise and operator &), but other concepts are possible. As a further example, the path could specify whether the transform coefficient is equal or unequal to zero. Examples, in C-style syntax, for the tables are given in Fig. 5 (these tables are identical to Table 1).
[0045] Fig. 5 shows a state transition table state_trans_table and the table setld, which specifies the quantization set associated with the states. The table given in C-style syntax represents the tables specified in Table 1.
[0046] Instead of using a table state_trans_table[][] for determining the next state, an arithmetic operation yielding the same result can be used. Similarly, the table setld[] could also be implemented using an arithmetic operation. Or the combination of the table look-up using the 1d table setld[] and the sign function could be implemented using an arithmetic operation.
[0047] In VVC, the level coding process begins with signaling whether the block contains nonzero levels. If non-zero levels are present, the position of the last significant scanning position is encoded. With this position established, a reverse scanning pattern is employed to encode the levels along the scanning path.
[0048] The processing of a block is divided into sub-blocks. For example, a 16x16 block is partitioned into four 4x4 sub-blocks. Each sub-block is fully processed, meaning that all level information is encoded for a sub-block before moving on to the next sub-block in reversed scanning order.
[0049] When conducting the coding of levels within a sub-block, for each scanning position (except the last position of the regular scanning pattern, which is the starting position), it is first signaled whether the position contains a non-zero level (sig). If a non-zero level is present, it is then signaled whether its absolute value is greater than one (gtl). If the value of gtl is equal to one, the parity of the level (par) and the greater than three flag (gt3) are coded. This process is applied to all scanning positions within the sub-block initially. In the second step, the remaining level information (rem) is coded when the already coded information for that scanning position indicates that the level is greater than three. In the final coding step, the signs of the non-zero scanning positions are transmitted.
[0050] For example, when encoding a number of absolute values of quantization indices, which represent a coding of levels, the bin sig may, e.g., indicate whether the encoded absolute value is 0 (e.g., sig=Q) or whether the absolute value is greater than 0 (e.g., sig=Q). The bin gtl may, e.g., indicate, whether the encoded absolute value is 0 or 1 (e.g., gtl=O) or whether the absolute value is greater than 1 (e.g., gtl=V). The parity bin par may, e.g., indicate for all absolute values greater than 1, whether an even absolute value (e.g., par=0) or an odd absolute value (e.g., / w=l ) is encoded. Moreover, the bin gt3 may, e.g., indicate, whether the encoded absolute value is between 0 and 3 (e.g., gl3=Q) or whether the absolute value is greater than 3 (e.g., r3=l).This allows to uniquely encode integer values between 0 and 3.
[0051] Higher values may, e.g., be encoded by employing, e.g., a non-binary remainder rem. which may, e.g., indicate a remainder for encoding higher integer value (for example, gt3= ,par=0, rem=0 may e.g., encode 4; gt3= ,par= , rem=0 may e.g., encode 5; gt3= , par=<P rem=l may e.g., encode 6; gt3=l, par=\, rem=l may e.g., encode 7; gt3=\, par=Q, rem=2 may e.g., encode 8; etc.).
[0052] In more detail, regarding entropy coding of transform coefficient levels, an important aspect of dependent scalar quantization is that there are different sets of admissible reconstruction levels (also called quantization sets) for the transform coefficients. The quantization set for a current transform coefficient is determined based on the values of the quantization index for preceding transform coefficients. If we consider the preferred example in Fig. 1 and compare the two quantization sets, it is obvious that the distance between the reconstruction level equal to zero and the neighboring reconstruction levels is larger in set 0 than in set 1. Hence, the probability that a quantization index is equal to 0 is larger if set 0 is used and it is smaller if set 1 is used. In VTM-7, this effect is exploited in the entropy coding by switching probability models based on the states that are used for a current quantization index.
[0053] Note that for a suitable switching of codeword tables or probability models, the path (binary function of the quantization index) of all preceding quantization indexes must be known when entropy decoding a current quantization index (or a corresponding binary decision of a current quantization index).
[0054] In VTM-7, the quantization indexes are coded using binary arithmetic coding similar to H.264 | MPEG-4 AVC or H.265 | MPEG-H HEVC. For that purpose, the non-binary quantization indexes are first mapped onto a series of binary decisions (which are commonly referred to as bins).
[0055] Regarding binarization, the quantization indexes are transmitted as absolute value and, for absolute values greater than 0, a sign. While the sign is transmitted as single bin, there are many possibilities for mapping the absolute values onto a series of binary decisions.
[0056] Table 2: Binarization of absolute values |q| of transform coefficient levels in VTM-7.
[0057] The binarization of absolute values used in VTM-7 is illustrated in Table 2. The following binary and non-binary syntax elements are transmitted:
[0058] • sig_flag specifies whether the absolute value |q| of the transform coefficient level is greater than 0;
[0059] • if sig_flag is equal to 1, gt1_flag specifies whether the absolute value |q| of the transform coefficient level is greater than 1;
[0060] • if gt1_f lag is equal to 1 , par_flag specifies the parity of the absolute value |q| of the transform coefficient level, and gt3_flag specifies whether the absolute value |q| of the transform coefficient level is greater than 3;
[0061] • if gt3_flag is equal to 1, the non-binary value rem specifies the remainder of the absolute level |q|. This syntax element is transmitted in bypass mode of the arithmetic coding engine, using a Golomb-Rice code. Not present syntax elements are inferred to be equal to 0. At the decoder side the absolute value of the transform coefficient levels is reconstructed as follows:
[0062] |q| = sig_flag + gt1_flag + par_flag + 2 * ( gt3_flag + rem ) For non-zero transform coefficient levels (indicated by sig_flag equal to 1), a sign_flag specifying the sign of the transform coefficient level is additionally transmitted in bypass mode.
[0063] It has been proven useful that, when limiting the number of context-coded bins, in the first coding phase a limitation of the number of context-coded bins to four bins is employed. Specifically, all bins of the first coding phase are encoded with adaptive context models, while the remainder (rem) and the signs are encoded with a fixed probability of 1 / 2, e.g., without adaptive context models. This approach, known as bypass-mode, is less complex than the variant with adaptive context-coded bins. Therefore, to enhance throughput in practical implementations, the number of context-coded bins is restricted as follows:
[0064] Without further limitations, the number of context-coded bins per sample equals four, or 4WH, with W being the width, and H being the height of the block. By introducing a budget that is lower than 4WH, the number of context-coded bins per sample can be decreased without scarifying coding performance as follows: After the coding of a context- coded bin, the budget is decreased by one. The first coding phase only starts for a subblock when the remaining budget is greater than or equal to four. For all remaining scanning positions of the block that are not processed by the first coding phase, the remainder dec represents the absolute value rather than the remaining level value. The absolute values (dec) are coded similarly to rem, e.g., using adaptive Rice codes.
[0065] In the following, context modeling and adaptive binarization are discussed. For each context-coded bin, the context modeling selects an adaptive context model. Due to their distinct statistical properties, different sets of context models are utilized for the four types of context-coded flags (sig, gtl, par, gt3). Additionally, separate context models are employed for blocks of the luma and chroma components. The context modeling process employs a local template, e.g., comprising five previously processed neighboring locations (right, right-right, bottom, bottom-bottom, and bottom-right, accessed by j in the following). A partial absolute sum is calculated and mapped to a context model (or context model offset to provide a more accurate description).
[0066] Let ?! denote the partial sum calculated in the first coding phase according to:
[0067] The context index offset csigfor szg may, e.g., be calculated according to: csig= min(3, T + 1) » 1).
[0068] For the remaining context-coded flags, the calculation of the context index offset cccmay, e.g., be calculated according to: ccc= min(4, T — ri), where n is the number of non-zero valued neighboring locations specified by the local template.
[0069] The formulas above describe the selection of a context model within a given set. In addition to the previously mentioned description of different context model sets, a further distinction may, e.g., be made based on the diagonal of the current scanning position.
[0070] The above is now explained in more detail.
[0071] Regarding the coding order of bins and bypass mode, the entropy of transform coefficient levels in VTM-7 shares several aspects with that of HEVC. But it also includes additional aspects by which the entropy coding is adapted to dependent quantization.
[0072] Fig. 6 shows a Signaling of the position of the first non-zero quantization index 120 in coding order (highlighted sample). In addition to the position of the first non-zero transform coefficient 120, only bins for the coefficients 122 succeeding the first non-zero quantization index 120 in coding order 102 are transmitted (e.g., blue-marked samples), the coefficients preceding the first non-zero quantization index 120 in coding order 102 are inferred to be equal to 0 (e.g., white-marked samples).
[0073] Each block of transform coefficients is partitioned into fixed-size subblocks (also called coefficient groups). In general, the subblocks have a size of 4x4 coefficients (see Fig. 6). In some cases, other sizes may be used (e.g., 2x2, 2x4, 4x2, 2x8, or 8x2).
[0074] Similar to HEVC, the coding of a block of transform coefficients proceeds as follows:
[0075] • First, a so-called coded_block_flag is transmitted which specifies whether there are any non-zero transform coefficient levels in the transform block. If coded_block_flag is equal to 0, all transform coefficient levels are equal to 0, and no further data is transmitted for the transform block. In some cases (e.g. in the Skip mode), the coded_block_flag may not be explicitly transmitted, but is inferred based on other syntax elements.
[0076] • If coded_block_flag is equal to 1 , the x and y coordinate of the first significant transform coefficient in coding order is transmitted (this is sometimes referred as last significant coefficient, since the actual scanning order specifies a scanning from high-frequency to low frequency components);
[0077] • The scanning of transform coefficients proceeds on the basis of subblocks (typically, 4x4 subblocks); all transform coefficient levels of a subblock are coded before any transform coefficient level of any other subblock is coded; the subblocks are processed in a pre-defined scanning order, starting with the subblock that contains the first significant coefficient in scanning order and ending with the subblock that contains the DC coefficient;
[0078] • The syntax includes a coded_subblock_flag, which indicates whether the subblock contains any non-zero transform coefficient levels; for the first and last subblock in scanning order (i.e., the subblock that contains the first significant coefficient in scanning order and the subblock that contains the DC coefficient), this flag is not transmitted, but inferred to be equal to 1 ; when the coded_subblock_flag is transmitted, it is typically transmitted at the start of a subblock;
[0079] • For each subblock with coded_subblock_flag equal to 1, the transform coefficient levels are transmitted in multiple scan passes over the scanning positions of a subblock.
[0080] In the following, we will initially neglect a special aspect (which will be described later). Without that special aspect, the coding of subblocks with coded_subblock_flag equal to 1 proceeds as follows:
[0081] • For the first subblock in coding order (i.e., the subblock that includes the first significant scanning position, for which the x and y coordinate are explicitly transmitted), the coding starts at the scan position firstSigScanldx of the first nonzero coefficient in scanning order (i.e., the scan position that corresponds to the explicitly transmitted x and y coordinate). For all other subblocks, the coding starts at the minimum scan index minSubblockScanldx inside the subblock. The scan ends at the maximum scan index maxSubblockScanldx inside the subblock. • In the first pass, the regular coded bins sig_flag, gt1_flag, par_flag, and gt3_flag are transmitted: o The sig_flag is not transmitted if it can be inferred to be equal to 1. It can be inferred to be equal to zero in the following two cases:
[0082] (1) The current scan position is equal to the scan index startScanldx of the first non-zero coefficient in scanning order (i.e., the scan position that corresponds to the explicitly transmitted x and y coordinate);
[0083] (2) The current subblock is a subblock for which a coded_subblock_flag equal to 1 was transmitted, the current scan index is the maximum scan index maxSubblockScanldx inside the subblock, and all previously transmitted sig_flag’s for the current subblock are equal to 1. o If sig_flag[k] for a scan index k is equal to 1 (transmitted or inferred), a gt1_flag is transmitted. Otherwise, gt1_flag is inferred to be equal to zero.
[0084] ° If gt1_flag[k] f°r a scanindex k is equal to 1 , a par_flag and a gt3_flag are transmitted. Otherwise, the par_flag and the gt3_flag are inferred to be equal to zero.
[0085] • In the second pass, the remainder values rem are coded for those scan indexes, for which gt3_flag equal to 1 was transmitted. For all other scan indexes, the remainder rem is inferred to be equal to 0. The remainder values rem are coded using a concatenation of a Rice code and exponential Golomb code, which is parameterized by a so-called Rice parameter. The bins of the corresponding codewords are coded in bypass mode. The Rice parameter is determined based on already coded syntax elements (see below).
[0086] • Finally, in the last pass, for all scan indexes k with sig_flag[k] equal to 1 (coded or inferred), a sign_flag is transmitted, which specifies whether the transform coefficient value is negative or positive.
[0087] The following pseudo code further illustrates the coding process for the transform coefficient levels inside a subblock. It is illustrated from a decoder perspective. The Boolean variable firstSubblock specifies whether the current subblock is the first subblock in coding order (inside the transform block). The firstSigScanldx specifies the scan index that corresponds to the position of the first significant transform coefficient in the transform block (the one that is explicitly signaled at the start of the transform block syntax). minSubblockScanldx and maxSubblockScanldx represent the minimum and maximum values of the scan indexes for the current subblock. Note that the first scan index for which any data is transmitted depends on whether the subblock includes the first significant coefficient in scanning order. If this is the case, the first scan pass starts at the scan index that corresponds to the first significant transform coefficient; otherwise, the first scan pass starts at the minimum scan index of the subblock. The variable coeff[k] for a scan index k represents the reconstructed transform coefficient level. The encoding proceeds analog to the decoding.
[0088] Pseudo-code specifying the decoding of subblock with coded subblock flag egual to 1 startScanldx = ( firstSubblock ? firstSigScanldx : minSubblockScanldx ) endScanldx = maxSubblockScanldx
[0089] / / first pass (regular coded bins) for( k = startScanldx; k <= endScanldx; k++ ) { coeff[k] = 0 if( sig_flag cannot be inferred to be egual to 1 ) { decode sig_flag[k]
[0090] } if( sig_flag[k] != 0 ) { decode gt1_flag[k] if( gt1_flag[k] != 0 ) { decode par_flag[k] decode gt3_flag[k]
[0091] } coeff[k] = 1 + gt1_flag[k] + par_flag[k] + 2 * gt3_flag[k]
[0092] } state = stateT ransT able[ state ][ coeff[k] & 1 ]
[0093] }
[0094] / / second pass (bypass coding of remainder) for( k = startScanldx; k <= endScanldx; k++ ) { if( gt3_flag[k] != 0 ) { decode rem[k] coeff[k] = coeff[k] + 2 * rem[k] }
[0095] }
[0096] / / third pass (bypass coding of signs) for( k = startScanldx; k <= endScanldx; k++ ) { if( coeff[k] != 0 ) { decode sign_flag[k] if( sign_flag[k] == 1 ) coeff[k] = -coeff[k] }
[0097] }
[0098] One disadvantage relative to HEVC was that the maximum number of regular coded bins per transform coefficient is increased. In order to circumvent this issue, the following concept is included in VTM-7:
[0099] • The maximum number of regular coded bins maxNumRegBins for a transform block is set equal to maxNumRegBins = 1.75 * width * height, where width and height represent the size of the transform block (more accurately, of the non-zero out region of the transform block, in cases in which transform coefficients are forced to be equal to zero).
[0100] • At the start of the transform block, a counter remRegBins, which specifies the available number of regular coded bins is set equal to remRegBins = maxNumRegBins.
[0101] • After a regular coded bin is encoded or decoded, the counter remRegBins is decreased by one.
[0102] • If, at the beginning of the coding for a scan position in the first pass, the counter is less than 4 (in which case, not all bins sig_flag, gt1_flag, par_flag and gt3_flag could be transmitted without exceeding the maximum number of regular coded bins), the first pass is terminated. In addition, the remainder values are only transmitted for such scan positions that were included in the first pass.
[0103] The absolute levels for which no data is coded in the first scan pass are completely coded in bypass mode. They are coded using the same class of codes as for the remainder rem. The Rice parameter as well as a variable posO are determined based on already coded syntax elements (see below). These absolute levels |q| are not directly coded, but they are first mapped to syntax elements absjevel, which are then coded using the parametric codes.
[0104] The mapping of an absolute value |q| to the syntax element absjevel depends on a variable posO. It is specified by absjevel = ( |q| == 0 ? posO : ( |q|<= posO ? |q|-1 : |q| )
[0105] At the decoder side, the mapping of the syntax element absjevel to the absolute value |q| is specified as follows
[0106] |q| = ( absjevel == posO ? 0 : ( absjevel < posO ? absjevel + 1 : absjevel )
[0107] It should be noted that limit on the number of regular coded bins is applied on transform block basis and not on a subblock basis. That means, the counter remRegBins is initialized at the beginning of a transform block and decreased for each regular coded bin. If the coding switches to the bypass mode in a particular subblock, all transform coefficient levels of all following subblocks in coding order are also coded in bypass mode.
[0108] The coding process for a subblock (with coded_subblockjlag equal to 1) is illustrated in the following pseudo code. As in the previous pseudo-code, the coding process is illustrated from a decoder perspective. For the first coded subblock, the initial value of the counter remRegBins is equal to maxNumRegBins. For all following subblocks, the initial value of the counter remRegBins is equal to the value that was obtained at the end of the previous coded subblock. The scan index startldxBypass is the first scan index at which the bypass coding (i.e. , the coding of absjevel) starts inside the subblock (if applicable).
[0109] Pseudo-code (114) specifying the decoding of subblock with coded subblock flag equal to 1 (with bypass mode): startScanldx = (firstSubblock ? firstSigScanldx : minSubblockScanldx) endScanldx = maxSubblockScanldx startldxBypass = (remRegBins >= 4 ? maxSubblockScanldx + 1 : startScanldx)
[0110] / / first pass (regular coded bins) for( k = startScanldx; k <= endScanldx && remRegBins >= 4; k++ ) { coeff[k] = 0 if( sigjlag cannot be inferred to be equal to 1 ) { decode sig_flag[k] remRegBins = remRegBins - 1
[0111] } if( sig_flag[k] != 0 ) { decode gt1_flag[k] remRegBins = remRegBins - 1 if( gt1_flag[k] != 0 ) { decode par_flag[k] decode gt3_flag[k] remRegBins = remRegBins - 2
[0112] } coeff[k] = 1 + gt1_flag[k] + par_flag[k] + 2 * gt3_flag[k]
[0113] } if( remRegBins < 4 ) { startldxBypass = k + 1
[0114] } state = stateTransTable[ state ][ coeff[k] & 1 ]
[0115] }
[0116] / / second pass (bypass coding of remainder) for( k = startScanldx; k < startldxBypass; k++ ) { if( gt3_flag[k] != 0 ) { decode rem[k] coeff[k] = coeff[k] + 2 * rem[k]
[0117] }
[0118] }
[0119] / / third pass (bypass coding of absolute levels) for( k = startldxBypass; k <= endScanldx; k++ ) { derive posO decode abs_level[k] if( abs_level[k] == posO ) coeff[k] = 0 else if ( abs_level[k] < posO) coeff[k] = abs_level[k] + 1 else coeff[k] = abs_level[k] state = stateTransTable[ state ][ coeff[k] & 1 ]
[0120] } / / fourth pass (bypass coding of signs) for( k = startScanldx; k <= endScanldx; k++ ) { if( coeff[k] != 0 ) { decode sign_flag[k] if( sign_flag[k] == 1 ) coeff[k] = -coeff[k]
[0121] }
[0122] }
[0123] In the following, context modelling is described:
[0124] For the regular coded bins sig_flag, gt1_flag, par_flag, and gt3_flag one of multiple probability models (or contexts) is selected for the actual coding.
[0125] Fig. 7 shows a local template used for selecting probability models for one or more bins. The black square 130 represents the current scan position and the highlighted samples 132 (e.g., the blue squares) represent the neighboring scan positions inside the template.
[0126] Significance flag sig_flag
[0127] For the significance flag, the selected probability model depends on the following:
[0128] • whether the current transform block is a luma or chroma transform block;
[0129] • the state for dependent quantization;
[0130] • the x and y coordinate of the current transform coefficient;
[0131] • the partially reconstructed absolute values (after the first pass) in a local neighborhood.
[0132] In the following, more details are given.
[0133] The variable state represents the state for a current transform coefficient used in dependent quantization. The state can take the values 0, 1, 2, or 3. Initially, the state is set equal to 0. as explained above, the state for a current transform coefficient is given by the state for the previous coefficient in coding order and the parity (or, more general, a binary function) of the previous transform coefficient level. With x and y being the coordinates of a current transform coefficient inside the transform block, let diag = x + y be diagonal position of the current coefficient. Given the diagonal position, a diagonal class index dsig is derived as follows: dsig = ( diag < 2 ? 2 : ( diag < 5 ? 1 : 0 ) ) for luma transform blocks and dsig = ( diag < 2 ? 1 : 0 ) for chroma transform blocks
[0134] The ternary operator ( c ? a : b ) represents an if-then-else statement. If the condition c is true, the value of a is used, otherwise (c is false), the value b is used.
[0135] The context also depends on the partially reconstructed absolute values inside a local neighborhood. In VTM-7, the local neighborhood is given by the template T shown in Fig. 7. But other templates are possible. The template used could also depend on whether a luma or a chroma block is coded. Let sumAbs be the sum of partially reconstructed absolute values (after the first pass) in the template T: sumAbs absl[ / c], where abs1[k] represents the partially reconstructed absolute level for a scan index k after the first pass. With the binarization of VTM-7, it is given by abs1[k] = sig_flag[k] + gt1_flag[k] + par_flag[k] + 2 * gt3_flag[k].
[0136] Note that abs1[k] is equal to the value coeff[k] obtained after the first scan pass (see above).
[0137] Let us assume that the possible probability models for the sig_flag are organized in a 1d array. And let ctxIdSig be an index that identifies the probability model used. According to VTM-7, the context index ctxIdSig could be derived as follows:
[0138] If the transform block is a luma block, ctxIdSig = min( (sumAbs+1)»1 , 3 ) + 4 * dsig + 12 * min( state-1 , 0 )
[0139] If the transform block is a chroma block, ctxIdSig = 36 + min( (sumAbs+1)»1 , 3 ) + 4 * dsig + 8 * min( state-1 , 0 ) Here, the operator “»” represents a bit shift to the right (in two-complement arithmetic). The operation is identical to a division by two and a downrounding (to the next integer) of the result.
[0140] Note that other organizations of the context models are possible. But in any case, the probability model chosen for coding a sig_flag depends on
[0141] • whether a luma or a chroma block is coded;
[0142] • the state variable, specifically min( state-1, 0 );
[0143] • the diagonal class dsig;
[0144] • the sum of partially reconstructed absolute values (after the first pass) inside the local template, specifically min( (sumAbs+1)»1, 3 ).
[0145] Flags gt1_flag, par_flag, and gt3_flag
[0146] For the flags gt1_flag, par_flag, and gt3_flag, the selected probability model depends on the following:
[0147] • whether the current transform block is a luma or chroma transform block;
[0148] • whether the current transform coefficient is the first non-zero coefficient in coding order inside the transform block;
[0149] • the x and y coordinate of the current transform coefficient;
[0150] • the partially reconstructed absolute values (after the first pass) in a local neighborhood.
[0151] Let firstCoeff be a variable that indicates whether the current scan index represents the scan index of the first non-zero transform coefficient level in coding order (i.e., the scan index or which the x and y location are explicitly coded). If the current scan index is equal to the scan index of the first non-zero transform coefficient level, firstCoeff is equal to 1; otherwise, firstCoeff is equal to 0.
[0152] With x and y being the coordinates of a current transform coefficient inside the transform block, let diag = x + y be diagonal position of the current coefficient. Given the diagonal position, a diagonal class index dclass is derived as follows: dclass = ( diag == 0 ? 3 : ( diag < 3 ? 2 : ( diag < 10 ? 1 : 0 ) ) ) for luma transform blocks and dclass = ( diag == 0 ? 1 : 0 ) for chroma transform blocks The context also depends on the partially reconstructed absolute values inside a local neighborhood. The variable sumAbs is derived as specified above. In addition, a variable numSig is derived, which specifies the number of non-zero transform coefficient levels inside the template T: numSig = ZfeeTsig_flag[ / c].
[0153] Similar as for the sig_flag, let us assume that the possible probability models are organized in a 1d array. And let ctxld be an index that identifies the probability model used. According to VTM-7, the context index ctxld could be derived as follows:
[0154] If the transform block is a luma block, ctxld = ( firstCoeff ? 0 : 1 + min( sumAbs-numSig, 4 ) + 5 * dclass )
[0155] If the transform block is a chroma block, ctxld = 21 + ( firstCoeff ? 0 : 1 + min( sumAbs-numSig, 4 ) + 5 * dclass )
[0156] Note that other organizations of the context models are possible. But in any case, the probability model chosen depends on
[0157] • whether a luma or a chroma block is coded;
[0158] • whether the current scan index is the scan index of the first non-zero transform coefficient level in coding order;
[0159] • the diagonal class dclass;
[0160] • the difference of the sum of partially reconstructed absolute values (after the first pass) inside the local template and the number of non-zero transform coefficient levels inside the local template, specifically min( sumAbs-numSig, 4 ).
[0161] It should be noted that the same context index ctxld is used for gt1_flag, par_flag, and gt3_flag. But for each of these flags, a different set of context models is used.
[0162] Rice parameter for remainder and bypass-coded transform coefficient levels
[0163] Just as with context modeling for context-coded bins, the identical local template is utilized to determine the Rice parameter for the adaptive binarization process. But in contrast to the first coding phase, the absolute values of the neighboring locations may, e.g., be reconstructed, and thus, available for the determination of the Rice parameter.
[0164] Let |x| denote the absolute level, and let T2denote the absolute sum calculated in the second coding phase according to: Then, the Rice parameter k for the current scanning position may, e.g., be derived according to: k = max(0, min (31, T2— 5 ■ m)).
[0165] In the above equation, m = 4 when coding rem, and m = 0 when coding dec.
[0166] In more detail both the remainder rem and the absolute values absjevel are coded using a parametric class of codes, for which the bins are coded in bypass mode (see above). The binarization actually used is determined by a so-called Rice parameter RP. Furthermore, for the absolute levels absjevel, a variable posO determines the mapping from the coded syntax element absjevel to the actual absolute values.
[0167] The Rice parameter for the remainder RPrem is derived based on the sum of reconstructed absolute values in a local template. In VTM-7, the same template as specified above is used. With T representing the template and coeff[k] representing the reconstructed coefficients for a scan index k, the sum sumAbs is derived according to sumAbs abs( coeff[ / c] )
[0168] Note that the completely reconstructed transform coefficient values coeff[k] are used. Let tabRPrem[ ] be a fixed table of size 32. The Rice parameter RPrem is derived according to
[0169] RPrem = tabRPrem[ max( min( sumAbs - 20, 31 ), 0 ) ].
[0170] Let tabRPabs[ ] be another fixed table of size 32. The Rice parameter RPabs used for coding absjevel is derived according to
[0171] RPabs = tabRPabs[ min( sumAbs, 31 ) ].
[0172] The variable posO depends on both, the sum of reconstructed absolute values in the local template and the state variable. The variable posO is derived from the quantization state and the Rice parameter RPabs according to posO = ( state < 2 ? 1 : 2 ) « RPabs. High-Level Signaling of the Quantization Method
[0173] VTM-7 supports dependent quantization with 4 states and, additionally, conventional independent scalar quantization. The quantization method chosen by an encoder is indicated in the bitstream as follows.
[0174] The picture header includes a syntax element pic_dep_quant_enabled_flag, which is coded using a single bit and indicates whether dependent quantization is used or not used for a current picture. If pic_dep_quant_enabled_flag is equal to 0, conventional independent quantization is used for the current picture. If pic_dep_quant_enabled_flag is equal to 1, the specified variant of dependent quantization is used for the current picture.
[0175] The picture header syntax element is not always present in the picture header. Its presence is controlled by another syntax element that is coded in the picture parameter set. The picture parameter set includes a syntax element pps_dep_quant_enabled_idc, which is coded using a fixed-length code of 2 bits and has the following meaning:
[0176] • the value of 0 specifies that the pic_dep_quant_enabled_flag is present in the picture header;
[0177] • the value of 1 specifies that the pic_dep_quant_enabled_flag is not present in the picture header, but is inferred to be equal to 0 for all pictures that refer to the picture parameter set;
[0178] • the value of 2 specifies that the pic_dep_quant_enabled_flag is not present in the picture header, but is inferred to be equal to 1 for all pictures that refer to the picture parameter set;
[0179] • the value of 3 is reserved for future use.
[0180] The pps_dep_quant_enabled_idc in the picture parameter set is present of a picture parameter set flag constant_slice_header_params_enabled_flag is equal to 1. If the picture parameter set flag constant_slice_header_params_enabled_flag is equal to 0, the picture parameter set syntax element pps_dep_quant_enabled_idc is not present in the picture parameter set, but it is inferred to be equal to 0 (in which case, the pic_dep_quant_enabled_flag is coded in the picture header).
[0181] The object of the present invention is to provide improved concepts for probability estimation of transform coefficients for video coding. The object of the present invention is solved by the subject-matter of the independent claims. Particular embodiments are provided in the dependent claims. A video encoder according to an embodiment is provided. The video encoder is configured to encode a video into a video data stream. Moreover, the video encoder is configured to generate the video data stream by encoding level information of each of a plurality of transform coefficients. Furthermore, the video encoder is configured to generate the video data stream such that the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0182] Moreover, a video encoder according to another embodiment is provided. The video encoder is configured to encode a video into a video data stream. Furthermore, the video encoder is configured to generate the video data stream by encoding level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0183] Moreover, a video encoder according to a further embodiment is provided. The video encoder is configured to encode a video into a video data stream. Moreover, the video encoder is configured to generate the video data stream by encoding level information of a plurality of transform coefficients. Furthermore, the video encoder is configured to select a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template. Moreover, the video encoder is configured to encode the level information using the context model and / or the Rice parameter being selected.
[0184] Furthermore, a video decoder for receiving a video data stream having a video encoded therein according to an embodiment is provided. The video decoder is configured to decode the video from the video data stream. To decode the video, the video decoder is configured to determine level information of a plurality of transform coefficients. The video decoder is configured to determine level information of a current transform coefficient of the plurality of transform coefficients depending on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available. Moreover, a video decoder for receiving a video data stream having a video encoded therein according to another embodiment is provided. The video decoder is configured to decode the video from the video data stream. To decode the video, the video decoder is configured to determine level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased- precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0185] Furthermore, a video decoder for receiving a video data stream having a video encoded therein according to another embodiment is provided. The video decoder is configured to decode the video from the video data stream. To decode the video, the video decoder is configured to determine level information of a plurality of transform coefficients by selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, and by using the context model and / or the Rice parameter being selected for determining the level information.
[0186] Moreover, a video data stream according to an embodiment is provided. The video data stream comprises, encoded therein, level information of each of a plurality of transform coefficients. Furthermore, the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0187] Furthermore, a video data stream according to another embodiment is provided. The video data stream comprises, encoded therein, level information of a plurality of transform coefficients that have been encoded using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level. Moreover, a video data stream according to a further embodiment is provided. The video data stream comprises, encoded therein, level information of a plurality of transform coefficients, wherein the level information is encoded depending on a context model being selected and / or a Rice parameter being selected.
[0188] Furthermore, a method for video encoding according to an embodiment is provided. The method comprises encoding a video into a video data stream. Moreover, the method comprises generating the video data stream by encoding level information of each of a plurality of transform coefficients. Generating the video data stream is conducted such that the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0189] Moreover, a method for video encoding according to an embodiment is provided. The method comprises encoding a video into a video data stream. Furthermore, the method comprises generating the video data stream by encoding level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0190] Furthermore, a method for video encoding according to an embodiment is provided. The method comprises encoding a video into a video data stream. Moreover, the method comprises generating the video data stream by encoding level information of a plurality of transform coefficients. Furthermore, the method comprises selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template. Moreover, he method comprises encoding the level information using the context model and / or the Rice parameter being selected.
[0191] Moreover, a method for video decoding according to an embodiment is provided. The method comprises for receiving a video data stream having a video encoded therein. Furthermore, the method comprises decoding the video from the video data stream. To decode the video, the method comprises determining level information of a plurality of transform coefficients. Moreover, the method comprises determining level information of a current transform coefficient of the plurality of transform coefficients depending on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0192] Furthermore, a method for video decoding according to an embodiment is provided. The method comprises for receiving a video data stream having a video encoded therein. Moreover, he method comprises decoding the video from the video data stream. To decode the video, the method comprises determining level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0193] Moreover, a method for video decoding according to an embodiment is provided. The method comprises for receiving a video data stream having a video encoded therein. Furthermore, the method comprises decoding the video from the video data stream. To decode the video, the method comprises determining level information of a plurality of transform coefficients by selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, and by using the context model and / or the Rice parameter being selected for determining the level information.
[0194] Moreover, computer programs are provided, wherein each of the computer programs is configured to implement one of the above-described methods when being executed on a computer or signal processor.
[0195] In the following, embodiments of the present invention are described in more detail with reference to the figures, in which:
[0196] Fig. 1 illustrates two different quantizers.
[0197] Fig. 2 illustrates a Pseudo-code illustrating a reconstruction process for transform coefficients. Fig. 3 illustrates a Pseudo-code illustrating an alternative implementation of a reconstruction process for transform coefficients.
[0198] Fig. 4 illustrates a Pseudo-code illustrating a reconstruction process of transform coefficients for a transform block.
[0199] Fig. 5 illustrates a state transition table.
[0200] Fig. 6 illustrates coding orders.
[0201] Fig. 7 illustrates a local template useable for selecting probability models for one or more bins.
[0202] Fig. 8 illustrates a system according to an embodiment.
[0203] A video encoder 100, illustrated in Fig. 8, according to an embodiment is provided. The video encoder 100 is configured to encode a video into a video data stream. Moreover, the video encoder 100 is configured to generate the video data stream by encoding level information of each of a plurality of transform coefficients. Furthermore, the video encoder 100 is configured to generate the video data stream such that the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0204] According to an embodiment, if the neighboring location is located outside a block of a picture of the video, in which the current transform coefficient is located, the coding information on the neighboring location is not available.
[0205] In an embodiment, if the neighboring location has not been processed by a previous level coding step, the coding information on the neighboring location is not available.
[0206] According to an embodiment, if the neighboring location is a scanning position beyond a last significant scanning position, the coding information on the neighboring location is not available. In an embodiment, if the neighboring location lies within a sub-block comprises zerovalued levels only, the coding information on the neighboring location is not available, and the video encoder 100 is configured to generate the video data stream such that the video data stream comprises a coded sub-block flag indication indicating that the sub-block comprises the zero-valued levels only.
[0207] According to an embodiment, if the coding information on the neighboring location is available, the video encoder 100 is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a first context model. If the coding information on the neighboring location is not available, the video encoder 100 is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
[0208] In an embodiment, if the coding information on all locations of two or more neighboring locations of the current transform coefficient is available, the video encoder 100 is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a first context model. If the coding information on at least one location of two or more neighboring locations of the current transform coefficient is not available, the video encoder 100 is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
[0209] According to an embodiment, if the coding information on the neighboring location is not available, the video encoder 100 is configured to modify the context model.
[0210] In an embodiment, the context model depends on a number of unavailable neighboring locations of the current transform coefficient.
[0211] According to an embodiment, the video encoder 100 is configured to select the context model out of a set of two or more context models depending on a number of unavailable neighboring locations of the current transform coefficient.
[0212] In an embodiment, the video encoder 100 is configured to select the context model out of a set of two or more context models as a selected context model depending on a number of unavailable neighboring locations of the current transform coefficient. The video encoder 100 is configured to adjust the selected context model depending on the number of unavailable neighboring locations of the current transform coefficient.
[0213] According to an embodiment, the video encoder 100 is configured to select a template configuration out of two or more template configurations, wherein each of the two or more template configurations defines two or more neighboring locations which are taken into account for encoding the current transform coefficient. The video encoder 100 is configured to select the template configuration out of two or more template configurations depending on an availability of coding information on the two or more neighboring locations. Moreover, the video encoder 100 is configured to determine the encoding of the level information of the current transform coefficient depending on the template configuration that has been selected.
[0214] In an embodiment, the video encoder 100 is configured to select the template configuration out of two or more template configurations such that coding information on all of the two or more neighboring locations being defined by the template configuration is available.
[0215] According to an embodiment, the video encoder 100 is configured to select the template configuration out of two or more template configurations depending on an unavailability pattern defined by neighboring locations of the two or more neighboring locations for which coding information is unavailable.
[0216] In an embodiment, a number of a set comprising the two or more template configuration varies depending on a rule, wherein the rule depends on the number of unavailable locations within a local template.
[0217] According to an embodiment, the rule defines that determining a context model index is conducted in a first way, if coding information on all of one or more neighboring locations is available. The rule defines that determining the context model index is conducted in a second way, being different from the first way, if coding information on at least one of one or more neighboring locations is unavailable.
[0218] In an embodiment, the video encoder 100 is configured to dynamically adjust a template configuration, which defines two or more neighboring locations which are taken into account for encoding the current transform coefficient, depending on an availability of coding information on the two or more neighboring locations. The video encoder 100 is configured to determine the encoding of the level information of the current transform coefficient depending on the template configuration that has been dynamically adjusted.
[0219] According to an embodiment, the video encoder 100 is configured to determine the encoding of the level information of the current transform coefficient depending on the context model such that the context model depends on a diagonal position of the current transform coefficient, a coding state and a number of unavailable locations within a template around with respect to a position of the current transform coefficient.
[0220] Moreover, a video encoder 100 according to another embodiment is provided. The video encoder 100 is configured to encode a video into a video data stream. Furthermore, the video encoder 100 is configured to generate the video data stream by encoding level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0221] According to an embodiment, the video encoder 100 is configured to generate intermediate level information of a current transform coefficient from an unquantized level of the current transform coefficient. The video encoder 100 is configured to generate quantized level information on the current transform coefficient from the intermediate level information on the current transform coefficient, the intermediate level information having a higher precision than the quantized level information. Moreover, the video encoder 100 is configured to conduct context modelling depending on the intermediate level information.
[0222] In an embodiment, the intermediate level information on the current transform coefficient indicates a level of the current transform coefficient having twice the size or twice the precision than a level of the current transform coefficient being indicated by the quantized level information.
[0223] According to an embodiment, the video encoder 100 is configured to determine the increased-precision information using information on a state of the trellis-coded quantization. In an embodiment, the video encoder 100 is configured for encoding the level information of the one of the plurality of transform coefficients using the context modelling without modifying the context modelling by halving one or more template sums and by adding a rounding offset to the one or more template sums after halving.
[0224] According to an embodiment, the video encoder 100 is configured for encoding the level information of the one of the plurality of transform coefficients using the context modelling without modifying the context modelling by generating one or more template sums of two or more processed intermediate values that have been obtained by halving and adding a rounding offset to each of two more initial intermediate values.
[0225] In an embodiment, the video encoder 100 is configured for encoding the level information of the one of the plurality of transform coefficients using the context modelling by adjusting the context modeling depending on one or more template sums of intermediate levels, wherein the video encoder 100 is configured to modify the context modeling and a Rice parameter derivation.
[0226] According to an embodiment, the video encoder 100 is configured for encoding the level information of the one of the plurality of transform coefficients using a state of the trelliscoded quantization and a partial sum of one or more templates for context modeling.
[0227] Moreover, a video encoder 100 according to a further embodiment is provided. The video encoder 100 is configured to encode a video into a video data stream. Moreover, the video encoder 100 is configured to generate the video data stream by encoding level information of a plurality of transform coefficients. Furthermore, the video encoder 100 is configured to select a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template. Moreover, the video encoder 100 is configured to encode the level information using the context model and / or the Rice parameter being selected.
[0228] In an embodiment, the reconstructed samples are represented in a spatial domain.
[0229] According to an embodiment, the video encoder 100 is configured to select the context model and / or the Rice parameter depending on one or more residual samples in the neighboring area of the current position defined by the template.
[0230] In an embodiment, the template is a non-local template. According to an embodiment, the video encoder 100 is configured to determine a correlation between a last significant scanning position and a residual energy of neighboring samples for context modeling.
[0231] In an embodiment, the video encoder 100 is configured to encode the level information using the context model depending on a plurality of neighboring samples of a current position, wherein the plurality of neighboring samples depends on a prediction mode.
[0232] According to an embodiment, the video encoder 100 is configured to encode the level information depending on a context modeling function which depends on a prediction mode.
[0233] In an embodiment, a number of levels employing context modeling using neighboring samples is fixed and is limited to lowest frequency locations.
[0234] According to an embodiment, a number of levels employing context modeling using neighboring samples is fixed with an actual number for a region or a transform block being determined depending on a prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream.
[0235] In an embodiment, neighboring samples of a current position are employed for context modeling of directly quantized residuals.
[0236] According to an embodiment, the video encoder 100 is configured to infer a transform skip flag depending on neighboring samples.
[0237] In an embodiment, the video encoder 100 is configured to operate in a transform skip mode. The video encoder 100, operating in the transform skip mode, is configured to select the context model and / or the Rice parameter depending on the neighboring reconstructed samples in the neighboring area of the current position defined by the template. Moreover, the video encoder 100 is configured to encode the level information using the context model and / or the Rice parameter being selected.
[0238] According to an embodiment, the video encoder 100 is configured to employ a different residual error coding concept, if the video encoder 100 is operating in the transform skip mode, compared to when the video encoder 100 is not operating in the transform skip mode. Furthermore, a video decoder 200, illustrated in Fig. 8, for receiving a video data stream having a video encoded therein according to an embodiment is provided. The video decoder 200 is configured to decode the video from the video data stream. To decode the video, the video decoder 200 is configured to determine level information of a plurality of transform coefficients. The video decoder 200 is configured to determine level information of a current transform coefficient of the plurality of transform coefficients depending on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0239] According to an embodiment, if the neighboring location is located outside a block of a picture of the video, in which the current transform coefficient is located, the coding information on the neighboring location is not available.
[0240] In an embodiment, if the neighboring location has not been processed by a previous level coding step, the coding information on the neighboring location is not available.
[0241] According to an embodiment, if the neighboring location is a scanning position beyond a last significant scanning position, the coding information on the neighboring location is not available.
[0242] In an embodiment, if the neighboring location lies within a sub-block comprises zerovalued levels only, the coding information on the neighboring location is not available, and wherein, to decode the video, the video decoder 200 is configured to determine the level information of the current transform coefficient depending on a coded sub-block flag indication in the video data stream, which indicates that the sub-block comprises the zerovalued levels only.
[0243] According to an embodiment, if the coding information on the neighboring location is available, the video decoder 200 is configured to determine the level information of the current transform coefficient depending on a first context model. If the coding information on the neighboring location is not available, the video decoder 200 is configured to determine the level information of the current transform coefficient depending on a second context model being different from the first context model.
[0244] In an embodiment, if the coding information on all locations of two or more neighboring locations of the current transform coefficient is available, the video decoder 200 is configured to determine the level information of the current transform coefficient depending on a first context model. If the coding information on at least one location of two or more neighboring locations of the current transform coefficient is not available, the video decoder 200 is configured to determine the level information of the current transform coefficient depending on a second context model being different from the first context model.
[0245] According to an embodiment, if the coding information on the neighboring location is not available, the video decoder 200 is configured to modify the context model.
[0246] In an embodiment, the context model depends on a number of unavailable neighboring locations of the current transform coefficient.
[0247] According to an embodiment, the video decoder 200 is configured to select the context model out of a set of two or more context models depending on a number of unavailable neighboring locations of the current transform coefficient.
[0248] In an embodiment, the video decoder 200 is configured to select the context model out of a set of two or more context models as a selected context model depending on a number of unavailable neighboring locations of the current transform coefficient. The video decoder 200 is configured to adjust the selected context model depending on the number of unavailable neighboring locations of the current transform coefficient.
[0249] According to an embodiment, the video decoder 200 is configured to select a template configuration out of two or more template configurations, wherein each of the two or more template configurations defines two or more neighboring locations which are taken into account for encoding the current transform coefficient. The video decoder 200 is configured to select the template configuration out of two or more template configurations depending on an availability of coding information on the two or more neighboring locations. The video decoder 200 is configured to determine the level information of the current transform coefficient depending on the template configuration that has been selected.
[0250] In an embodiment, the video decoder 200 is configured to select the template configuration out of two or more template configurations such that coding information on all of the two or more neighboring locations being defined by the template configuration is available. According to an embodiment, the video decoder 200 is configured to select the template configuration out of two or more template configurations depending on an unavailability pattern defined by neighboring locations of the two or more neighboring locations for which coding information is unavailable.
[0251] In an embodiment, a number of a set comprising the two or more template configuration varies depending on a rule, wherein the rule depends on the number of unavailable locations within a local template.
[0252] According to an embodiment, the rule defines that determining a context model index is conducted in a first way, if coding information on all of one or more neighboring locations is available. The rule defines that determining the context model index is conducted in a second way, being different from the first way, if coding information on at least one of one or more neighboring locations is unavailable.
[0253] In an embodiment, the video decoder 200 is configured to dynamically adjust a template configuration, which defines two or more neighboring locations which are taken into account for encoding the current transform coefficient, depending on an availability of coding information on the two or more neighboring locations. The video decoder 200 is configured to determine the level information of the current transform coefficient depending on the template configuration that has been dynamically adjusted.
[0254] According to an embodiment, video decoder 200 is configured to determine the level information of the current transform coefficient depending on the context model such that the context model depends on a diagonal position of the current transform coefficient, a coding state and a number of unavailable locations within a template around with respect to a position of the current transform coefficient.
[0255] Moreover, a video decoder 200 for receiving a video data stream having a video encoded therein according to another embodiment is provided. The video decoder 200 is configured to decode the video from the video data stream. To decode the video, the video decoder 200 is configured to determine level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0256] According to an embodiment, the video decoder 200 is configured to generate intermediate level information of a current transform coefficient from an unquantized level of the current transform coefficient, wherein the intermediate level information has a higher precision than the quantized level information, wherein the video decoder 200 is configured to conduct context modelling depending on the intermediate level information.
[0257] In an embodiment, the intermediate level information on the current transform coefficient indicates a level of the current transform coefficient having twice the size or twice the precision than a level of the current transform coefficient being indicated by the quantized level information.
[0258] According to an embodiment, the video decoder 200 is configured to determine the increased-precision information using information on a state of the trellis-coded quantization.
[0259] In an embodiment, the video decoder 200 is configured to use the context modelling without modifying the context modelling by halving one or more template sums and by adding a rounding offset to the one or more template sums after halving.
[0260] According to an embodiment, the video decoder 200 is configured to use the context modelling without modifying the context modelling by generating one or more template sums of two or more processed intermediate values that have been obtained by halving and adding a rounding offset to each of two more initial intermediate values.
[0261] In an embodiment, the video decoder 200 is configured to use the context modelling by adjusting the context modeling depending on one or more template sums of intermediate levels, wherein the video decoder 200 is configured to modify the context modeling and a Rice parameter derivation.
[0262] According to an embodiment, the video decoder 200 is configured to use a state of the trellis-coded quantization and a partial sum of one or more templates for context modeling.
[0263] Furthermore, a video decoder 200 for receiving a video data stream having a video encoded therein according to another embodiment is provided. The video decoder 200 is configured to decode the video from the video data stream. To decode the video, the video decoder 200 is configured to determine level information of a plurality of transform coefficients by selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, and by using the context model and / or the Rice parameter being selected for determining the level information.
[0264] According to an embodiment, the reconstructed samples are represented in a spatial domain.
[0265] In an embodiment, the video decoder 200 is configured to select the context model and / or the Rice parameter depending on one or more residual samples in the neighboring area of the current position defined by the template.
[0266] According to an embodiment, the template is a non-local template.
[0267] In an embodiment, the video decoder 200 is configured to determine a correlation between a last significant scanning position and a residual energy of neighboring samples for context modeling.
[0268] According to an embodiment, the video decoder 200 is configured to determine the level information using the context model depending on a plurality of neighboring samples of a current position, wherein the plurality of neighboring samples depends on a prediction mode.
[0269] In an embodiment, the video decoder 200 is configured to determine the level information depending on a context modeling function which depends on a prediction mode.
[0270] According to an embodiment, a number of levels employing context modeling using neighboring samples is fixed and is limited to lowest frequency locations.
[0271] In an embodiment, a number of levels employing context modeling using neighboring samples is fixed with an actual number for a region or a transform block being determined depending on a prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream. According to an embodiment, neighboring samples of a current position are employed for context modeling of directly quantized residuals.
[0272] In an embodiment, the video decoder 200 is configured to receive a transform skip flag, and is configured to determine the level information depending on the transform skip flag.
[0273] According to an embodiment, the video decoder 200 is configured to operate in a transform skip mode. The video decoder 200, operating in the transform skip mode, is configured to select the context model and / or the Rice parameter depending on the neighboring reconstructed samples in the neighboring area of the current position defined by the template. Moreover, the video decoder 200 is configured to encode the level information using the context model and / or the Rice parameter being selected.
[0274] In an embodiment, the video decoder 200 is configured to employ a different residual error coding concept, if the video encoder 200 is operating in the transform skip mode, compared to when the video encoder 200 is not operating in the transform skip mode.
[0275] Fig. 8 illustrates a system according to an embodiment. The system comprises a video encoder 100 as described above. The video encoder 100 is configured to encode a video into a video data stream. Moreover, the system comprises a video decoder 200 as described above. The video decoder 200 is configured to receive the video data stream having the video encoded therein, wherein the video decoder 200 is configured to decode the video from the video data stream.
[0276] Moreover, a video data stream according to an embodiment is provided. The video data stream comprises, encoded therein, level information of each of a plurality of transform coefficients. Furthermore, the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
[0277] According to an embodiment, if the neighboring location is located outside a block of a picture of the video, in which the current transform coefficient is located, the coding information on the neighboring location is not available. In an embodiment, if the neighboring location has not been processed by a previous level coding step, the coding information on the neighboring location is not available.
[0278] According to an embodiment, if the neighboring location is a scanning position beyond a last significant scanning position, the coding information on the neighboring location is not available.
[0279] In an embodiment, if the neighboring location lies within a sub-block comprises zerovalued levels only, the coding information on the neighboring location is not available, and the video data stream comprises a coded sub-block flag indication indicating that the subblock comprises the zero-valued levels only.
[0280] According to an embodiment, if the coding information on the neighboring location is available, the encoding of the level information of the current transform coefficient depends on a first context model. If the coding information on the neighboring location is not available, the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
[0281] In an embodiment, if the coding information on all locations of two or more neighboring locations of the current transform coefficient is available, the encoding of the level information of the current transform coefficient depends on a first context model. If the coding information on at least one location of two or more neighboring locations of the current transform coefficient is not available, the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
[0282] According to an embodiment, if the coding information on the neighboring location is not available, the context model is a modified context model.
[0283] In an embodiment, the context model depends on a number of unavailable neighboring locations of the current transform coefficient.
[0284] According to an embodiment, the context model has been selected out of a set of two or more context models depending on a number of unavailable neighboring locations of the current transform coefficient.
[0285] In an embodiment, the context model has been selected out of a set of two or more context models as a selected context model depending on a number of unavailable neighboring locations of the current transform coefficient. The selected context model has been adjusted depending on the number of unavailable neighboring locations of the current transform coefficient.
[0286] According to an embodiment, the encoding of the level information of the current transform coefficient depends on the context model, wherein the context model depends on a diagonal position of the current transform coefficient, a coding state and a number of unavailable locations within a template around with respect to a position of the current transform coefficient.
[0287] Furthermore, a video data stream according to another embodiment is provided. The video data stream comprises, encoded therein, level information of a plurality of transform coefficients that have been encoded using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
[0288] According to an embodiment, the intermediate level information depends on an unquantized level of the current transform coefficient, wherein the quantized level information on the current transform coefficient depends on the intermediate level information on the current transform coefficient, the intermediate level information having a higher precision than the quantized level information, wherein the context modelling depends on the intermediate level information.
[0289] In an embodiment, the intermediate level information on the current transform coefficient indicates a level of the current transform coefficient having twice the size or twice the precision than a level of the current transform coefficient being indicated by the quantized level information.
[0290] According to an embodiment, the increased-precision information depends on information on a state of the trellis-coded quantization.
[0291] In an embodiment, the level information of the one of the plurality of transform coefficients depends on the context modelling without being modified and depends on halving one or more template sums and adding a rounding offset to the one or more template sums after halving. According to an embodiment, the level information of the one of the plurality of transform coefficients depends on the context modelling without being modified and depends on one or more template sums of two or more processed intermediate values that have been obtained by halving and adding a rounding offset to each of two more initial intermediate values.
[0292] In an embodiment, the level information of the one of the plurality of transform coefficients depends on the context modeling being adjusted, and depends on one or more template sums of intermediate levels.
[0293] According to an embodiment, the level information of the one of the plurality of transform coefficients depends on a state of the trellis-coded quantization and a partial sum of one or more templates for context modeling.
[0294] Moreover, a video data stream according to a further embodiment is provided. The video data stream comprises, encoded therein, level information of a plurality of transform coefficients, wherein the level information is encoded depending on a context model being selected and / or a Rice parameter being selected.
[0295] According to an embodiment, the reconstructed samples are represented in a spatial domain.
[0296] In an embodiment, the context model and / or the Rice parameter has been selected depending on one or more residual samples in the neighboring area of the current position defined by the template.
[0297] According to an embodiment, the template is a non-local template.
[0298] In an embodiment, the context model depends on a plurality of neighboring samples of a current position, wherein the plurality of neighboring samples depends on a prediction mode.
[0299] According to an embodiment, the level information depends on a context modeling function which depends on a prediction mode.
[0300] In an embodiment, a number of levels employing context modeling using neighboring samples is fixed and is limited to lowest frequency locations. According to an embodiment, a number of levels employing context modeling using neighboring samples is fixed with an actual number for a region or a transform block being determined depending on a prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream.
[0301] In an embodiment, neighboring samples of a current position are employed for context modeling of directly quantized residuals.
[0302] According to an embodiment, the video data stream comprises a transform skip flag depending on neighboring samples.
[0303] In the following, principles and particular embodiments of the present invention are described.
[0304] Context modeling is a pivotal component influencing the coding efficiency of entropy coding in hybrid video coding, particularly within the CABAC (Context-Adaptive Binary Arithmetic Coding) framework.
[0305] In embodiments, previously unidentified aspects of context modeling are presented that contribute to improved compression efficiency.
[0306] A first aspect of embodiments relates to the non-availability of a neighboring location in context modeling.
[0307] According to the first aspect of an embodiment, the context modeling, particularly the evaluation of the local template, is modified to include information about the availability of neighboring locations.
[0308] If neighboring locations are not available, this does not necessarily indicate that the position covered by the local template is outside the current block; it may also represent positions that have not been processed by the coding itself.
[0309] For example, a neighboring location could be a scanning position beyond the last significant scanning position. Another example is when a position lies within a sub-block consisting of zero-valued levels only. In this scenario, the so-called coded sub-block flag is coded with a value of zero, and all scanning positions within the sub-block are skipped by the coding process. Considering the unavailability of neighboring locations enhances coding efficiency by reducing uncertainty when those locations are unavailable. Consequently, this improvement enhances the precision of the estimation of statistical dependencies when the locations become available.
[0310] Different configurations are possible in the case of unavailability of a neighboring location given the local template configuration. For example, a different set of context models may be employed when one neighboring location is unavailable, effectively modeling the scenario where the local template size is smaller than usual.
[0311] In the following, particular embodiments of the present invention are described.
[0312] According to an embodiment, if, given a static template configuration, a location is outside of the block, it may be concluded that the neighboring position is unavailable.
[0313] According to an embodiment, scanning positions are considered unavailable that have not been processed by level coding.
[0314] In a particular embodiment, all scanning positions starting from the DC frequency to the last significant scanning position are considered available, given a static template configuration.
[0315] According to an embodiment, the context modeling is modified when one or more unavailable locations are present.
[0316] In a particular embodiment, the context modeling itself may depend on the number of unavailable locations.
[0317] In an embodiment, different context model sets are employed depending on the number of unavailable locations within the local template.
[0318] According to an embodiment, the combination of both previously mentioned embodiments is considered. In particular, the context modeling is adjusted when one or more unavailable locations are present, and different context model sets are utilized based on the number of unavailable locations within the local template. This comprehensive approach ensures adaptive and efficient context modeling in various scenarios. According to an embodiment, given a set of template configurations, the template configuration may, e.g., be selected that has no unavailable neighboring locations. Each template configuration may specify a different context modeling derived from the template sum.
[0319] In an embodiment, the local template is dynamically adjusted based on the real-time availability of neighboring locations. This dynamic adjustment allows for the optimization of context modeling by either expanding or contracting the template size to include only the available neighboring locations, thereby enhancing coding efficiency. In this embodiment, different context model sets are employed for the local template configuration.
[0320] According to an embodiment, context models are employed based on the characteristics of the unavailability patterns of neighboring locations. This selective activation is governed by predefined rules that consider the spatial distribution of unavailable locations, ensuring that the most appropriate context model is employed for each unique scenario.
[0321] In an embodiment, a hybrid context modeling approach is introduced for scenarios where both available and unavailable neighboring locations are present. This approach combines different context modeling strategies to exploit the available information fully while compensating for the lack of information due to unavailable locations.
[0322] According to an embodiment, the decision on context model sets, such as for coding sig, combines the diagonal (e.g., a diagonal position d = x + y), the state, and the count of unavailable locations within the local template.
[0323] In an embodiment, the number of context models within the context model sets varies. This variation may, e.g., depend on a rule that determines the context model set to use, which in turn depends on the number of unavailable locations within the local template.
[0324] Regarding unavailability, for example, the context modeling as described in VVC is considered.
[0325] In embodiments, unavailable locations may, e.g., be those that have not been processed by level coding. When an unavailable location occurs within the template, the context modeling switches to a different context model set for coding sig. Assuming all context models used for coding significance are organized in a single array, the final context model index for coding significance in luma in VVC may, for example, be determined as follows with d being the diagonal offset:
[0326] For example as: csig= min (3, (Ti + 1) » 1) + 4d, or for example as:csifl= 12 • max(0,S — 1) + min(3, (T±+ 1) » 1) + 4d.
[0327] When one neighboring location is marked as unavailable, a different context model set may, e.g., be used. In this scenario, the formula mentioned above may, e.g., be extended as follows: csig= 24 • max(0,S — 1) + min(3, (Tr+ 1) » 1) + 4d + a, where a is the activation offset applied when unavailability occurs.
[0328] A second aspect according to embodiments relates to the intermediate levels reconstructed during trellis-coded quantization (TCQ). When these intermediate levels are incorporated into context modeling, they provide a more accurate description of statistical dependencies, thereby enhancing coding efficiency.
[0329] According to the second aspect of an embodiment, context modeling with TCQ may, e.g., be improved by considering intermediate levels instead of decoded levels.
[0330] A key consideration is that, on average, the absolute value of intermediate levels is twice that of decoded levels. The additional precision is stored in the state and incorporated into the intermediate levels.
[0331] In an embodiment, halving the local template sum prior to context modeling may, e.g., be conducted, which does not require modifications to existing context modeling. However, this approach does not exploit the increased precision offered by intermediate levels.
[0332] Various options exist to leverage this increased precision and thus enhance coding efficiency, either without altering existing context modeling or by developing a new context modeling approach based on the statistical properties of intermediate levels.
[0333] According to an embodiment, the existing context modeling remains unmodified, and both ?! and T2are adapted to the existing context modeling by halving the template sums. However, to maintain precision, a rounding offset is added to the sums, so that the higher precision provided by the intermediate levels is retained after halving the template sums.
[0334] In an embodiment, the existing context modeling remains unmodified, while the adaptation of the value range is already performed during the calculation of the sums using the templates. In this preferred embodiment, a rounding offset is added to each intermediate level before halving the value. This approach ensures that the precision is maintained throughout the calculation process.
[0335] According to an embodiment, the context modeling is adjusted to the template sums with intermediate levels. Here, the context modeling and the Rice parameter derivation is modified by establishing updated thresholds that are suitable for the range of the intermediate levels.
[0336] In an embodiment, only the state and a partial sum of the templates are used for context modeling. This approach avoids the need to fully reconstruct intermediate levels, maintaining the value range for existing context modeling.
[0337] An example configuration for using the intermediate levels, without altering the existing context modeling, involves modifying the template sums for sig as follows:
[0338] For the derivation of par, gtl, and gt3, the adjustment is as follows: ccc= min (4, (Tx— ri) » 1).
[0339] It should be noted that in the equation mentioned above, the template sum is first reduced by the number of non-zero positions before being right-shifted by one. Specifically, subtracting the number of non-zero positions effectively remove the significance information for context modeling. Therefore, the proper method to incorporate intermediate levels while preserving the original value range is to subtract this number from the template before halving its value.
[0340] For the Rice parameter derivation, an approach according to an embodiment is: However, an alternative approach according to another embodiment that offers higher precision may, e.g., be: k = max(0, min(31, (T2— 10 ■ m) » 1 )), where m = 8 or m = 7 + (|x|&3) to accommodate for the maximum partial value coded in the first coding pass. It should be noted that for the bypass-only pass, no modification is necessary, i.e., m = 0.
[0341] In a third aspect of embodiments, context modeling utilizes neighboring reconstructed or residual samples. The above embodiments can be integrated into a comprehensive concept that leverages all of the above embodiments depending on the statistical dependencies identified during the parsing process. This concept cannot be only applied to regular transformed and quantized levels, but also to levels resulting from the identity transform. For example, the neighboring locations may indicate a content that is suitable for the identity transform so that the transform skip flag is inferred instead of being signaled.
[0342] According to embodiments, context modeling employs neighboring reconstructed or residual samples. Although the use of neighboring reconstructed samples for entropy coding is known, e.g., in conjunction with neural networks, their application in traditional context modeling has not been firmly established. To address this, a non-local template is used to define the area around reconstructed neighboring samples. Based on the analysis of this designated area, an appropriate context model and / or Rice parameter is selected. This method can operate independently or in conjunction with an existing local template. Unlike the local template, which is applied to the reconstructed levels of the current block and hence are all in the frequency domain, the non-local template covers reconstructed samples, i.e., the samples are in the spatial domain. For the sake of simplicity, consider only the context modeling for the DC frequency. Accordingly, a function is required to calculate the mean value of the reconstructed samples in this consideration. This mean value represents the DC in the frequency domain and can provide a significant input to the context modeling in conjunction with the chosen prediction mode.
[0343] According to an embodiment, reconstructed neighboring samples are utilized for context modeling.
[0344] In an embodiment, only neighboring residual samples are used for context modeling. According to an embodiment, the context modeling using neighboring samples involves the last significant scanning position. In this embodiment, the correlation between the last significant scanning position and the residual energy of the neighboring samples are measured for context modeling.
[0345] In an embodiment, the context modeling considers the prediction mode in conjunction with the neighboring samples. In this embodiment, the prediction mode may enable the consideration of the neighboring samples and / or lead to the selection of a different context modeling function.
[0346] According to an embodiment, the number of levels employing context modeling using neighboring samples is fixed and is limited to the lowest frequency locations.
[0347] In an embodiment, the number of levels employing context modeling using neighboring samples is fixed with the actual number for a region or a transform block is determined by the prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream.
[0348] According to an embodiment, the neighboring samples are employed for context modeling of directly quantized residuals, i.e. , when the identity transform is applied.
[0349] In an embodiment, the transform skip flag is inferred depending on the neighboring samples.
[0350] Although some aspects have been described in the context of an apparatus, it is clear that these aspects also represent a description of the corresponding method, where a block or device corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or an electronic circuit. In some embodiments, one or more of the most important method steps may be executed by such an apparatus.
[0351] Depending on certain implementation requirements, embodiments of the invention can be implemented in hardware or in software or at least partially in hardware or at least partially in software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium may be computer readable.
[0352] Some embodiments according to the invention comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.
[0353] Generally, embodiments of the present invention can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code may for example be stored on a machine readable carrier.
[0354] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0355] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0356] A further embodiment of the inventive methods is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) comprising, recorded thereon, the computer program for performing one of the methods described herein. The data carrier, the digital storage medium or the recorded medium are typically tangible and / or non-transitory.
[0357] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be configured to be transferred via a data communication connection, for example via the Internet.
[0358] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein. A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0359] A further embodiment according to the invention comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.
[0360] In some embodiments, a programmable logic device (for example a field programmable gate array) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.
[0361] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0362] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0363] The above described embodiments are merely illustrative for the principles of the present invention. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the impending patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.
Claims
Claims1. A video encoder (100), wherein the video encoder (100) is configured to encode a video into a video data stream, wherein the video encoder (100) is configured to generate the video data stream by encoding level information of each of a plurality of transform coefficients, wherein the video encoder (100) is configured to generate the video data stream such that the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
2. A video encoder (100) according to claim 1 , wherein, if the neighboring location is located outside a block of a picture of the video, in which the current transform coefficient is located, the coding information on the neighboring location is not available.
3. A video encoder (100) according to claim 1 or 2, wherein, if the neighboring location has not been processed by a previous level coding step, the coding information on the neighboring location is not available.
4. A video encoder (100) according one of the preceding claims, wherein, if the neighboring location is a scanning position beyond a last significant scanning position, the coding information on the neighboring location is not available.
5. A video encoder (100) according one of the preceding claims,wherein, if the neighboring location lies within a sub-block comprises zero-valued levels only, the coding information on the neighboring location is not available, and the video encoder (100) is configured to generate the video data stream such that the video data stream comprises a coded sub-block flag indication indicating that the sub-block comprises the zero-valued levels only.
6. A video encoder (100) according one of the preceding claims, wherein, if the coding information on the neighboring location is available, the video encoder (100) is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a first context model, and wherein, if the coding information on the neighboring location is not available, the video encoder (100) is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
7. A video encoder (100) according one of claims 1 to 5, wherein, if the coding information on all locations of two or more neighboring locations of the current transform coefficient is available, the video encoder (100) is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a first context model, and wherein, if the coding information on at least one location of two or more neighboring locations of the current transform coefficient is not available, the video encoder (100) is configured to generate the video data stream such that the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
8. A video encoder (100) according to one of claims 1 to 5, wherein, if the coding information on the neighboring location is not available, the video encoder (100) is configured to modify the context model.
9. A video encoder (100) according to one of claims 1 to 5,wherein the context model depends on a number of unavailable neighboring locations of the current transform coefficient.
10. A video encoder (100) according to one of claims 1 to 9, wherein the video encoder (100) is configured to select the context model out of a set of two or more context models depending on a number of unavailable neighboring locations of the current transform coefficient.
11. A video encoder (100) according to one of claims 1 to 9, wherein the video encoder (100) is configured to select the context model out of a set of two or more context models as a selected context model depending on a number of unavailable neighboring locations of the current transform coefficient, and wherein the video encoder (100) is configured to adjust the selected context model depending on the number of unavailable neighboring locations of the current transform coefficient.
12. A video encoder (100) according to one of the preceding claims, wherein the video encoder (100) is configured to select a template configuration out of two or more template configurations, wherein each of the two or more template configurations defines two or more neighboring locations which are taken into account for encoding the current transform coefficient, wherein the video encoder (100) is configured to select the template configuration out of two or more template configurations depending on an availability of coding information on the two or more neighboring locations, wherein the video encoder (100) is configured to determine the encoding of the level information of the current transform coefficient depending on the template configuration that has been selected.
13. A video encoder (100) according to claim 12,wherein the video encoder (100) is configured to select the template configuration out of two or more template configurations such that coding information on all of the two or more neighboring locations being defined by the template configuration is available.
14. A video encoder (100) according to claim 12 or 13, wherein the video encoder (100) is configured to select the template configuration out of two or more template configurations depending on an unavailability pattern defined by neighboring locations of the two or more neighboring locations for which coding information is unavailable.
15. A video encoder (100) according to one of claims 12 to 14, wherein a number of a set comprising the two or more template configuration varies depending on a rule, wherein the rule depends on the number of unavailable locations within a local template.
16. A video encoder (100) according to claim 15, wherein the rule defines that determining a context model index is conducted in a first way, if coding information on all of one or more neighboring locations is available, and wherein the rule defines that determining the context model index is conducted in a second way, being different from the first way, if coding information on at least one of one or more neighboring locations is unavailable.
17. A video encoder (100) according to one of the preceding claims, wherein the video encoder (100) is configured to dynamically adjust a template configuration, which defines two or more neighboring locations which are taken into account for encoding the current transform coefficient, depending on an availability of coding information on the two or more neighboring locations, wherein the video encoder (100) is configured to determine the encoding of the level information of the current transform coefficient depending on the template configuration that has been dynamically adjusted.
18. A video encoder (100) according to one of the preceding claims, wherein the video encoder (100) is configured to determine the encoding of the level information of the current transform coefficient depending on the context model such that the context model depends on a diagonal position of the current transform coefficient, a coding state and a number of unavailable locations within a template around with respect to a position of the current transform coefficient.
19. A video encoder (100), wherein the video encoder (100) is configured to encode a video into a video data stream, wherein the video encoder (100) is configured to generate the video data stream by encoding level information of a plurality of transform coefficients using trelliscoded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
20. A video encoder (100) according to claim 19, wherein the video encoder (100) is configured to generate intermediate level information of a current transform coefficient from an unquantized level of the current transform coefficient, wherein the video encoder (100) is configured to generate quantized level information on the current transform coefficient from the intermediate level information on the current transform coefficient, the intermediate level information having a higher precision than the quantized level information, wherein the video encoder (100) is configured to conduct context modelling depending on the intermediate level information.
21. A video encoder (100) according to claim 20,wherein the intermediate level information on the current transform coefficient indicates a level of the current transform coefficient having twice the size or twice the precision than a level of the current transform coefficient being indicated by the quantized level information.
22. A video encoder (100) according to one of claims 19 to 21, wherein the video encoder (100) is configured to determine the increased- precision information using information on a state of the trellis-coded quantization.
23. A video encoder (100) according to one of claims 19 to 22, wherein the video encoder (100) is configured for encoding the level information of the one of the plurality of transform coefficients using the context modelling without modifying the context modelling by halving one or more template sums and by adding a rounding offset to the one or more template sums after halving.
24. A video encoder (100) according to one of claims 19 to 22, wherein the video encoder (100) is configured for encoding the level information of the one of the plurality of transform coefficients using the context modelling without modifying the context modelling by generating one or more template sums of two or more processed intermediate values that have been obtained by halving and adding a rounding offset to each of two more initial intermediate values.
25. A video encoder (100) according to one of claims 19 to 22, wherein the video encoder (100) is configured for encoding the level information of the one of the plurality of transform coefficients using the context modelling by adjusting the context modeling depending on one or more template sums of intermediate levels, wherein the video encoder (100) is configured to modify the context modeling and a Rice parameter derivation.
26. A video encoder (100) according to one of claims 19 to 22,wherein the video encoder (100) is configured for encoding the level information of the one of the plurality of transform coefficients using a state of the trellis-coded quantization and a partial sum of one or more templates for context modeling.
27. A video encoder (100) according to one of claims 19 to 26, wherein the video encoder (100) is also a video encoder (100) according to one of claims 1 to 18.
28. A video encoder (100), wherein the video encoder (100) is configured to encode a video into a video data stream, wherein the video encoder (100) is configured to generate the video data stream by encoding level information of a plurality of transform coefficients, wherein the video encoder (100) is configured to select a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, wherein the video encoder (100) is configured to encode the level information using the context model and / or the Rice parameter being selected.
29. A video encoder (100) according to claim 28, wherein the reconstructed samples are represented in a spatial domain.
30. A video encoder (100) according to claim 28 or 29, wherein the video encoder (100) is configured to select the context model and / or the Rice parameter depending on one or more residual samples in the neighboring area of the current position defined by the template.
31. A video encoder (100) according to one of claims 28 to 30, wherein the template is a non-local template.
32. A video encoder (100) according to one of claims 28 to 31,wherein the video encoder (100) is configured to determine a correlation between a last significant scanning position and a residual energy of neighboring samples for context modeling.
33. A video encoder (100) according to one of claims 28 to 32, wherein the video encoder (100) is configured to encode the level information using the context model depending on a plurality of neighboring samples of a current position, wherein the plurality of neighboring samples depends on a prediction mode.
34. A video encoder (100) according to one of claims 28 to 33, wherein the video encoder (100) is configured to encode the level information depending on a context modeling function which depends on a prediction mode.
35. A video encoder (100) according to one of claims 28 to 34, wherein a number of levels employing context modeling using neighboring samples is fixed and is limited to lowest frequency locations.
36. A video encoder (100) according to one of claims 28 to 35, wherein a number of levels employing context modeling using neighboring samples is fixed with an actual number for a region or a transform block being determined depending on a prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream.
37. A video encoder (100) according to one of claims 28 to 36, wherein neighboring samples of a current position are employed for context modeling of directly quantized residuals.
38. A video encoder (100) according to one of claims 28 to 37, wherein the video encoder (100) is configured to infer a transform skip flag depending on neighboring samples.
39. A video encoder (100) according to one of claims 28 to 38, wherein the video encoder (100) is configured to operate in a transform skip mode, wherein the video encoder (100), operating in the transform skip mode, is configured to select the context model and / or the Rice parameter depending on the neighboring reconstructed samples in the neighboring area of the current position defined by the template, wherein the video encoder (100) is configured to encode the level information using the context model and / or the Rice parameter being selected.
40. A video encoder (100) according to claim 39, wherein the video encoder (100) is configured to employ a different residual error coding concept, if the video encoder (100) is operating in the transform skip mode, compared to when the video encoder (100) is not operating in the transform skip mode.
41. A video encoder (100) according to one of claims 28 to 40, wherein the video encoder (100) is also a video encoder (100) according to one of claims 1 to 27.
42. A video decoder (200) for receiving a video data stream having a video encoded therein, wherein the video decoder (200) is configured to decode the video from the video data stream, wherein, to decode the video, the video decoder (200) is configured to determine level information of a plurality of transform coefficients, wherein the video decoder (200) is configured to determine level information of a current transform coefficient of the plurality of transform coefficients depending on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
43. A video decoder (200) according to claim 42, wherein, if the neighboring location is located outside a block of a picture of the video, in which the current transform coefficient is located, the coding information on the neighboring location is not available.
44. A video decoder (200) according to claim 42 or 43, wherein, if the neighboring location has not been processed by a previous level coding step, the coding information on the neighboring location is not available.
45. A video decoder (200) according one of claims 42 to 44, wherein, if the neighboring location is a scanning position beyond a last significant scanning position, the coding information on the neighboring location is not available.
46. A video decoder (200) according one of claims 42 to 45, wherein, if the neighboring location lies within a sub-block comprises zero-valued levels only, the coding information on the neighboring location is not available, and wherein, to decode the video, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on a coded sub-block flag indication in the video data stream, which indicates that the subblock comprises the zero-valued levels only.
47. A video decoder (200) according one of claims 42 to 46, wherein, if the coding information on the neighboring location is available, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on a first context model, and wherein, if the coding information on the neighboring location is not available, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on a second context model being different from the first context model.
48. A video decoder (200) according one of claims 42 to 47,wherein, if the coding information on all locations of two or more neighboring locations of the current transform coefficient is available, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on a first context model, and wherein, if the coding information on at least one location of two or more neighboring locations of the current transform coefficient is not available, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on a second context model being different from the first context model.
49. A video decoder (200) according one of claims 42 to 47, wherein, if the coding information on the neighboring location is not available, the video decoder (200) is configured to modify the context model.
50. A video decoder (200) according one of claims 42 to 47, wherein the context model depends on a number of unavailable neighboring locations of the current transform coefficient.
51. A video decoder (200) according to one of claims 42 to 50, wherein the video decoder (200) is configured to select the context model out of a set of two or more context models depending on a number of unavailable neighboring locations of the current transform coefficient.
52. A video decoder (200) according to one of claims 42 to 50, wherein the video decoder (200) is configured to select the context model out of a set of two or more context models as a selected context model depending on a number of unavailable neighboring locations of the current transform coefficient, and wherein the video decoder (200) is configured to adjust the selected context model depending on the number of unavailable neighboring locations of the current transform coefficient.
53. A video decoder (200) according to one of claims 42 to 52, wherein the video decoder (200) is configured to select a template configuration out of two or more template configurations, wherein each of the two or more template configurations defines two or more neighboring locations which are taken into account for encoding the current transform coefficient, wherein the video decoder (200) is configured to select the template configuration out of two or more template configurations depending on an availability of coding information on the two or more neighboring locations, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on the template configuration that has been selected.
54. A video decoder (200) according to claim 53, wherein the video decoder (200) is configured to select the template configuration out of two or more template configurations such that coding information on all of the two or more neighboring locations being defined by the template configuration is available.
55. A video decoder (200) according to claim 53 or 54, wherein the video decoder (200) is configured to select the template configuration out of two or more template configurations depending on an unavailability pattern defined by neighboring locations of the two or more neighboring locations for which coding information is unavailable.
56. A video decoder (200) according to one of claims 53 to 55, wherein a number of a set comprising the two or more template configuration varies depending on a rule, wherein the rule depends on the number of unavailable locations within a local template.
57. A video decoder (200) according to claim 56,wherein the rule defines that determining a context model index is conducted in a first way, if coding information on all of one or more neighboring locations is available, and wherein the rule defines that determining the context model index is conducted in a second way, being different from the first way, if coding information on at least one of one or more neighboring locations is unavailable.
58. A video decoder (200) according to one of claims 42 to 57, wherein the video decoder (200) is configured to dynamically adjust a template configuration, which defines two or more neighboring locations which are taken into account for encoding the current transform coefficient, depending on an availability of coding information on the two or more neighboring locations, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on the template configuration that has been dynamically adjusted.
59. A video decoder (200) according to one of claims 42 to 58, the video decoder (200) is configured to determine the level information of the current transform coefficient depending on the context model such that the context model depends on a diagonal position of the current transform coefficient, a coding state and a number of unavailable locations within a template around with respect to a position of the current transform coefficient.
60. A video decoder (200) for receiving a video data stream having a video encoded therein, wherein the video decoder (200) is configured to decode the video from the video data stream, wherein, to decode the video, the video decoder (200) is configured to determine level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein theincreased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
61. A video decoder (200) according to claim 60, wherein the video decoder (200) is configured to generate intermediate level information of a current transform coefficient from an unquantized level of the current transform coefficient, wherein the intermediate level information has a higher precision than the quantized level information, wherein the video decoder (200) is configured to conduct context modelling depending on the intermediate level information.
62. A video decoder (200) according to claim 61 , wherein the intermediate level information on the current transform coefficient indicates a level of the current transform coefficient having twice the size or twice the precision than a level of the current transform coefficient being indicated by the quantized level information.
63. A video decoder (200) according to one of claims 60 to 62, wherein the video decoder (200) is configured to determine the increased- precision information using information on a state of the trellis-coded quantization.
64. A video decoder (200) according to one of claims 60 to 63, wherein the video decoder (200) is configured to use the context modelling without modifying the context modelling by halving one or more template sums and by adding a rounding offset to the one or more template sums after halving.
65. A video decoder (200) according to one of claims 60 to 63,wherein the video decoder (200) is configured to use the context modelling without modifying the context modelling by generating one or more template sums of two or more processed intermediate values that have been obtained by halving and adding a rounding offset to each of two more initial intermediate values.
66. A video decoder (200) according to one of claims 60 to 63, wherein the video decoder (200) is configured to use the context modelling by adjusting the context modeling depending on one or more template sums of intermediate levels, wherein the video decoder (200) is configured to modify the context modeling and a Rice parameter derivation.
67. A video decoder (200) according to one of claims 60 to 63, wherein the video decoder (200) is configured to use a state of the trellis-coded quantization and a partial sum of one or more templates for context modeling.
68. A video decoder (200) according to one of claims 60 to 67, wherein the video decoder (200) is also a video decoder (200) according to one of claims 42 to 59.
69. A video decoder (200) for receiving a video data stream having a video encoded therein, wherein the video decoder (200) is configured to decode the video from the video data stream, wherein, to decode the video, the video decoder (200) is configured to determine level information of a plurality of transform coefficients by selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, and by using the context model and / or the Rice parameter being selected for determining the level information.
70. A video decoder (200) according to claim 69, wherein the reconstructed samples are represented in a spatial domain.
71. A video decoder (200) according to claim 69 or 70, wherein the video decoder (200) is configured to select the context model and / or the Rice parameter depending on one or more residual samples in the neighboring area of the current position defined by the template.
72. A video decoder (200) according to one of claims 69 to 71, wherein the template is a non-local template.
73. A video decoder (200) according to one of claims 69 to 72, wherein the video decoder (200) is configured to determine a correlation between a last significant scanning position and a residual energy of neighboring samples for context modeling.
74. A video decoder (200) according to one of claims 69 to 73, wherein the video decoder (200) is configured to determine the level information using the context model depending on a plurality of neighboring samples of a current position, wherein the plurality of neighboring samples depends on a prediction mode.
75. A video decoder (200) according to one of claims 69 to 74, wherein the video decoder (200) is configured to determine the level information depending on a context modeling function which depends on a prediction mode.
76. A video decoder (200) according to one of claims 69 to 75, wherein a number of levels employing context modeling using neighboring samples is fixed and is limited to lowest frequency locations.
77. A video decoder (200) according to one of claims 69 to 76, wherein a number of levels employing context modeling using neighboring samples is fixed with an actual number for a region or a transform block beingdetermined depending on a prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream.
78. A video decoder (200) according to one of claims 69 to 77, wherein neighboring samples of a current position are employed for context modeling of directly quantized residuals.
79. A video decoder (200) according to one of claims 69 to 78, wherein the video decoder (200) is configured to receive a transform skip flag, and is configured to determine the level information depending on the transform skip flag.
80. A video decoder (200) according to one of claims 69 to 79, wherein the video decoder (200) is configured to operate in a transform skip mode, wherein the video decoder (200), operating in the transform skip mode, is configured to select the context model and / or the Rice parameter depending on the neighboring reconstructed samples in the neighboring area of the current position defined by the template, wherein the video decoder (200) is configured to encode the level information using the context model and / or the Rice parameter being selected.
81. A video decoder (200) according to claim 80, wherein the video decoder (200) is configured to employ a different residual error coding concept, if the video encoder (200) is operating in the transform skip mode, compared to when the video encoder (200) is not operating in the transform skip mode.
82. A video decoder (200) according to one of claims 69 to 81, wherein the video decoder (200) is also a video decoder (200) according to one of claims 42 to 68.
83. A system, comprising:a video encoder (100) according to one of claims 1 to 41 , wherein the video encoder (100) is configured to encode a video into a video data stream, and a video decoder (200) according to one of claims 42 to 82, wherein the video decoder (200) is configured to receive the video data stream having the video encoded therein, wherein the video decoder (200) is configured to decode the video from the video data stream.
84. A video data stream, wherein the video data stream comprises, encoded therein, level information of each of a plurality of transform coefficients, wherein the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
85. A video data stream according to claim 84, wherein, if the neighboring location is located outside a block of a picture of the video, in which the current transform coefficient is located, the coding information on the neighboring location is not available.
86. A video data stream according to claim 84 or 85, wherein, if the neighboring location has not been processed by a previous level coding step, the coding information on the neighboring location is not available.
87. A video data stream according one of claims 84 to 86, wherein, if the neighboring location is a scanning position beyond a last significant scanning position, the coding information on the neighboring location is not available.
88. A video data stream according one of claims 84 to 87, wherein, if the neighboring location lies within a sub-block comprises zero-valued levels only, the coding information on the neighboring location is not available, and the video data stream comprises a coded sub-block flag indication indicating that the sub-block comprises the zero-valued levels only.
89. A video data stream according one of claims 84 to 88, wherein, if the coding information on the neighboring location is available, the encoding of the level information of the current transform coefficient depends on a first context model, and wherein, if the coding information on the neighboring location is not available, the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
90. A video data stream according one of claims 84 to 88, wherein, if the coding information on all locations of two or more neighboring locations of the current transform coefficient is available, the encoding of the level information of the current transform coefficient depends on a first context model, and wherein, if the coding information on at least one location of two or more neighboring locations of the current transform coefficient is not available, the encoding of the level information of the current transform coefficient depends on a second context model being different from the first context model.
91. A video data stream according one of claims 84 to 88, wherein, if the coding information on the neighboring location is not available, the context model is a modified context model.
92. A video data stream according one of claims 84 to 88,wherein the context model depends on a number of unavailable neighboring locations of the current transform coefficient.
93. A video data stream according to one of claims 84 to 92, wherein the context model has been selected out of a set of two or more context models depending on a number of unavailable neighboring locations of the current transform coefficient.
94. A video data stream according to one of claims 84 to 92, wherein the context model has been selected out of a set of two or more context models as a selected context model depending on a number of unavailable neighboring locations of the current transform coefficient, and wherein the selected context model has been adjusted depending on the number of unavailable neighboring locations of the current transform coefficient.
95. A video data stream according to one of claims 84 to 94, wherein the encoding of the level information of the current transform coefficient depends on the context model, wherein the context model depends on a diagonal position of the current transform coefficient, a coding state and a number of unavailable locations within a template around with respect to a position of the current transform coefficient.
96. A video data stream, wherein the video data stream comprises, encoded therein, level information of a plurality of transform coefficients that have been encoded using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
97. A video data stream according to claim 96,wherein the intermediate level information depends on an unquantized level of the current transform coefficient, wherein the quantized level information on the current transform coefficient depends on the intermediate level information on the current transform coefficient, the intermediate level information having a higher precision than the quantized level information, wherein the context modelling depends on the intermediate level information.
98. A video data stream according to claim 97, wherein the intermediate level information on the current transform coefficient indicates a level of the current transform coefficient having twice the size or twice the precision than a level of the current transform coefficient being indicated by the quantized level information.
99. A video data stream according to one of claims 96 to 98, wherein the increased-precision information depends on information on a state of the trellis-coded quantization.
100. A video data stream according to one of claims 96 to 99, wherein the level information of the one of the plurality of transform coefficients depends on the context modelling without being modified and depends on halving one or more template sums and adding a rounding offset to the one or more template sums after halving.
101. A video data stream according to one of claims 96 to 99, wherein the level information of the one of the plurality of transform coefficients depends on the context modelling without being modified and depends on one or more template sums of two or more processed intermediate values that have been obtained by halving and adding a rounding offset to each of two more initial intermediate values.
102. A video data stream according to one of claims 96 to 99, wherein the level information of the one of the plurality of transform coefficients depends on the context modeling being adjusted, and depends on one or more template sums of intermediate levels.
103. A video data stream according to one of claims 96 to 99, wherein the level information of the one of the plurality of transform coefficients depends on a state of the trellis-coded quantization and a partial sum of one or more templates for context modeling.
104. A video data stream according to one of claims 96 to 103, wherein the video data stream is also a video data stream according to one of claims 84 to 95.
105. A video data stream, wherein the video data stream comprises, encoded therein, level information of a plurality of transform coefficients, wherein the level information is encoded depending on a context model being selected and / or a Rice parameter being selected.
106. A video data stream according to claim 105, wherein the reconstructed samples are represented in a spatial domain.
107. A video data stream according to claim 105 or 106, wherein the context model and / or the Rice parameter has been selected depending on one or more residual samples in the neighboring area of the current position defined by the template.
108. A video data stream according to one of claims 105 to 107, wherein the template is a non-local template.
109. A video data stream according to one of claims 105 to 108,wherein the context model depends on a plurality of neighboring samples of a current position, wherein the plurality of neighboring samples depends on a prediction mode.
110. A video data stream according to one of claims 105 to 109, wherein the level information depends on a context modeling function which depends on a prediction mode.
111. A video data stream according to one of claims 105 to 110, wherein a number of levels employing context modeling using neighboring samples is fixed and is limited to lowest frequency locations.
112. A video data stream according to one of claims 105 to 111 , wherein a number of levels employing context modeling using neighboring samples is fixed with an actual number for a region or a transform block being determined depending on a prediction mode, the quantization parameter, and / or a value signaled forward adaptive in a header of the bitstream.
113. A video data stream according to one of claims 105 to 112, wherein neighboring samples of a current position are employed for context modeling of directly quantized residuals.
114. A video data stream according to one of claims 105 to 113, wherein the video data stream comprises a transform skip flag depending on neighboring samples.
115. A video data stream according to one of claims 105 to 114, wherein the video data stream is also a video data stream according to one of claims 84 to 104.
116. A method for video encoding, wherein the method comprises encoding a video into a video data stream,wherein the method comprises generating the video data stream by encoding level information of each of a plurality of transform coefficients, wherein generating the video data stream is conducted such that the video data stream comprises an encoding of level information of a current transform coefficient of the plurality of transform coefficients, wherein the encoding of the level information of the current transform coefficient depends on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
117. A method for video encoding, wherein the method comprises encoding a video into a video data stream, wherein the method comprises generating the video data stream by encoding level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
118. A method for video encoding, wherein the method comprises encoding a video into a video data stream, wherein the method comprises generating the video data stream by encoding level information of a plurality of transform coefficients, wherein the method comprises selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, wherein the method comprises encoding the level information using the context model and / or the Rice parameter being selected.
119. A method for video decoding,wherein the method comprises for receiving a video data stream having a video encoded therein, wherein the method comprises decoding the video from the video data stream, wherein, to decode the video, the method comprises determining level information of a plurality of transform coefficients, wherein the method comprises determining level information of a current transform coefficient of the plurality of transform coefficients depending on a context model, wherein the context model depends on whether or not coding information on a neighboring location of the current transform coefficient is available.
120. A method for video decoding, wherein the method comprises for receiving a video data stream having a video encoded therein, wherein the method comprises decoding the video from the video data stream, wherein, to decode the video, the method comprises determining level information of a plurality of transform coefficients using trellis-coded quantization depending on a context modelling using increased-precision information on levels of the plurality of transform coefficients, wherein the increased-precision information on a level of one of the plurality of transform coefficients has a higher precision than a quantized level or a decoded level of said one of the plurality of transform coefficients, the decoded level being obtained from the quantized level.
121. A method for video decoding, wherein the method comprises for receiving a video data stream having a video encoded therein, wherein the method comprises decoding the video from the video data stream,wherein, to decode the video, the method comprises determining level information of a plurality of transform coefficients by selecting a context model and / or a Rice parameter depending on neighboring reconstructed samples in a neighboring area of a current position defined by a template, and by using the context model and / or the Rice parameter being selected for determining the level information.
122. A computer program for implementing the method of one of claims 116 to 121 when being executed on a computer or signal processor.
Citation Information
Patent Citations
Coefficient level coding in video coding
US20170064336A1