Decoder and method supporting adaptive dependent quantization of transform coefficient levels
By supporting multiple variants of dependent quantization and entropy coding methods, the trade-off between coding efficiency and complexity is resolved, improving the efficiency and flexibility of video coding and enabling flexible reconstruction and coding at the transform coefficient level.
Patent Information
- Application Number
- CN202080097352.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-20
- Filing Date
- 2020-12-18
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2040-12-18
AI Technical Summary
Existing video coding technologies struggle to find a suitable balance between coding efficiency and implementation complexity when processing the quantization of transform coefficients. Furthermore, traditional independent quantization methods cannot effectively utilize the dependencies between transform coefficients, resulting in low coding efficiency.
By employing a dependent quantization method that supports multiple variations and combining it with entropy coding under a unified architecture, and through state transition tables and context modeling of entropy coding, flexible reconstruction and coding at the transform coefficient level are achieved. This supports switching between independent quantization and dependent quantization, thereby improving coding efficiency.
By employing various variants of dependent quantization and entropy coding methods, coding efficiency is improved and coding complexity is reduced, achieving an optimal trade-off between coding efficiency and implementation complexity, thereby enhancing both coding efficiency and decoder flexibility.
Smart Images

Figure CN115152215B_ABST
Abstract
Description
1. TECHNICAL FIELD
[0001] Embodiments according to the present application relate to a decoder, an encoder and a method supporting adaptive dependent quantization of transform coefficient levels.
[0002] The present application is applicable to lossy coding of blocks of residual samples. A residual sample represents a difference between a sample of an original block of samples and a sample of a prediction signal (the prediction signal can be obtained by intra-picture prediction or inter-picture prediction, or by a combination of intra- and inter-picture prediction, or by any other means; in a special case, the prediction signal can be set equal to zero).
[0003] A block of residual samples is transformed using a signal transform. Typically, a linear and separable transform is used (a linear transform is linear, but can incorporate additional rounding of transform coefficients). Typically, an integer approximation of DCT-II or an integer approximation of another transform of the DCT / DST family is used. Different transforms can be used in horizontal or vertical direction. The transform is not limited to a linear and separable transform. Any other transform (linear and non-separable, or non-linear) can be used. In particular, a second order non-separable transform can be applied to the block of transform coefficients or to a part of the block of transform coefficients after a first separable transform. Due to the signal transform, a block of transform coefficients is obtained, which represents the original block of residual samples in a different signal space. In a special case, the transform can be equal to the identity transform (i.e. the block of transform coefficients can be equal to the block of residual samples). The block of transform coefficients is encoded using lossy coding. At the decoder side, the block of reconstructed transform coefficients is inverse transformed to obtain a reconstructed block of residual samples. Finally, a reconstructed block of picture samples is obtained by adding the prediction signal.
[0004] The present application describes a concept of lossy coding of blocks of transform coefficients. It particularly describes a method to realize a plurality of variants of dependent quantization and an adaptive selection of the applicable variant.
[0005] In this context, dependent quantization refers to a quantization method with the following properties: At the encoder side, a block of transform coefficients is mapped to a block of transform coefficient levels (i.e. quantization indices) which represent the transform coefficients with reduced fidelity. At the decoder side, the quantization indices are mapped to reconstructed transform coefficients (which differ from the original transform coefficients due to quantization). In contrast to traditional scalar quantization, the transform coefficients are not quantized independently. Rather, the set of admissible reconstructed levels for some transform coefficients depends on the quantization indices selected for other transform coefficients.
[0006] There are multiple variants of dependent quantization, which mainly differ in the number of possible so-called quantization states. Increasing the number of quantization states typically improves coding efficiency, but also increases encoder complexity. Therefore, it is reasonable to support multiple variants of dependent quantization in order to give the encoder the freedom to choose the appropriate trade-off between coding efficiency and implementation complexity. But for the decoder, the implementation complexity increases with the number of supported variants for quantization.
[0007] The present invention describes a method for supporting independent quantization as well as multiple variants of dependent quantization in a unified architecture. The reconstruction procedure and the entropy coding of transform coefficient levels are designed in such a way that the decoder can support all variants in a single implementation. The only aspect that depends on the chosen variant is the selection of the state transition table. The present invention comprises the following aspects:
[0008] • In addition to the traditional independent quantization, two or more variants of dependent quantization are supported;
[0009] • In particular, a variant of dependent quantization with 8 states is added;
[0010] • High-level signaling of the used quantization method;
[0011] • A unified procedure for reconstructing transform coefficient levels from quantization indices;
[0012] • A unified procedure for entropy coding of quantization indices (including a maximum number of standard for context-coded bins);
[0013] • A unified state-dependent context selection for the significance flag. 2. BACKGROUND
[0014] The following text describes the concepts of dependent quantization and entropy coding at the associated transform coefficient level. The description particularly emphasizes the design aspects specified in draft 7 for the Versatile Video Coding standard (VTM-7). However, it should be noted that many variations of this particular design are possible.
[0015] 2.1 Dependency Quantification
[0016] Dependent quantization of transform coefficients involves the following concept: the set of available reconstruction levels for transform coefficients depends on the quantization index (within the same transform block) selected for the preceding transform coefficient in the reconstruction order. Based on the quantization index used for the preceding transform coefficient (in encoding order), multiple sets of reconstruction levels are predefined, and one of these predefined sets is selected to reconstruct the current transform coefficient.
[0017] 2.1.1 Set of Refactoring Levels
[0018] In the union of predefined sets of reconstruction levels (two or more sets), a set of permissible reconstruction levels is selected for the current transform coefficient (based on the quantization index used for the preceding transform coefficients in coding order). The values of the reconstruction levels in the set of reconstruction levels are parameterized by block-based quantization parameters. The block-based quantization parameter (QP) determines the quantization step size Δ, and all reconstruction levels (within the entire set of reconstruction levels) represent integer multiples of the quantization step size Δ. This is used for a specific transform coefficient t. k (k represents the reconstruction order) quantization step size Δ k It can be determined not only by the block quantization parameter QP, but also for a specific transform coefficient t. k Quantization step size Δ k It may also be determined by the quantization weight matrix and block quantization parameters. Typically, it is used for the transform coefficients t. k Quantization step size Δ k It is used for the transformation coefficient t k weighting coefficient w k (Specified by the quantization weight matrix) and the block quantization step size Δ b The product of lock (specified by the block quantization parameter) is given.
[0019] Δ k =w k ·Δ block .
[0020] In VTM-7, dependent scalar quantization for transform coefficients uses two completely different sets of reconstruction levels. The transform coefficients t in these two sets... k The total reconstruction level is represented by the quantization step size Δ of the transform coefficients here. k The quantization step size is an integer multiple of the quantization step size (which is determined at least in part by the block-based quantization parameters). Note that the quantization step size Δ... k This precisely represents the scaling factor used for the permissible reconstructed values in both sets. Besides the scaling factor used for different transformation coefficients t within the transform block... k Possible individual quantization step size Δ k (And therefore, apart from the individual scaling factor) the same two sets of reconstruction levels are used for all transformation coefficients.
[0021] Figure 1 This displays two sets of reconstruction levels. Hollow and filled circles represent two distinct subsets within the set of reconstruction levels; these subsets can be used to determine the set of reconstruction levels used for the next transformation coefficients in the reconstruction order.
[0022] The two sets at the refactoring level are in Figure 1 The text shows:
[0023] • Included in the first quantization set (denoted as Figure 1 The reconstruction level in the set 0) represents an even integer multiple of the quantization step size.
[0024] • Second quantized set (labeled as) Figure 1 The set 1) contains all odd integer multiples of the quantization step size and additionally contains reconstruction levels equal to zero.
[0025] It should be noted that both reconstructed sets are symmetric to zero. Both reconstructed sets contain reconstructed levels equal to zero; otherwise, these reconstructed sets would be disjoint. The union of both reconstructed sets includes all integer multiples of the quantization step size.
[0026] 2.1.2 Representation of Refactoring Levels
[0027] The reconstruction level selected by the encoder from the permissible reconstruction levels must be indicated within the bitstream. As in traditional independent scalar quantization, this is achieved using a so-called quantization index (also known as a transform coefficient level). The quantization index (or transform coefficient level) uniquely identifies an integer within the quantization set (i.e., within the set of reconstruction levels) of the available reconstruction levels. The quantization index is transmitted to the decoder as part of the bitstream (using any entropy coding technique). On the decoder side, the reconstructed transform coefficients can be uniquely computed based on the current set of reconstruction levels (determined by the previous quantization index in encoding / reconstruction order) and the transmitted quantization indexes used for the current transform coefficients.
[0028] Figure 1 The reconstruction levels in the table are denoted by the associated quantization indices (the quantization indices are given by the numbers below the circles, which represent the reconstruction levels). The quantization index equal to 0 is assigned to the reconstruction level equal to 0. The quantization index equal to 1 is assigned to the smallest reconstruction level greater than 0, the quantization index equal to 2 is assigned to the next reconstruction level greater than 0 (i.e., the second smallest reconstruction level greater than 0)... and so on. Or, in other words, the reconstruction levels greater than 0 are denoted by integers greater than 0 in increasing order of numerical values (i.e., by 1, 2, 3,...). Likewise, the quantization index -1 is assigned to the largest reconstruction level less than 0, the quantization index -2 is assigned to the next reconstruction level less than 0 (i.e., the second largest)... and so on. Or, in other words, the reconstruction levels less than 0 are denoted by integers less than 0 in decreasing order of numerical values (i.e., by -1, -2, -3,...).
[0029] The use of reconstruction levels (which represent integer multiples of the quantization step size) will allow low-complexity algorithms for the computation of the reconstructed transform coefficients at the decoder side. This is shown in the preferred example based on Figure 1 The first set of quantizations includes all even integer multiples of the quantization step size, and the second set of quantizations includes all odd integer multiples of the quantization step size plus the reconstruction level equal to 0 (which is included in both sets of quantizations). The reconstruction procedure of the transform coefficients can be implemented with an algorithm represented by the pseudo-code similar to Figure 2
[0030] Figure 2 A pseudo-code illustrating the reconstruction procedure of the transform coefficients is shown. k represents the index indicating the reconstruction order of the current transform coefficient, the quantization index for the current transform coefficient is denoted by level[k], the quantization step size k is denoted by quant_step_size[k], and trec[k] represents the value of the reconstructed transform coefficient t' k The variable setId[k] indicates the set of reconstruction levels applied to the current transform coefficient. It is determined based on the preceding transform coefficients in the reconstruction order; the possible values of setId[k] are 0 and 1. The variable n indicates the integer factor of the quantization step size; it is given by the set of reconstruction levels selected (i.e., the value of setId[k]) and the transmitted quantization index level[k].
[0031] In the pseudo-code of Figure 2 , level[k] represents the quantization index for the transform coefficient t k The transmitted quantization index, and setId[k] (equal to 0 or 1) represents the identifier of the set of the current reconstruction level (which is determined based on the previous quantization index in reconstruction order, as detailed below). The variable n represents an integer multiple of the quantization step size given by the quantization index level[k] and the set identifier setId[k]. If using a quantization step size Δ... k If the transform coefficients are encoded using the first set of reconstruction levels that are even integer multiples of the quantization index (setId[k] == 0), then the variable n is twice the transmitted quantization index. If the transform coefficients are encoded using the second set of reconstruction levels (setId[k] == 1), then there are three cases: (a) if level[k] equals 0, then n also equals 0; (b) if level[k] is greater than 0, then n equals twice the quantization index level[k] minus 1; and (c) if level[k] is less than 0, then n equals twice the quantization index level[k] plus 1. This can be represented using a sign function.
[0032]
[0033] Next, if the second quantization set is used, the variable n is equal to twice the quantization index level[k] minus the sign function sign(level[k]) of the quantization index.
[0034] Once the variable n (an integer factor representing the quantization step size) is determined, then n is multiplied by the quantization step size Δ k And obtain the reconstruction transformation coefficients t′ k .
[0035] Figure 3 Explanation shown Figure 2 The pseudocode is an alternative implementation of the pseudocode in the transform block. The main change is the use of an integer implementation to represent the multiplication operation with the quantization step size, which uses scaling and shift parameters. Typically, for the transform block, the shift parameter (denoted by shift) is fixed, and only the scaling parameter (given by scale[k]) can depend on the position of the transform coefficients. The variable add represents the rounding offset, which is usually set to equal add = (1 << (shift - 1)). In Δ k Given the nominal quantization step size of the transform coefficients, select the parameters shift and scale[k] to make Δ k ≈scale[k]·2 -shift .
[0036] As mentioned above, the reconstruction transform coefficients t′ can be obtained through integer approximation. k To replace the quantization step size Δ k Exact multiplication. This is inFigure 3 The pseudocode in the code explains this. Here, the variable `shift` represents a bit shift to the right. Its value usually depends only on the block's quantization parameters (but it's possible that the shift parameters can change for different transform coefficients within the block). The variable `scale[k]` represents the transform coefficient `t`. k The scaling factor; besides the block quantization parameter, it can also depend, for example, on the corresponding elements of the quantization weight matrix. The variable `add` represents the rounding offset, which is usually set to equal `add = (1 << (shift - 1))`. It should be noted that, in addition to rounding, Figure 3 The integer arithmetic represented by the pseudocode in the last line is equivalent to the quantization step size Δ. k (Given the following) multiplication operations.
[0037] Δ k =scale[k]·2 -shift .
[0038] Compared to Figure 2 , Figure 3 Another change (on the surface) is the use of the ternary if-then-else operator (a?b:c) to switch between two sets at the refactoring level, and thus, for example, the C programming language can be known to use the ternary if-then-else operator (a?b:c).
[0039] 2.1.3 Dependent Reconstruction of Transformation Coefficients
[0040] Another important design aspect of dependent scalar quantization is the algorithm used to switch between defined sets of quantizations (sets of reconstructed levels). The algorithm used determines the "packing density" achievable in the N-dimensional space of the transform coefficients (and therefore also in the N-dimensional space of the reconstructed samples). A higher packing density ultimately results in increased coding efficiency.
[0041] In VTM-7, the transitions between quantization sets (set 0 and set 1) are determined by state variables (or quantization states). For the first transform coefficient in the reconstruction sequence, the state variable is set to a predetermined value. Typically, this predetermined value is 0. The state variables for subsequent transform coefficients in the encoding sequence are determined by an update procedure. The state of a particular transform coefficient depends only on the state and value of the previous transform coefficients in the reconstruction sequence.
[0042] The state variable has four possible values (0, 1, 2, 3). On the one hand, the state variable indicates the quantization set used for the current transform coefficient. When (and only when) the state variable is equal to 0 or 1, quantization set 0 is used; and when (and only when) the state variable is equal to 2 or 3, quantization set 1 is used. On the other hand, the state variable also indicates the possible transition between quantization sets.
[0043] The state of a particular transform coefficient depends only on the state of the previous transform coefficient in the reconstruction order and a binary function of the value of the previous transform coefficient. The binary function is referred to as the path in the following. In VTM-7, the following state transition table is used, where "path" refers to the binary function of the previous transform coefficient level in the reconstruction order.
[0044] Table 1: State transition table in VTM-6.
[0045]
[0046] In VTM-7, the path is given by the parity of the quantization index. In case level[k] is the transform coefficient level, it can be determined according to
[0047] path = (level[k] & 1),
[0048] where the operator "&" denotes the bitwise "and" in two's complement integer arithmetic.
[0049] As an alternative, the path can also represent other binary functions of level[k]. For example, it can indicate whether the transform coefficient level is equal to or not equal to 0:
[0050]
[0051] The concept of state transition for dependent scalar quantization allows for a low complexity implementation of the reconstruction of transform coefficients in the decoder. A preferred example of the reconstruction procedure for the transform coefficients of a single transform block is shown in the pseudo code in C language form in Figure 4
[0052] Figure 4 The pseudo code illustrates the reconstruction procedure for the transform coefficients of a transform block. The array level represents the transmitted transform coefficient levels (quantization indices) for the transform block, and the array trec represents the corresponding reconstructed transform coefficients. The two-dimensional table state_trans_table represents the state transition table, and the table setId represents the quantization sets related to the states. The function path() represents the binary function of the transform coefficient level.
[0053] In Figure 4 In the pseudo code, the index k denotes the reconstruction order of the transform coefficients. It is noted that in the example code, the index k is decreasing in reconstruction order. The index of the last transform coefficient is equal to k = 0. The first index kstartdenotes the reconstruction index (or more precisely the inverse reconstruction index) of the first reconstructed transform coefficient. The variable kstartmay be set equal to the number of transform coefficients in the transform block minus 1, or it can be set equal to the index of the first non-zero quantized index in the coding / reconstruction order (e.g. if the position of the first non-zero quantized index is transmitted in the applied entropy coding method). In the latter case, all preceding transform coefficients (for which k > kstart) are inferred to be equal to 0. The reconstruction procedure for each transform coefficient is identical to the example of Figure 3 Figure 3 In the example of Figure 3 , the quantized index is denoted by level[k] and the associated reconstructed transform by trec[k]. The state variable is denoted by state. The one-dimensional table setId[] denotes the set of quantizations associated with different values of the state variable, and the two-dimensional table state_trans_table[][] denotes the state transition, given the current state (first parameter) and the path (second parameter). As an example, the path can be given by the parity of the quantized index (using the bitwise AND operator &), but other concepts are possible. As another example, the path can denote whether the transform coefficient is equal or not equal to zero. Examples of the tables (in C language syntax) are given in Figure 5
[0054] Figure 5 The state transition table state_trans_table and the table setId are shown, which denote the set of quantizations associated with a state. The tables given in C language syntax denote the tables indicated in Table 1.
[0055] Arithmetic operations yielding the same results can be used instead of using the table state_trans_table[][] to determine the next state. Likewise, arithmetic operations can also be used to implement the table setId[]. Alternatively, a combination of table look-up using the one-dimensional table setId[] and the sign function can be implemented using arithmetic operations.
[0056] 2.2 Entropy coding of transform coefficient levels
[0057] A main aspect of dependent scalar quantization is that there are different sets of admissible reconstruction levels (also referred to as sets of quantizations) for the transform coefficients. The value of the quantized index of a preceding transform coefficient is used to determine the set of quantizations for the current transform coefficient. If we consider Figure 1It is apparent that the distance between the reconstructed level equal to zero and the neighboring reconstructed levels is larger in set 0 than in set 1. Therefore, the probability that the quantization index is equal to 0 is larger if set 0 is used and smaller if set 1 is used. In VTM-7, this effect is exploited in the entropy coding by switching the probability model based on the state used for the current quantization index.
[0058] It should be noted that for proper switching of the codeword table or the probability model, all the paths of the previous quantization indices (binary functions of the quantization indices) must be known when entropy decoding the current quantization index (or the corresponding binary decision of the current quantization index).
[0059] In VTM-7, the quantization indices are coded using binary arithmetic coding similar to H.264 | MPEG-4 AVC or H.265 | MPEG-H HEVC. For this purpose, the non-binary quantization indices are first mapped to a series of binary decisions, which are usually called bins.
[0060] 2.2.1 Binarization
[0061] The quantization indices are transmitted in absolute values and a sign (for the case that the absolute value is larger than 0). When the sign is transmitted in a single bin, there are many possibilities to map the absolute value to a series of bins.
[0062] Table 2: Binarization of the absolute value |q| of a transform coefficient level in VTM-7.
[0063]
[0064]
[0065] The binarization of the absolute value used by VTM-7 is shown in Table 2. The following binarized and non-binarized syntax elements are transmitted:
[0066] • sig_flag indicates whether the absolute value |q| of a transform coefficient level is larger than 0;
[0067] • If sig_flag is equal to 1, gt1_flag indicates whether the absolute value |q| of a transform coefficient level is larger than 1;
[0068] • If gt1_flag is equal to 1, par_flag indicates the parity of the absolute value |q| of a transform coefficient level and gt3_flag indicates whether the absolute value |q| of a transform coefficient level is larger than 3;
[0069] • If gt3_flag is equal to 1, the non-binary value rem indicates the remainder of the absolute level |q|. This syntax element is transmitted in bypass mode of the arithmetic coder, using a Golomb-Rice code.
[0070] Non-existing syntax elements are inferred to be equal to 0. At the decoder side, the absolute value of the transform coefficient level is reconstructed as follows:
[0071] |q| = sig_flag + gt1_flag + par_flag + 2 * (gt3_flag + rem)
[0072] For non-zero transform coefficient levels (indicated by sig_flag being equal to 1), sign_flag indicates the sign of the transform coefficient level which is transmitted additionally in bypass mode.
[0073] 2.2.2 Coding order and bypass mode of the binarized symbols
[0074] The entropy of the transform coefficient levels in VTM-7 has several aspects in common with HEVC. But it also includes additional aspects for the entropy coding of the dependent quantization.
[0075] Figure 6 Signaling of the position of the first non-zero quantization index 120 in coding order (highlighted samples). In addition to the position of the first non-zero transform coefficient 120, only the binarized symbols of the coefficients 122 following the first non-zero quantization index 120 in coding order 102 are transmitted (e.g. samples marked in blue), the coefficients preceding the first non-zero quantization index 120 in coding order 102 are inferred to be equal to 0 (e.g. samples marked in white).
[0076] Each block of transform coefficients is split into sub-blocks of fixed size (also called coefficient groups). In general, the sub-blocks have a size of 4x4 coefficients (see Figure 6 ). In some cases, other sizes (e.g. 2x2, 2x4, 4x2, 2x8, or 8x2) can be used.
[0077] Similar to HEVC, the coding of a block of transform coefficients is done as follows:
[0078] • First, a so-called coded_block_flag is transmitted which indicates whether there is any non-zero transform coefficient level in the transform block. If coded_block_flag is equal to 0, all transform coefficient levels are equal to 0 and no further data for the transform block is transmitted. In some cases (e.g. in Skip mode), the coded_block_flag can not be explicitly transmitted but inferred based on other syntax elements.
[0079] • if coded_block_flag is equal to 1, the x and y coordinates of the first significant transform coefficient in coding order are transmitted (this is sometimes referred to as the last significant coefficient, because the actual scan order indicates a scan from high frequency to low frequency components);
[0080] • the scan of transform coefficients is performed on a subblock basis (typically 4x4 subblocks); all transform coefficient levels of a subblock are coded before any transform coefficient level of any other subblock is coded; the subblocks are processed in a predetermined scan order, which starts with the subblock that contains the first significant coefficient in the scan order and ends with the subblock that contains the DC coefficient;
[0081] • the syntax includes a coded_subblock_flag that indicates whether a subblock contains any non-zero transform coefficient level; for the first and last subblock in the scan order (i.e. the subblock that contains the first significant coefficient in the scan order and the subblock that contains the DC coefficient), this flag is not transmitted and is inferred to be equal to 1; when coded_subblock_flag is transmitted, it is typically transmitted at the beginning of the subblock;
[0082] • for each subblock for which coded_subblock_flag is equal to 1, the transform coefficient levels are transmitted in multiple scan passes over the scan positions of the subblock.
[0083] In the following, the special aspect (which will be described later) will initially be ignored. Without this special aspect, the coding of a subblock for which coded_subblock_flag is equal to 1 proceeds as follows:
[0084] • for the first subblock in coding order (i.e. the subblock that contains the first significant scan position, whose x and y coordinates are explicitly transmitted), the coding starts at the scan position firstSigScanIdx of the first non-zero coefficient in the scan order (i.e. the scan position that corresponds to the explicitly transmitted x and y coordinates). For all other subblocks, the coding starts at the smallest scan index minSubblockScanIdx within the subblock. The scan ends at the largest scan index maxSubblockScanIdx within the subblock.
[0085] • in the first pass, the regular coding bins sig_flag, gt1_flag, par_flag, and gt3_flag are transmitted:
[0086] o If sig_flag can be inferred to be equal to 1, sig_flag is not transmitted. In the following two cases, sig_flag can be inferred to be equal to zero:
[0087] (1) The current scan position is equal to the scan index startScanIdx of the first non-zero coefficient in scan order (i.e. the scan position corresponding to the explicitly transmitted x and y coordinates);
[0088] (2) The current subblock is a subblock for which coded_subblock_flag has been transmitted equal to 1, the current scan index is the maximum scan index maxSubblockScanIdx within the subblock, and all previously transmitted sig_flag for the current subblock are equal to 1.
[0089] o If sig_flag[k] for scan index k is equal to 1 (transmitted or inferred), gt1_flag is transmitted. In other cases, gt1_flag is inferred to be equal to zero.
[0090] o If gt1_flag[k] for scan index k is equal to 1, par_flag and gt3_flag are transmitted. In other cases, par_flag and gt3_flag are inferred to be equal to zero.
[0091] • In a second pass, the remainder rem is coded for those scan indices for which gt3_flag has been transmitted equal to 1. For all other scan indices, the remainder rem is inferred to be equal to 0. The remainder rem is coded using a concatenation of a Rice code and an exponential Golomb code, which is parameterized by a so-called Rice parameter. The binary codeword for the corresponding code is coded in bypass mode. The Rice parameter is determined based on already coded syntax elements (see below).
[0092] • Finally, in a final pass, sign_flag is transmitted for all scan indices k for which sig_flag[k] is equal to 1 (coded or inferred), which indicates whether the transformed coefficient value is negative or positive.
[0093] The following pseudo-code further shows the encoding procedure for the transform coefficient levels within a subblock. It is presented from the decoder point of view. The Boolean variable firstSubblock indicates whether the current subblock is the first subblock in the encoding order (within the transform block). firstSigScanldx indicates the scan index which corresponds to the position of the first significant transform coefficient in the transform block (one of the transform block syntaxes is explicitly signaled at the beginning of the transform block syntax). minSubblockScanldx and maxSubblockScanldx represent the minimum and maximum values of the scan indices of the current subblock. It should be noted that the first scan index for which any data has been transmitted depends on whether the subblock includes the first significant coefficient in the scan order. If this is the case, the first scan pass starts at the scan index corresponding to the first significant transform coefficient; otherwise, the first scan pass starts at the minimum scan index of the subblock. The variable coeff[k] for a scan index k represents the reconstructed transform coefficient level. The encoding is done in simulation of the decoding.
[0094] Pseudo-code specifying the decoding of a sub-block for which coded_subblock_flag is equal to 1:
[0095]
[0096]
[0097] A drawback with respect to HEVC is that the maximum number of regular coded bins per transform coefficient is increased. To circumvent this problem, the following concepts are included in VTM-7:
[0098] • The maximum number of regular coded bins maxNumRegBins for a transform block is set equal to
[0099] maxNumRegBins = 1.75 * width * height,
[0100] where width and height represent the size of the transform block (more precisely, the size of the non-zero outer region of the transform block in case the transform coefficients are forced to be equal to zero).
[0101] • At the beginning of the transform block, the counter remRegBins (which represents the number of regular coded bins available) is set equal to remRegBins = maxNumRegBins.
[0102] • After a regular coded bin is encoded or decoded, the counter remRegBins is decreased by 1.
[0103] • At the beginning of the encoding of the scan position in the first pass, if the counter is less than 4 (in this case, not all binary codes sig_flag, gt1_flag, par_flag and gt3_flag can be transmitted without exceeding the maximum number of regular encoding binary codes), the first pass is ended. Moreover, for such scan position, only the rest of the first pass is transmitted.
[0104] • The absolute levels that are not encoded with data in the first scan pass are encoded in bypass mode only. They are encoded using the same class of codes as for the rest rem. The Rice parameter and the variable posO are determined based on the already encoded syntax elements (see below). These absolute levels |q| are not directly encoded, but first mapped to the syntax element abs_level, which is then encoded using the parameter code.
[0105] The mapping of the absolute value |q| to the syntax element abs_level depends on the variable posO. It is expressed as:
[0106] abs_level = (|q| == 0? posO : (|q| <= posO? |q| - 1 : |q|)
[0107] At the decoder side, the mapping of the syntax element abs_level to the absolute value |q| is expressed as follows:
[0108] |q| = (abs_level == posO? 0 : (abs_level < posO? abs_level + 1 : abs_level)
[0109] It should be noted that the number limitation of regular encoding binary codes is applied on transform block basis, not on sub-block basis. This means that the counter remRegBins is initialized at the beginning of a transform block and decremented for each regular encoding binary code. If the encoding is switched to bypass mode in a particular sub-block, all transform coefficient levels of all subsequent sub-blocks in the encoding order will also be encoded in bypass mode.
[0110] The following pseudo-code shows the encoding procedure for a sub-block (coded_subblock_flag equal to 1). Like the previous pseudo-code, this encoding procedure is presented from the decoder point of view. For the first encoding sub-block, the initial value of the counter remRegBins is equal to maxNumRegBins. For all subsequent sub-blocks, the initial value of the counter remRegBins is equal to the value obtained at the end of the previous encoding sub-block. The scan index startIdxBypass is the first scan index at which the bypass encoding (i.e. the encoding of abs_level) starts within a sub-block (if applicable).
[0111] Pseudo-code specifying the decoding (in bypass mode) of a sub-block for which coded_subblock_flag is equal to 1 (114):
[0112]
[0113]
[0114]
[0115] 2.2.3 Context Modelling
[0116] For the regular coded binarized symbols sig_flag, gt1_flag, par_flag, and gt3_flag, one of a number of probability models (or contexts) is selected for the actual coding.
[0117] Figure 7 A local template is shown to select the probability model for one or more binarized symbols. The black square 130 represents the current scan position, and the highlighted sample 132 (e.g. blue square) represents a neighboring scan position within the template.
[0118] 2.2.3.1 Significant flag sig_flag
[0119] For the significant flag, the selected probability model depends on:
[0120] • whether the current transform block is a luma or chroma transform block;
[0121] • the state of dependent quantization;
[0122] • the x and y coordinates of the current transform coefficient;
[0123] • the absolute value of the partial reconstruction in the local neighborhood (after the first pass).
[0124] In the following, more details will be given.
[0125] The variable state represents the state of the current transform coefficient used in dependent quantization. The state can take the values 0, 1, 2, or 3. Initially, the state is set equal to 0. As explained above, the state of the current transform coefficient is given by the state of the previous coefficient in coding order and the parity (or more generally a binary function) of the previous transform coefficient level.
[0126] In case x and y are the coordinates of the current transform coefficient within the transform block, let diag = x + y be the diagonal position of the current coefficient. Given the diagonal position, then the diagonal class index dsig is derived as follows:
[0127] dsig = (diag < 2? 2 : (diag < 5? 1 : 0)) for luma transform blocks
[0128] and
[0129] dsig = (diag < 2? 1 : 0) for chroma transform blocks
[0130] The ternary operator (c? a : b) represents an if-then-else statement. If the condition c is true, then the value a is used, otherwise (c is false) the value b is used.
[0131] The context also depends on the absolute values of the partial reconstructions within the local neighboring region. In VTM-7, the local neighboring region is given by the template T shown in Fig. 1. But other templates are possible as well. The template used can also depend on whether a luma or chroma block is coded. Let sumAbs be the sum of the absolute values of the partial reconstructions (after the first pass) in the template T: Figure 7
[0132]
[0133] where abs1[k] denotes the absolute level of the partial reconstruction of scan index k after the first pass. In case of the binarization of VTM-7, it is given by:
[0134] abs1[k] = sig_flag[k] + gt1_flag[k] + par_flag[k] + 2*gt3_flag[k].
[0135] Note that abs1[k] is equal to the value coeff[k] obtained after the first scan pass (see above).
[0136] It is assumed that the possible probability models for sig_flag are of the structure of a one-dimensional array. And let ctxIdSig be the index identifying the probability model used. According to VTM-7, the context index ctxIdSig can be derived as follows:
[0137] If the transform block is a luma block,
[0138] ctxIdSig = min((sumAbs + 1) » 1, 3) + 4*dsig + 12*min(state - 1, 0)
[0139] If the transform block is a chroma block,
[0140] ctxldSig = 36 + min((sumAbs + 1) » 1, 3) + 4 * dsig + 8 * min(state - 1, 0)
[0141] Here, the operator "»" denotes bit shift right (in two's complement arithmetic). This operation is the same as dividing by two and rounding the result down (to the next integer).
[0142] It is noted that other structures of the context model are possible. In any case, the selection of the probability model used for encoding sig_flag depends on:
[0143] • whether a luma or chroma block is coded;
[0144] • the state variable, in particular min(state - 1, 0);
[0145] • the diagonal class dsig;
[0146] • the sum of absolute values of the partial reconstruction within the local neighborhood (after the first pass), in particular min((sumAbs + 1) » 1, 3).
[0147] 2.2.3.2 Flags gt1_flag, par_flag, and gt3_flag
[0148] For the flags gt1_flag, par_flag, and gt3_flag, the selected probability model depends on:
[0149] • whether the current transform block is a luma or chroma transform block;
[0150] • whether the current transform coefficient is the first non-zero coefficient in coding order within the transform block;
[0151] • the x and y coordinates of the current transform coefficient;
[0152] • the absolute values of the partial reconstruction in the local neighborhood (after the first pass).
[0153] Let firstCoeff be a variable indicating whether the current scan index represents the scan index of the first non-zero transform coefficient level in coding order (i.e., a scan index that explicitly encodes the x and y positions). If the current scan index is equal to the scan index of the first non-zero transform coefficient level, then firstCoeff is equal to 1; otherwise, firstCoeff is equal to 0.
[0154] In case x and y are the coordinates of the current transform coefficient within the transform block, let diag = x + y be the diagonal position of the current coefficient. Given the diagonal position, the diagonal class index dclass is derived as follows:
[0155] dclass = (diag == 0? 3 : (diag < 3? 2 : (diag < 10? 1 : 0))) for luma transform blocks
[0156] and
[0157] dclass = (diag == 0? 1 : 0) for chroma transform blocks
[0158] The context also depends on the absolute values of the partial reconstruction within the local neighboring region. The variable sumAbs is derived as shown above. Furthermore, the variable numSig, which represents the number of non-zero transform coefficient levels within the template T, is derived as follows:
[0159]
[0160] Similarly for sig_flag, the assumed probability model is a structure of one-dimensional arrays. And let ctxld be the index used to identify the probability model. According to VTM-7, the context index ctxld can be derived as follows:
[0161] In case the transform block is a luma block,
[0162] ctxld = (firstCoeff? 0 : 1 + min(sumAbs - numSig, 4) + 5 * dclass)
[0163] In case the transform block is a chroma block,
[0164] ctxld = 21 + (firstCoeff? 0 : 1 + min(sumAbs - numSig, 4) + 5 * dclass)
[0165] It should be noted that other structures of the context model are possible. But in any case, the chosen probability model depends on:
[0166] • whether a luma or chroma block is coded;
[0167] • whether the current scan index is the scan index of the first non-zero transform coefficient level in coding order;
[0168] • the diagonal class dclass;
[0169] • The difference between the sum of the absolute values of the partial reconstruction within the local sample (after the first pass) and the number of non-zero transform coefficient levels within the local template, specifically min(sumAbs-numSig,4).
[0170] It should be noted that the same context index ctxld is used for gt1_flag, par_flag, and gt3_flag. But for each of these flags, a different set of context models is used.
[0171] 2.2.3.3 Rice parameters for transform coefficient levels of the remaining part and of the bypass coded
[0172] Both the remaining part rem and the absolute value abs_level are coded using a parameter class code, where the binary code is coded in bypass mode (see above). The actual binarization used is determined by a so-called Rice parameter RP. Furthermore, for the absolute level abs_level, the variable posO determines the mapping from the coded syntax element abs_level to the actual absolute value.
[0173] The Rice parameter RPrem for the remaining part is derived based on the sum of the absolute values of the reconstruction in the local template. In VTM-7, the same template as described above is used. In case T denotes the template and coeff[k] denotes the reconstructed coefficient at scan index k, the sum sumAbs is derived as follows
[0174]
[0175] It should be noted that the fully reconstructed transform coefficient values coeff[k] are used. Let tabRPrem[] be a fixed table of size 32. The Rice parameter RPrem is derived as follows:
[0176] RPrem = tabRPrem[max(min(sumAbs - 20, 31), 0)].
[0177] Let tabRPabs[] be another fixed table of size 32. The Rice parameter RPabs for coding abs_level is derived as follows:
[0178] RPabs = tabRPabs[min(sumAbs, 31)].
[0179] The variable posO depends on both the sum of the absolute values of the reconstruction in the local template and a state variable. The variable posO is derived from the quantization state and the Rice parameter RPabs as follows:
[0180] posO = (state < 2? 1 : 2) « RPabs.
[0181] 2.3 High level signaling of quantization method
[0182] VTM-7 supports dependent quantization with 4 states and additionally supports the legacy independent scalar quantization. The quantization method selected by the encoder is signaled in the bitstream as described below.
[0183] The picture header includes the syntax element pic_dep_quant_enabled_flag, which is coded with a single bit and indicates whether the current picture uses or does not use dependent quantization. If pic_dep_quant_enabled_flag is equal to 0, the current picture uses the legacy independent quantization. If pic_dep_quant_enabled_flag is equal to 1, the current picture uses a particular variant of dependent quantization.
[0184] The picture header syntax element is not always present in the picture header. Its presence is controlled by another syntax element coded in the picture parameter set. The picture parameter set includes the syntax element pps_dep_quant_enabled_idc, which is coded with a fixed length code of 2 bits and has the following significance:
[0185] • the value 0 indicates that pic_dep_quant_enabled_flag is present in the picture header;
[0186] • the value 1 indicates that pic_dep_quant_enabled_flag is not present in the picture header, but pic_dep_quant_enabled_flag is inferred to be equal to 0 for all pictures referring to the picture parameter set;
[0187] • the value 2 indicates that pic_dep_quant_enabled_flag is not present in the picture header, but pic_dep_quant_enabled_flag is inferred to be equal to 1 for all pictures referring to the picture parameter set;
[0188] • the value 3 is reserved for future use.
[0189] If the picture parameter set flag constant_slice_header_params_enabled_flag is equal to 1, the pps_dep_quant_enabled_idc is present in the picture parameter set. If the picture parameter set flag constant_slice_header_params_enabled_flag is equal to 0, the picture parameter set syntax element pps_dep_quant_enabled_idc is not present in the picture parameter set, but it is inferred to be equal to 0 (in this case, the pic_dep_quant_enabled_flag is coded in the picture header).
[0190] Therefore, it is desirable to provide concepts that make image and / or video coding more efficient to support adaptive dependent quantization at the level of transform coefficients. Additionally, or alternatively, it is desirable to reduce the bitstream and thus the signalization cost.
[0191] This is the object of the independent claims of the present application.
[0192] Further embodiments according to the application are defined by the objects of the dependent claims of the present application. 3. SUMMARY
[0193] According to a first aspect of the present application, the inventors have understood that when trying to use dependent quantization / dequantization at the level of transform coefficients, a problem is encountered, which arises from the fact that providing a large number of dependent quantization / dequantization modes involves a high implementation complexity. According to the first aspect of the present application, this difficulty is overcome by unifying the architecture for the various variants of independent quantization / dequantization and dependent quantization / dequantization. The inventors have found that it is advantageous to select the quantizer independently of the quantization mode information. The only aspect that depends on the quantization mode information is the update of the current conversion state, e.g. the selection of the conversion table that can be used for the update. This is based on the idea that a unified procedure for all supported quantization modes introduces almost no additional implementation complexity at the decoder side with respect to supporting a single quantization mode. But if different approaches have to be supported, the decoder complexity increases. Thus, allowing the use of different quantization modes, thereby giving the encoder the freedom to choose the variant, provides the most appropriate trade-off between coding efficiency and implementation complexity for the field of application. Furthermore, the selection of the quantizer (independent of the quantization mode) allows the use of dependent quantization modes with more than 4 conversion states at the encoder / decoder, thereby enabling an improved coding efficiency because of a more dense compactness of the admissible reconstruction points in the N-dimensional signal space.
[0194] Thus, according to a first aspect of the present application, a decoder is configured to decode from a data stream residual levels, the residual levels representing prediction residuals, and to dequantize the residual levels sequentially by: depending on a current conversion state, selecting a quantizer from a set of default quantizers; dequantizing a current residual level using the quantizer to obtain a dequantized residual value; applying a binary function to the current residual level to obtain a property (e.g. a parity check) of the current residual level, and depending on the property of the current residual level, updating the current conversion state with a conversion table to be used. The dequantization of the residual levels is performed sequentially, which means that the above-mentioned dequantization steps are performed in a loop (e.g. until all residual values are dequantized), and thus the updating of the current conversion state can be understood as a first step or a final step in this sequence. The updating step of the current conversion state connects the successive dequantizations of the residual levels, e.g. the dequantization of a previous residual level is connected to the dequantization of a current residual level by the updating of the current conversion state. Furthermore, e.g. in the case of the dequantization of a first residual level, the current conversion state can be set to a default / predetermined state (e.g. zero), and depending on the property of the current residual level (i.e. the first residual level property), the current conversion state is updated with a conversion table to be used (e.g. to initialize the sequential dequantization of the residual levels). The decoder is configured to use the dequantized residual values to reconstruct a media signal. Furthermore, the decoder is configured to select, depending on quantization mode information comprised in the data stream, a default conversion table from a set of default conversion tables as the conversion table to be used, each default conversion table representing a surjective mapping from a domain of a combination of a set of one or more conversion states having the property of the current residual level to the set of one or more conversion states, wherein the default conversion tables differ in the cardinality of the set of one or more conversion states. The surjective mapping e.g. maps elements (i.e. conversion states) of the set of one or more conversion states to elements (i.e. conversion states) of the set of one or more conversion states. Both the domain and the codomain of the surjective mapping are e.g. associated with the same set of one or more conversion states. The selection of the conversion table to be used e.g. can be performed for each residual level at the updating of the current conversion state; can be performed for each transform block comprising a plurality of residual levels; or can be performed per image. Furthermore, the decoder is configured to perform the selection of the quantizer by mapping a predetermined number of bits of the current conversion state to the default quantizers using a first mapping, regardless of which of the default conversion tables is selected as the conversion table to be used; the predetermined number and the first mapping are equal regardless of which of the default conversion tables is selected as the conversion table to be used. The predetermined number of bits represents e.g. a number of bits (e.g. LSBs or bits located at the second but not the last significant bit position) located at one or more predetermined bit positions.
[0195] Thus, according to a first aspect of the present application, a decoder is configured to decode residual levels from a data stream, the residual levels representing prediction residuals, and to inversely quantize the residual levels sequentially by: selecting a quantizer from among a set of default quantizers in dependence on a current transition state; inversely quantizing a current residual level using the quantizer to obtain an inversely quantized residual value; and updating the current transition state in dependence on a property of the current residual level obtained by applying a binary function to the current residual level (e.g. a parity check), and in dependence on quantization mode information included in the data stream. As mentioned above, the inverse quantization of the residual levels is performed sequentially, which means that the above-mentioned inverse quantization steps are performed in a loop (e.g. until all residual values are quantized), and thus the updating of the current transition state can be understood as a first step or a final step in this sequence. The updating of the current transition state connects consecutive inverse quantizations of residual levels, e.g. the inverse quantization of a previous residual level is connected to the inverse quantization of a current residual level by updating the current transition state. Furthermore, e.g. in the case of the inverse quantization of a first residual level, the current transition state can be set to a default / predetermined state (e.g. zero), and this current transition state is updated (e.g. to initialize the sequential inverse quantization of the residual levels) in dependence on a property of the current residual level (i.e. a first residual level property). The current transition state is transformed (e.g. at the time of updating) from a domain of combinations of one or more transition states of a set of one or more transition states having a property of the current residual level to the set of one or more transition states according to a surjective mapping depending on the quantization mode information, wherein the cardinality of the set of one or more transition states differs in dependence on the quantization mode information. Irrespective of the quantization mode information (e.g. independent of the quantization mode information, or independent of the quantization mode information), the decoder is configured to perform the selection of the quantizer by mapping a predetermined number of bits of the current transition state (e.g. one or more bits at one or more predetermined bit positions, e.g. LSBs or bits at the second but not the last significant bit position) to a default quantizer using a first mapping; the predetermined number and the first mapping are equal irrespective of the quantization mode information. Furthermore, the decoder is configured to reconstruct a media signal using the inversely quantized residual values.
[0196] Thus, according to a first aspect of the present application, an encoder is configured to: predictively encode a media signal to obtain a residual signal; and sequentially quantize residual values representing the residual signal to obtain residual levels by: selecting a quantizer from among a set of default quantizers in dependence on a current conversion state; quantizing a current residual value using the quantizer to obtain a current residual level; applying a binary function to the current residual level to obtain a property (e.g. a parity check) of the current residual level; and updating the current conversion state in dependence on the property of the current residual level using a conversion table to be used. The quantization of the residual levels is performed sequentially, which means that the above-mentioned quantization steps are performed in a loop (e.g. until all residual values are quantized), and thus the updating of the current conversion state can be understood as a first step or a final step in this sequence. The updating step of the current conversion state connects the sequential quantization of the residual levels, e.g. the quantization of a previous residual level is connected with the quantization of a current residual level by the updating of the current conversion state. Furthermore, e.g. in case of the quantization of a first residual level, the current conversion state can be set to a default / predetermined state (e.g. zero), and this current conversion state is updated in dependence on the property of the first residual level (i.e. the first residual level property) using the conversion table to be used (e.g. to initialize the sequential dequantization of the residual levels). The encoder is configured to encode the residual levels into a data stream. Furthermore, the encoder is configured to select a default conversion table from among a set of default conversion tables as the conversion table to be used in dependence on quantization mode information transmitted in the data stream, each default conversion table representing a surjective mapping from a domain of a combination of a set of one or more conversion states having the property of the current residual level to the set of one or more conversion states, wherein the default conversion tables differ in the cardinality of the set of one or more conversion states. The surjective mapping e.g. maps elements of the set of one or more conversion states (i.e. conversion states) to elements of the set of one or more conversion states (i.e. conversion states). Both the domain and the codomain of the surjective mapping are e.g. associated with the same set of one or more conversion states. The selection of the conversion table to be used e.g. can be performed for each residual level at the updating of the current conversion state; can be performed for each transform block comprising a plurality of residual levels; or can be performed per image. Furthermore, the encoder is configured to perform the selection of the quantizer by mapping a predetermined number of bits of the current conversion state to the default quantizers using a first mapping, regardless of which of the default conversion tables is selected as the conversion table to be used; the predetermined number and the first mapping are equal regardless of which of the default conversion tables is selected as the conversion table to be used. The predetermined number of bits represents e.g. a number of one or more bits located at one or more predetermined bit positions (e.g. LSBs or bits located at the second but not the last significant bit position).
[0197] Thus, according to a first aspect of the application, an encoder is configured to predictively encode a media signal to obtain a residual signal, and to sequentially quantize residual values representing the residual signal to obtain residual levels by: selecting a quantizer from a set of default quantizers in dependence on a current transition state; quantizing a current residual value using the quantizer to obtain a current residual level; and updating the current transition state in dependence on a property (e.g. parity) of the current residual level obtained by applying a binary function to the current residual level, and in dependence on quantization mode information transmitted in a data stream. As mentioned above, the quantization of the residual levels is performed sequentially, which means that the quantization steps described above are performed in a loop (e.g. until all residual values are quantized), and thus the updating of the current transition state can be understood as a first step or a final step in this sequence. The updating step of the current transition state connects the sequential quantization of the residual levels, e.g. the quantization of a previous residual level is connected to the quantization of a current residual level by the updating of the current transition state. Furthermore, e.g. in case of the quantization of a first residual level, the current transition state can be set to a default / predetermined state (e.g. zero), and this current transition state is updated (e.g. to initialize the sequential dequantization of the residual levels) in dependence on the property of the current residual level (i.e. the first residual level property). The current transition state is transformed (e.g. at the time of the updating) from a domain of combinations of one or more transition states of a set of one or more transition states having the property of the current residual level to the set of one or more transition states according to an onto mapping depending on the quantization mode information, wherein the cardinality of the set of one or more transition states differs in dependence on the quantization mode information. Irrespective of the quantization mode information (e.g. independent of the quantization mode information, or independent of the quantization mode information), a decoder is configured to perform the selection of the quantizer by mapping a predetermined number of bits of the current transition state (e.g. one or more bits at one or more predetermined bit positions, e.g. the LSBs or the bits at the second but not the last significant bit position) to a default quantizer using a first mapping; the predetermined number and the first mapping are equal irrespective of the quantization mode information. Furthermore, the encoder is configured to encode the residual levels into the data stream.
[0198] Thus, according to a first aspect of the application, a method comprises decoding from a data stream a residual level, the residual level representing a prediction residual; and sequentially dequantizing the residual level by: selecting a quantizer from a set of default quantizers in dependence on a current transition state; dequantizing the current residual level using the quantizer to obtain a dequantized residual value; and updating the current transition state in dependence on a property of the current residual level (e.g. a parity) obtained by applying a binary function to the current residual level, and in dependence on quantization mode information included in the data stream. Further, the method comprises reconstructing the media signal using the dequantized residual value. The current transition state is converted from a domain of combinations of a set of one or more transition states having the property of the current residual level to the set of one or more transition states according to a surjective mapping depending on the quantization mode information, wherein a cardinality of the set of one or more transition states differs in dependence on the quantization mode information. Further, the method comprises performing the selection of the quantizer by mapping a predetermined number of bits of the current transition state (e.g. one or more bits at one or more predetermined bit positions, e.g. LSBs or bits at a second but not the last significant bit position) to the default quantizers using a first mapping, regardless of the quantization mode information (e.g. independent of the quantization mode information). The predetermined number and the first mapping are equal regardless of the quantization mode information.
[0199] Thus, according to a first aspect of the application, a method comprises predictively encoding a media signal to obtain a residual signal; and sequentially quantizing residual values representing the residual signal to obtain residual levels by: selecting a quantizer from a set of default quantizers in dependence on a current transition state; quantizing the current residual value using the quantizer to obtain the current residual level; and updating the current transition state in dependence on a property of the current residual level (e.g. a parity) obtained by applying a binary function to the current residual level, and in dependence on quantization mode information transmitted in a data stream. Further, the method comprises encoding the residual levels into the data stream. The current transition state is converted from a domain of combinations of a set of one or more transition states having the property of the current residual level to the set of one or more transition states according to a surjective mapping depending on the quantization mode information, wherein a cardinality of the set of one or more transition states differs in dependence on the quantization mode information. Further, the method comprises performing the selection of the quantizer by mapping a predetermined number of bits of the current transition state (e.g. one or more bits at one or more predetermined bit positions, e.g. LSBs or bits at a second but not the last significant bit position) to the default quantizers using a first mapping, regardless of the quantization mode information (e.g. independent of the quantization mode information). The predetermined number and the first mapping are equal regardless of the quantization mode information.
[0200] The above described methods are based on the same considerations as the aforementioned encoder / decoder. Incidentally, the methods can be done with all the features and functionalities, which are also described in the way with respect to the encoder / decoder.
[0201] An embodiment relates to a data stream having an image or video encoded therein with the encoding method described herein.
[0202] An embodiment relates to a computer program having a program code which, when the program code is executed on a computer, performs the method described herein. 4. BRIEF DESCRIPTION OF DRAWINGS
[0203] The accompanying drawings are not necessarily drawn to scale; instead, the drawings are generally intended to illustrate the principles of the application; in the following description, various embodiments of the application are described with reference to the drawings, in which:
[0204] Figure 1 Two different quantizers are shown;
[0205] Figure 2 Pseudo code illustrating the reconstruction procedure of the transform coefficients is shown;
[0206] Figure 3 Pseudo code illustrating an alternative implementation of the reconstruction procedure of the transform coefficients is shown;
[0207] Figure 4 Pseudo code illustrating the reconstruction procedure of the transform coefficients for a transform block is shown;
[0208] Figure 5 A state transition table is shown;
[0209] Figure 6 An encoding order is shown;
[0210] Figure 7 A local template for a probability model that can be used to select one or more binary codes is shown;
[0211] Figure 8 A decoder according to an embodiment is shown;
[0212] Figure 9 An embodiment of context-adaptive binary arithmetic coding for significant value binary codes is shown;
[0213] Figure 10 A conversion table according to an embodiment is shown;
[0214] Figure 11 Pseudo code illustrating the reconstruction procedure of the transform coefficients for a transform block according to an embodiment is shown;
[0215] Figure 12 An image and / or video encoder is shown;
[0216] Figure 13 An image and / or video decoder is shown; and
[0217] Figure 14 The relationship between the reconstructed signal and the combination of the prediction residual signal and the prediction signal is shown. 5. DETAILED DESCRIPTION
[0218] In the following description, identical or similar components or components having identical or similar functions are designated by identical component symbols (even if they appear in different drawings).
[0219] In the following description, numerous details are set forth to provide a more complete explanation of embodiments of the present application. It will be apparent, however, to one skilled in the art, that embodiments of the present application can be practiced without some or all of these specific details. In other instances, well known structures and devices are not shown in detail or are shown in block diagram form in order to avoid obscuring embodiments of the present application. Further, different embodiments of the described features can be combined with each other in any manner.
[0220] Embodiments described in the following will mainly explain features and functions in the perspective of a decoder. However, it is clear that an encoder can comprise identical or similar features and functions, e.g. a decoding performed by a decoder can correspond to an encoding by an encoder. Further, an encoder can comprise identical features as described with respect to a decoder in a feedback loop, e.g. in the prediction stage 36.
[0221] Figure 8 Embodiments of a decoder 20 are shown. The decoder 20 can optionally comprise features and / or functions as described in section 2 (e.g. one or more features described in sections 2.1 to 2.3).
[0222] The decoder 20 is configured to decode 210 the residual level 142 (which represents the prediction residual) from the data stream 14. The residual level 142 can be a quantized level of a transform coefficient representing the prediction residual. As shown, the residual level 142 can represent an integer multiple of a quantization step size 144. Figure 1
[0223] The decoder 20 is configured to inverse quantize 150 the residual levels 142 sequentially. The residual levels 142 can be inverse quantized 150 along the encoding order (i.e. the scan order or the reconstruction order). Different encoding orders are possible and the encoding order can depend on the media signal, e.g. an image signal, a video signal or a sound signal to be decoded by the decoder 20. Figure 6 An example of an encoding order 102 that can be used for decoding an image or video signal is shown. According to an embodiment, the decoder 20 can inverse quantize 150 the residual levels 142 starting with inverse quantizing 150 the first non-zero transform coefficient (i.e. the first transform coefficient in the encoding order whose residual level is not equal to zero) and continuing with inverse quantizing 150 the residual levels 142 of the subsequent transform coefficients in the encoding order until the last transform coefficient in the encoding order. The residual levels of the transform coefficients in the encoding order preceding the first non-zero quantized index can be inferred by the decoder 20 to be equal to zero.
[0224] The decoder 20 is configured to inverse quantize 150 the residual levels 142 sequentially by: selecting 152 a quantizer 153 from the set of default quantizers 151 depending on a current transform state 220; inverse quantizing 154 the current residual level using the quantizer 153 to obtain an inverse quantized residual value 104; and updating 158 the current transform state depending on a characteristic 157 of the current residual level obtained by applying 156 a binary function to the current residual level and depending on the quantization mode information 180 included in the data stream 14. The current residual level is inverse quantized 154 to obtain the inverse quantized residual value 104, wherein the quantizer 153 used for inverse quantizing 154 depends on the current transform state 220. After this inverse quantization 154, the current transform state is updated 158 to obtain a new current transform state, which is used for example for inverse quantizing the next residual level or for inverse quantizing the current residual level again using the updated current transform state (see for example the inverse quantization 150 of the first residual level below). The new current transform state depends for example on the current transform state and on the characteristic 157 of the current residual level. The characteristic 157 of the current residual level can represent a characteristic of the current residual level and not of the next residual level. The transform state of a particular transform coefficient can depend on the transform state of a preceding transform coefficient in the encoding order and on the characteristic 157 of the preceding transform coefficient.
[0225] At update 158, the current conversion state is converted from a domain of a combination 163 of a set 161 of one or more conversion states with a characteristic 157 of a current residual level to the set 161 of one or more conversion states according to a surjective mapping that depends on the quantization mode information 180, wherein a cardinality l, m, n of the set 161 of one or more conversion states differs depending on the quantization mode information 180. For example, in case the quantization mode information 180 indicates a first quantization mode, the set 161a of one or more conversion states has a cardinality l; in case the quantization mode information 180 indicates a second quantization mode, the set 161b of one or more conversion states has a cardinality m; and in case the quantization mode information 180 indicates a third quantization mode, the set 161c of one or more conversion states has a cardinality n. According to an embodiment, the decoder 20 can be configured to use the set 161a of one or more conversion states with a cardinality l (e.g. l = 1) or to use the set 161b of one or more conversion states with a cardinality m (e.g. m = 4) depending on the quantization mode information 180. Figure 10 Three possible surjective mappings 106a, 106b and 106c are shown.
[0226] According to an embodiment, the binary function is a parity.
[0227] In Figure 8 , the binary function can be referred to as a "path" and the first residual level characteristic can relate to "path 0" and the second residual level characteristic can relate to "path 1". According to an embodiment, the decoder 20 is configured to obtain a parity of the residual level as the residual level characteristic 157. The decoder can be configured to determine whether the residual level has an even or an odd parity by applying 156 a binary function to the current residual level. For example, "path 0" can correspond to an even parity and "path 1" can correspond to an odd parity.
[0228] As Figure 8 shown, the decoder 20 is configured to perform a selection 152 of a quantizer 153 (independent of the quantization mode information 180, e.g. irrespective of the quantization mode information 180) by mapping a predetermined number of bits of the current conversion state 220 to a default quantizer using a first mapping 200, wherein the predetermined number and the first mapping 200 are equal irrespective of the quantization mode information 180. The predetermined number of bits represents one or more bits of the current conversion state 220 located at one or more predetermined bit positions, e.g. the LSB or a bit located at a second but not the last significant bit position.
[0229] According to an embodiment, the set 151 of default quantizers (as shown in Figure 1 is composed of a first default quantizer 1511 and a second default quantizer 1512, wherein the first default quantizer 1511 comprises even integer multiples of the quantization step size 144 as reconstruction levels 142, and the second default quantizer 1512 comprises odd integer multiples of the quantization step size 144 and zero as reconstruction levels 142.
[0230] According to an embodiment, the first mapping 200 maps a current transform state 220 equal to one or less than one to the first default quantizer 1511 and maps a current transform state 220 greater than one to the second default quantizer 1512. The predetermined number of bits can be one, and the decoder 20 can be configured to map a second bit (e.g. the second LSB) of the current transform state 220 to the first default quantizer or to the second default quantizer.
[0231] Figure 8 The decoder 20 shown is configured to reconstruct 170 the media signal using the dequantized residual values 104. According to an embodiment, the decoder 20 is configured to perform the reconstruction 170 by obtaining a predicted version of the media signal by performing a prediction and correcting the predicted version with the dequantized residual values 104; or to reconstruct the media signal by obtaining a predicted version of the media signal by performing a prediction with the dequantized residual values 104; inverse transforming the dequantized residual values 104 to obtain a media residual signal; and correcting the predicted version with the media residual signal.
[0232] As mentioned above, the residual levels 142 are dequantized 150 in order. There are different choices to dequantize 150 the first residual level (i.e. the residual level of the first non-zero transform coefficient), wherein the decoder 20 can be configured to use one of these choices.
[0233] The decoder 20 can be configured to set 190 the current conversion state for the current residual level being the first residual level to a predetermined value. For example, the predetermined value can be zero. Based on this predetermined current conversion state, the selection 152 of the quantizer 153 can be performed and the first residual level is dequantized 154, e.g., using the selected quantizer 153. The decoder 20 can be configured to apply 156 the binary function to the first residual level to obtain the first residual level characteristic as the characteristic 157 of the current residual level and to update 158 the current conversion state in dependence on the characteristic 157 of the current residual level. According to a first option, the decoder 20 can be configured to use the updated conversion state as the current conversion state for performing the dequantization 150 for the next residual level. According to a second option, the decoder 20 is configured to use the updated conversion state as the current conversion state for performing the dequantization 150 of the first residual level again before performing the dequantization 150 of the next residual level. Alternatively, according to a third option, the decoder 20 can be configured to set 190 the current conversion state for the current residual level being the first residual level to the predetermined value and to apply 156 the binary function to the first residual level to obtain the first residual level characteristic as the characteristic 157 of the current residual level. For example, the predetermined value can be zero. Further, the decoder 20 is configured to update 158 the predetermined current conversion state in dependence on the characteristic 157 of the current residual level to obtain the current conversion state. Based on this current conversion state, the selection 152 of the quantizer 153 can be performed and the first residual level is dequantized 154, e.g., using the selected quantizer 153. Further, the decoder 20 can be configured to perform the updating 158 again to obtain the current conversion state for the next residual level.
[0234] According to an embodiment, the decoder 20 is configured to update 158 the current conversion state with the conversion table 106 to be used. For example, the decoder 20 is configured to inversely quantize 150 the residual level 142 sequentially by: selecting 152 a quantizer 153 from the set 151 of default quantizers in dependence on the current conversion state; inversely quantizing 154 the current residual level using the quantizer 153 to obtain an inversely quantized residual value 104; applying 156 a binary function to the current residual level to obtain a property 157 of the current residual level; and updating 158 the current conversion state with the conversion table 106 to be used in dependence on the property 157 of the current residual level. The decoder can be configured to select the default conversion table to be used as the conversion table 106 from the set 159 of default conversion tables 106a, b, c, each of which represents a surjective mapping from a domain of combinations 163 of the set 161 of one or more conversion states with the property 157 of the current residual level to the set 161 of one or more conversion states, in dependence on the quantization mode information 180 comprised in the data stream 14, wherein the default conversion tables differ in the cardinality / , m, n of the set 161 of one or more conversion states. Furthermore, the decoder 20 is, for example, configured to perform the selection 152 of the quantizer by mapping a predetermined number of bits of the current conversion state 220 to the default quantizer using a first mapping 200, which is equal, regardless of which default conversion table is selected as the conversion table to be used, as well as the predetermined number and the first mapping 200.
[0235] According to an embodiment, the set 159 of default conversion tables 106a, b, c comprises two or more of the following:
[0236] a first default conversion table 106, wherein the cardinality of the set of one or more conversion states is one,
[0237] a second default conversion table 106, wherein the cardinality of the set of one or more conversion states is four,
[0238] a third default conversion table 106, wherein the cardinality of the set of one or more conversion states is eight.
[0239] According to an embodiment, the set 159 of default conversion tables 106a, b, c comprises: a first default conversion table 106a, wherein the cardinality of the set of one or more conversion states is one; and a second default conversion table 106b, wherein the cardinality of the set of one or more conversion states is four.
[0240] In other words, depending on the quantization mode information 180, the cardinalities / , m, n of the surjective mapping and the set of one or more transition states 161 can be selected among two or more of the following: the cardinality / of the set of one or more transition states 161a is one, the cardinality m of the set of one or more transition states 161b is four, and the cardinality n of the set of one or more transition states 161c is eight. According to an embodiment, the cardinalities / , m, n of the surjective mapping and the set of one or more transition states 161 can be selected among the cardinality / of the set of one or more transition states 161a being one and the cardinality m of the set of one or more transition states 161b being four.
[0241] According to an embodiment, the decoder 20 is configured to manage the set 159 of default transition tables 106a, b, c as different parts of a uniform transition table; within the uniform transition table, for each default transition table 106a, b, c, each of the set of one or more transition states 161 is indexed using a state index which is different from the state index used to index any other default transition table within the uniform transition table (e.g. the default transition tables are disjoint from each other for the transition states of the different parts of the uniform transition table).
[0242] According to an embodiment, the decoder 20 is configured to perform the update 158 of the current transition state 220, e.g. depending on the characteristic 157 of the current residual level, by looking up in the uniform transition table an entry corresponding to the combination 163 of the characteristic 157 of the current residual level and the current transition state 220, using the transition table 106 of the next transition state 162 to be used for the next dequantized residual level.
[0243] According to an embodiment, the quantization mode information 180 comprises a syntax element in the data stream, the syntax element indicating a starting transition state for dequantizing a first residual level in a dequantization order (e.g. sequential dequantization is performed along this order), by fitting to the state index of the part of the uniform transition table only for the transition table to be used, such that the syntax element indicates the transition table to be used from the set of default transition tables.
[0244] According to an embodiment, the media signal is a video, and the decoder 20 is configured to read the quantization mode information 180 from the data stream 14, and to perform the selection of a default transition table 106 from the set of default transition tables in one of the following ways: once for the video, per image block, on a per image basis, on a per slice basis, and on a per image sequence basis.
[0245] In other words, the decoder 20 can be configured to read the quantization mode information 180, which controls the update to be performed in one of the following ways: once for the video, per image block, on a per image basis, on a per slice basis, and on a per image sequence basis.
[0246] According to an embodiment, the media signal is a color video, and the decoder 20 is configured to use different versions of the set 151 of default quantizers for different color components.
[0247] The decoder 20 and / or a corresponding encoder can comprise features and / or functionalities as described in one or more of the following embodiments.
[0248] The present invention refers to a decoder, an encoder and a method to support multiple variants of dependent quantization in a unified architecture. The present invention specifically describes a concept in which the entropy decoding and the reconstruction procedure (e.g. the reconstruction 170 of the media signal) is basically the same for all supported quantization methods, i.e. independent of the quantization mode information 180, e.g. only the state transition table 106 is selected based on the quantization method selected by the encoder and signaled in the bitstream (i.e. the data stream 14). In a preferred embodiment of the present invention, three quantization methods are supported, which can be indicated by the quantization mode information 180:
[0249] • a conventional independent scalar quantization;
[0250] • a dependent quantization with 4 states (i.e. transition states); and
[0251] • a dependent quantization with 8 states (i.e. transition states).
[0252] In another embodiment, an additional variant of a dependent quantization with more than 8 states (e.g. 16 states) is supported. In another embodiment of the present invention, only the conventional quantization and versions of the dependent quantization with more than 4 states are supported. In another embodiment, two or more variants of a dependent quantization with more than 4 states are supported. According to another aspect, only the conventional quantization and the dependent quantization with 4 states are supported.
[0253] An aspect of the present invention is to design all decoder operations to be independent of the actually used quantization method. In this aspect, the selection of the state transition table 106 can depend on the selected quantization method. Once the state transition table 106 is given, the decoding procedure is independent of the selected quantization method.
[0254] The advantages of the present invention are as follows:
[0255] • By supporting dependent quantization with more than 4 states, the coding efficiency can be improved due to achieving a denser packing of admissible reconstruction points in the N-dimensional signal space (this has been shown in the literature).
[0256] • By supporting different variants of dependent quantization (which differ in the number of states), the encoder is given the freedom to choose the variant which provides the most appropriate trade-off between coding efficiency and implementation complexity in its field of application. It should be noted that the encoder complexity depends on the number of quantization states supported, i.e. on the cardinality l, m, n of the set 161 of one or more transition states. While dependent quantization with 8 states yields higher coding efficiency than dependent quantization with 4 states, it also requires higher encoder complexity.
[0257] • By designing a unified decoder program for all supported quantization methods, essentially no additional implementation complexity is required at the decoder side. It should be noted that the decoding complexity is almost independent of the number of quantization states. But if different methods have to be supported, the decoder complexity will increase. By unifying the decoding program for all supported quantization methods, the implementation complexity of the decoder is almost constant with respect to supporting a single method of dependent quantization.
[0258] 5.1 Basic concept of the present application
[0259] In VTM-7, dependent quantization with 4 states is specified and the following state transition table 106, e.g. state transition table 106b in case the cardinality m of the set 161b of one or more transition states is four:
[0260]
[0261]
[0262] In VTM-5, the path is given by the parity check of the current transform coefficient level, i.e. the current residual level (or quantization index) in coding order 102. In case level[k] denotes the current transform coefficient level, the variable path is determined as follows:
[0263] path = (level[k] & 1),
[0264] where "&" denotes the bitwise "and" operator in binary complement arithmetic. As mentioned above, the path can be any binary function of the quantization index. Nevertheless, using the parity check is preferred.
[0265] In the decoding program of VTM-7, some operations are dependent on the state variable, i.e. the transition state. In particular, the following three aspects:
[0266] • The quantizer 153 selected for the reconstruction of the transform coefficient depends on the current conversion state 220. In VTM-7, the quantizer identifier (quantizer id) which can be zero or one is given by Qld = state » 1.
[0267] Given the current value of the state variable state and the current quantization index level[k], the reconstructed transform coefficient trec[k] (i.e. the dequantized residual value 104) is obtained as follows:
[0268]
[0269] where the second line within the if-branch represents an integer variable which is the product with the quantization step size 144. The expression (state » 1) indicates the used quantizer 153, e.g. the expression (state » 1) indicates the first mapping 200. The operator "»" represents a bit right shift.
[0270] • See Figure 9 , the context model (i.e. the context 242) used for encoding the significant flag (i.e. the significance bin 92) depends on the state variable (i.e. the current conversion state 220). In VTM-7, it is a function of the value 230 max(state-1,0) and optionally of other variables. It can be interpreted as using a set of three different context models 242. The first set is used for transform coefficient levels (i.e. residual levels 142) where the state variable 220 is equal to 0 or 1; the second set is used for transform coefficient levels 142 where the state variable 220 is equal to 2; and the third set is used for transform coefficient levels 142 where the state variable 220 is equal to 3.
[0271] • The parameter pos0 (e.g. representing the binarization parameter 302) used for encoding the syntax element dec_abs_level (transform coefficient level encoded in bypass mode) depends on the quantization state (i.e. the conversion state assigned to the individual residual level). In VTM-7, the variable pos0 is determined as follows:
[0272] pos0 = (state < 2? 1 : 2) « RPabs
[0273] where RPabs represents the Rice parameter selected based on the sum of absolute transform coefficients in the local neighboring region.
[0274] 5.1.1 Another variant of increasing dependent quantization
[0275] At this point, for example, an additional variant of dependent quantization with 8 states should be supported in the bitstream syntax as well as in the decoding process. This variant can use the following state transition table 106c (e.g. a set of one or more transition states 161c in case the base n is eight):
[0276]
[0277] For this variant, the quantizer 153 is given as Qld = state » 2.
[0278] Conceptually, two quantizers would be preferred that deviate from the chosen quantizers in VTM-7.
[0279] See Figure 9 For the coding significance flag 92, more than 3 different context sets 242 would be preferred. Using the same relation as for the 4-state variant (i.e. min(state - 1, 0)), 7 context sets 242 would be needed. The number of context sets 242 can be limited to 3, but then different rules than for the 4-state variant would have to be used.
[0280] Furthermore, the derivation of the posO variable (i.e. the binarization parameter 302) in VTM-7 (for coding dec_abs_level) depends on the used quantizer 153. For the first quantizer 1511, it is set to equal
[0281] posO = 1 « RPabs,
[0282] and for the second quantizer 1512, it is set to equal
[0283] posO = 2 « RPabs.
[0284] Therefore, for the 8-state variant, a different implementation would be needed.
[0285] 5.1.2 Unification of the decoding process
[0286] The idea of the proposed concept is to unify the decoding process for all quantization variants supported. The quantization variant of the residual levels can be indicated by the quantization mode information 180. The first observation that allows unification is that the quantization states (i.e. the transition states) can be relabeled without affecting the result of the reconstruction process. For example, the VTM-7 state transition table given by
[0287]
[0288] can be rewritten as
[0289]
[0290] Here, the meaning of states 1 and 2 is exchanged. For the original table, the quantizer 153 is given by Qld = state » 1. For the reformulated table 106, the quantizer 153 is given by Qld = state & 1, where "&" denotes the bitwise "and" operator.
[0291] In a similar way, the state transition table 106 for the 8-state version (or any other variant of dependent quantization) can be reformulated such that the same relationship (i.e., the first mapping 200) can be used to derive the quantizer identifier (i.e., the quantizer 153 to be used).
[0292] In order to further unify the decoding process, a compromise between coding efficiency and implementation complexity has to be made. For example, for the entropy coding of the significant flag 92, it is not possible to choose the best version that is adapted to all supported variants of the quantizer, but a unified version can still be designed, even if this is a suboptimal way for some of the supported variants of dependent quantization.
[0293] 5.1.3 Signaling of the selected quantization variant
[0294] Finally, the selected quantization variant (i.e., the quantizer 153 to be used) has to be indicated in the bitstream (i.e., the data stream 14). For this purpose, various variants are possible:
[0295] • it can be selected on block level (e.g., per coding tree unit or coding unit); in this case, a dedicated syntax element has to be transmitted for each corresponding block;
[0296] • it can be selected on slice or picture level (or, tile or tile group level); in this case, a dedicated syntax element has to be transmitted in the slice or picture header (or, tile or tile group);
[0297] • it can be selected on sequence level; in this case, a dedicated syntax element has to be transmitted in the sequence parameter set or similar syntax structure.
[0298] In one implementation, different quantizers can be used for different color components (e.g., different quantizers can be used for luma and chroma).
[0299] 5.2 Description of specific implementations
[0300] In the following, implementations of the present application will be described in more detail. However, it should be noted that other implementations are possible as well (as described above and in the following description). The present application is not limited to the specific designs described in the following.
[0301] It is a feature of the present application that two or more of the following aspects are combined: • the first mapping 200 is used to derive the quantizer identifier (i.e., the quantizer 153 to be used);
[0302] • Bitstream syntax and decoding procedure supporting one or more variants of independent quantization as well as dependent quantization, wherein at least one variant of dependent quantization uses 4 or more quantization states;
[0303] • The used quantization variant is indicated by a dedicated syntax element transmitted at block, slice, picture, or sequence level. In a preferred version, it is transmitted in the picture header.
[0304] • The quantization state of a transform coefficient is derived by a state transition procedure (i.e. surjective mapping) consisting of the following steps:
[0305] o The state (i.e. transition state) of the first transform coefficient in coding order (i.e. first residual level) is set to a predetermined value; in a preferred version, it is set equal to 0.
[0306] o The state of the current transform coefficient (i.e. current residual level) is derived based on the state of the preceding transform coefficient in coding order and a binary function of the value of the quantization index of the preceding transform coefficient (i.e. binary function of the value of the preceding residual level). In a preferred version, the binary function represents a parity check, i.e. the current state is given by a parity check of the preceding state and the preceding quantization index (i.e. preceding residual level).
[0307] o The state transition procedure (i.e. surjective mapping) which determines the current state 220 based on the preceding state and the binary function (preferably a parity check) of the preceding quantization index is indicated by the state transition table 106.
[0308] o The state transition table 106 is selected based on the actually selected quantization variant (which is indicated in the bitstream, e.g. by the quantization mode information 180). However, for a given state transition table 106, the state transition procedure (e.g. the used binary function) is independent of the selected quantizer variant.
[0309] • Support Figure 1 The two scalar quantizers (e.g. quantizer 1511 and quantizer 1512) shown, and the quantizer 153 used for the quantization of the current transform coefficient is uniquely determined by the quantizer state:
[0310] o The same mapping from the state 220 to the quantizer identifier (0 or 1) (i.e. first mapping 200) is used for all supported quantization variants. In a preferred version, the quantizer identifier is given by qld = state & 1
[0311]
[0312] o The admissible reconstruction levels 142 for the first quantizer 1511 (qid equal to 0) comprise all even integer multiples of the quantization step size 144.
[0313] o The admissible reconstruction levels 142 for the second quantizer 1512 (qid equal to 1) comprise all odd integer multiples of the quantization step size 144 and additionally comprise the reconstruction level equal to 0.
[0314] • see e.g. Figure 9 The quantization index (i.e. the residual level 142) is encoded using binary arithmetic coding, wherein the binarization 300 comprises at least one binary symbol (preferably the significance flag 92, which indicates whether the quantization index is zero or non-zero), for which the selected probability model (or context model 242) depends on the quantization state (i.e. the current transform state 220) of the corresponding transform coefficient (i.e. the current residual level).
[0315] Thus, the same method for deriving the context model 242 is used for all supported quantization variants. This means that the context model derivation for the corresponding binary symbol depends on the state 220, but it does not depend on the selected quantization variant, i.e. is independent of the quantization mode information 180.
[0316] In the following, specific embodiments will be described, which support independent quantization, dependent quantization with 4 states, and dependent quantization with 8 states.
[0317] 5.2.1 High-level signaling
[0318] According to an embodiment, as shown in Figure 8 The media signal is a video, and the quantization mode information 180 in the data stream 14 comprises a first syntax element 182 (e.g. pps_dep_quant_enabled_idc) indicating whether the video or a portion of the video is subject to dependent quantization or not. The quantization mode information 180 in the data stream 14 comprises a second syntax element 184 (e.g. pic_dep_quant_idc) controlling which of the set 159 of default transform tables 106a, b, c is selected as one of the default transform tables 106 for the video or the portion of the video.
[0319] In other words, the media signal is a video, and the quantization mode information 180 in the data stream 14 includes a first syntax element 182 (e.g., pps_dep_quant_enabled_idc) that indicates whether the video or a portion of the video is to be updated 158 with which of the surjective mappings and which of the cardinalities l, m, n of the set 161 of one or more transition states is to be selected for use in the update 158 via a second syntax element 184 (e.g., pic_dep_quant_idc) of the quantization mode information 180 in the data stream 14.
[0320] According to an embodiment, in case the portion is a video, the first syntax element 182 is contained within a video parameter set in the data stream 14, and the second syntax element 184 controls selection of one of the set 159 of default transform tables 106 in units of image sequence, picture, tile, slice, coding tree block, coding block, or residual transform block. Alternatively, in case the portion is an image sequence, the first syntax element 182 is contained within a sequence parameter set in the data stream 14, and the second syntax element 184 controls selection of one of the set 159 of default transform tables 106 in units of picture, tile, slice, coding tree block, coding block, or residual transform block. Alternatively, in case the portion is one or more pictures, the first syntax element 182 is contained within a picture parameter set in the data stream 14, and the second syntax element 184 controls selection of one of the set 159 of default transform tables 106 in units of picture, tile, slice, coding tree block, coding block, or residual transform block. Alternatively, in case the portion is one picture, the first syntax element 182 is contained within a picture header in the data stream 14, and the second syntax element 184 controls selection of one of the set 159 of default transform tables 106 in units of tile, slice, coding tree block, coding block, or residual transform block. Alternatively, in case the portion is one coding tree block, the first syntax element 182 is contained within a parameter set in the data stream 14, and the second syntax element 184 controls selection of one of the set 159 of default transform tables 106 in units of coding block, or residual transform block. In case of a coding tree block, the selection can be controlled in units of, for example, a unit in which a picture is uniformly pre-subdivided before each coding tree block is split into coding blocks by means of recursive multi-tree subdivisioning. In case of a coding block, the selection can be controlled in units of, for example, a unit in which an intra / inter mode decision is made.
[0321] In other words, in the case where the portion is a video, the first syntax element 182 is contained within a video parameter set in the data stream 14, and the second syntax element controls the update 158 in units of a picture sequence, a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block. Alternatively, in the case where the portion is a picture sequence, the first syntax element 182 is contained within a sequence parameter set in the data stream 14, and the second syntax element controls the update 158 in units of a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block. Alternatively, in the case where the portion is one or more pictures, the first syntax element 182 is contained within a picture parameter set in the data stream 14, and the second syntax element controls the update 158 in units of a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block. Alternatively, in the case where the portion is one picture, the first syntax element 182 is contained within a picture header in the data stream 14, and the second syntax element controls the update 158 in units of a tile, a slice, a coding tree block, a coding block, or a residual transform block. Alternatively, in the case where the portion is one coding tree block, the first syntax element 182 is contained within a parameter set in the data stream 14, and the second syntax element controls the update 158 in units of a coding block, or a residual transform block. In the case of a coding tree block, the update 158 can be controlled in units of, for example, a unit in which a picture is uniformly pre-subdivided before each coding tree block is split into coding blocks by means of recursive multi-tree subdivision. In the case of a coding block, the update can be controlled in units of, for example, a unit in which an intra / inter mode determination is made.
[0322] In one version, the selected quantization method is indicated by a syntax element (i.e., the second syntax element 184) encoded in a picture header. This syntax element 184 can be referred to as pic_dep_quant_idc, and can take on one of the following three values, for example, indicating which of the surjective mappings is to be selected and which of the cardinalities l, m, n of the set 161 of one or more transition states is to be used for the update 158 (e.g., indicating which of the set 159 of default transform tables 106a, b, c is to be selected as one default transform table 106 for the portion of the video):
[0323] • a value of 0 indicates that conventional independent scalar quantization is to be used for the current picture (e.g., indicating that the set 161a has a cardinality l equal to to 1);
[0324] • a value of 1 indicates that dependent quantization with 4 states is to be used for the current picture (e.g., indicating that the set 161b has a cardinality m equal to 4);
[0325] • value 2 indicates that dependent quantization with 8 states is used for the current picture (e.g. indicates that the set 161c has a cardinality n equal to 8).
[0326] Furthermore, similarly to VTM-7, the presence of the picture header syntax element is indicated by the picture parameter set (PPS, picture parameter set) syntax element pps_dep_quant_enabled_idc. This syntax element (i.e. the first syntax element 182) can be coded with a fixed length code of 2 bits and can have the following semantics:
[0327] • value 0 indicates that pic_dep_quant_idc is present in the picture header (e.g. indicates that the update 158 is controlled within the video or said part of the video by the second syntax element 184);
[0328] • value 1 indicates that pic_dep_quant_idc is not present in the picture header but is inferred to be equal to 0 (independent scalar quantization) for all pictures referring to the picture parameter set (e.g. indicates that the cardinality I of the set 161a is equal to 1);
[0329] • value 2 indicates that pic_dep_quant_idc is not present in the picture header but is inferred to be equal to 1 (dependent quantization with 4 states) for all pictures referring to the picture parameter set (e.g. indicates that the cardinality m of the set 161b is equal to 4);
[0330] • value 3 indicates that pic_dep_quant_idc is not present in the picture header but is inferred to be equal to 2 (dependent quantization with 8 states) for all pictures referring to the picture parameter set (e.g. indicates that the cardinality n of the set 161c is equal to 8).
[0331] Similarly to VTM-7, if the picture parameter set flag constant_slice_header_params_enabled_flag is equal to 1, the pps_dep_quant_enabled_idc 182 is present in the picture parameter set. If the picture parameter set flag constant_slice_header_params_enabled_flag is equal to 0, the picture parameter set syntax element pps_dep_quant_enabled_idc 182 is not present in the picture parameter set but is inferred to be equal to 0 (in this case, the pic_dep_quant_idc 184 is coded in the picture header).
[0332] 5.2.2 State transition table
[0333] Based on the value (transmitted value or inferred value) of the picture header syntax element pic_dep_quant_idc (or any similar syntax element), a state transition table 106 can be selected. As shown, three state transition tables 106a, b, c can support 3 supported quantization variants. The state transition tables 106a, b, c are shown as examples (the actually used state transition table can deviate from these examples). Figure 10
[0334] According to one embodiment, the decoder 20 and / or the corresponding encoder only uses the state transition tables 106a and 106b.
[0335] 5.2.3 State transition
[0336] For each transform block 112, the initial state is set equal to 0.
[0337] The transform coefficients (i.e. residual levels 142) of a transform block 112 are processed in a predetermined encoding order 102. Given the state of the current transform coefficient (currState) (i.e. the current transform state 220) and the quantization index of the current transform coefficient (currLevel) (also referred to as transform coefficient level or residual level), the state of the next transform coefficient in the encoding / processing order (nextState) is derived as follows:
[0338] nextState = stateTransTab[currState][currLevel & 1],
[0339] e.g. representing an onto mapping, where stateTransTab represents the state transition table 106 determined by the selected quantization variant. Here, the parity of the quantization index (currLevel & 1) determines the path variable (see above state transition tables). In case of using the parity, any other binary function of the quantization index (see above) can also be used.
[0340] 5.2.4 Reconstruction of transform coefficients
[0341] The quantizer 153 for the current transform coefficient (i.e. the current residual level) is uniquely indicated by the value of the state variable (i.e. the current transform state 220). The quantizer identifier Qld is derived as follows:
[0342] Qld = state & 1,
[0343] denotes the first mapping 200, where & denotes the bitwise "and" operator. Quantizer 1511 with Qld = 0 includes even integer multiples of quantization step size 144 as admissible reconstruction levels 142. Quantizer 1512 with Qld = 1 includes odd integer multiples of quantization step size 144 as admissible reconstruction levels 142, as well as the value equal to 0.
[0344] A preferred example of a reconstruction procedure for transform coefficients of a single transform block is shown in Figure 11 in pseudo code using C language form.
[0345] Figure 11 Pseudo code is shown that illustrates a reconstruction procedure for transform coefficients of a transform block 112. Array level denotes the transmitted transform coefficient levels (i.e., the transmitted residual levels 142 (quantization indices)) for the transform block 112, and array trec denotes the corresponding reconstructed transform coefficients 104. Two-dimensional table state_trans_table denotes the state transition table 106 (e.g., selected based on a high-level syntax element (e.g., dep_quant_idc)).
[0346] In the pseudo code of Figure 11 , index k indicates the reconstruction order 102 of the transform coefficients (i.e., the residual levels 142). It should be noted that in the example code, index k is decremented in the reconstruction order. The index of the last transform coefficient is equal to k = 0. First index kstart indicates the reconstruction index (or more precisely, the inverse reconstruction index) of the first reconstructed transform coefficient. Variable kstart can be set equal to the number of transform coefficients in the transform block minus 1, or it can be set equal to the index of the first non-zero quantization index in the encoding / reconstruction order 102 (e.g., if the position of the first non-zero quantization index is transmitted in the applied entropy coding method). In the latter case, all preceding transform coefficients (in case of index k > kstart) are inferred to be equal to 0. The quantization index (i.e., the residual level 142) is denoted by level[k], and the associated reconstructed transform coefficient (i.e., the inverse quantized residual level 104) is denoted by trec[k]. The state variable (i.e., the current transition state 220) is denoted by state. The state transition table 106 is denoted by the two-dimensional array state_trans_table[][]. In case of using the table state_trans_table[][] for determining the next state, an arithmetic operation yielding the same result can be used instead.
[0347] quant_step_size[k] represents the quantization step size 144 for the transform coefficient at index k. It is noted that different quantization step sizes 144 can be used for different transform coefficient positions k (as given by the combination of the quantization weight matrix and the block quantization parameter). As mentioned above, the multiplication by a non-integer quantization step size is typically implemented using integer arithmetic, e.g.:
[0348] trec[k] = (n * scale[k] + add) » shift
[0349] where the quantization step size 144 is essentially given by: k = scale[k] · 2 -shift .
[0350] 5.2.5 Entropy coding of the quantized indices (or transform coefficient levels)
[0351] In the preferred version (see Figure 9 ), the quantized indices are coded using a binary arithmetic coding similar to H.264 | MPEG-4 AVC or H.265 | MPEG-H HEVC.
[0352] For this purpose, the non-binary quantized indices 90 are first mapped to a series of binary decisions, which are often referred to as bins. The quantized indices are transmitted in absolute values and, for absolute values larger than 0, also in sign. In the preferred version, the same binarization as in VTM-7 is used.
[0353]
[0354] The following binary and non-binary syntax elements are transmitted:
[0355] • sig_flag 92 indicates whether the absolute value |q| of the transform coefficient level is larger than 0;
[0356] • If sig_flag 92 is equal to 1, gt1_flag 96 indicates whether the absolute value |q| of the transform coefficient level is larger than 1;
[0357] • If gt1_flag 96 is equal to 1, par_flag 94 indicates the parity of the absolute value |q| of the transform coefficient level and gt3_flag 98 indicates whether the absolute value |q| of the transform coefficient level is larger than 3;
[0358] • If gt3_flag 98 is equal to 1, the non-binary value rem 99 indicates the remainder of the absolute level |q|. This syntax element is transmitted using a Golomb-Rice code in bypass mode of the arithmetic coder.
[0359] A non-existing syntax element is inferred to be equal to 0. At the decoder side, the absolute value of a transform coefficient level is reconstructed as follows:
[0360] |q| = sig_flag + gt1_flag + par_flag + 2*(gt3_flag + rem)
[0361] For a non-zero transform coefficient level (indicated by sig_flag 92 being equal to 1), then an additional sign_flag is transmitted in bypass mode, which represents the sign of the transform coefficient level.
[0362] The coding order 102 of the binary symbols and the context model are basically the same as in VTM-7, except for the two aspects described below. In addition, the same concept of worst-case usage to limit the number of context-coded binary symbols is used.
[0363] Figure 9 An initial value range 90 for the absolute value of a residual level 142 is shown. This initial value range 90 can include all integer values between zero and some maximum value. The initial value range 90 can also be an interval open to larger values. The number of integer values in the initial value range 90 need not be a power of two. In addition, Figure 9 The various binary symbol types involved in representing an individual residual level 142, i.e. in representing its absolute value, are shown. There is a sig_flag 92 type, which indicates whether the absolute value of a particular residual level is zero or non-zero. That is, the sig_flag 92 bisects the initial value range 90 into two sub-portions; i.e. one portion includes only zero, while the other portion includes all other possible values. That is, if the latter is exactly zero (as Figure 9 The non-zero values of the initial value range 90 form a value range 96, which is further bisected by a binary symbol type gt1_flag 96; i.e. one portion includes only one, while the other portion includes all other possible values. The next flag, i.e. par_flag 94, divides the resulting value range (including all values greater than one) into one portion of odd values and another portion of even values. The par_flag 94 need only exist for a particular residual level 142 if the latter is greater than one. The par_flag 94 does not create uniqueness with respect to one of the halves into which it also bisects the value range 94. It indicates one of the halves as the next resulting (recursively defined) value range, and thus the next flag, i.e. gt3_flag 98, further bisects this resulting value range after the par_flag 94. As Figure 9As shown below, this means that if the latter is exactly zero, only the sig_flag 92 is encoded for the residual level of a particular transform coefficient; and if the latter is exactly one, the sig_flag 92 and the gt1_flag 96 are encoded for the residual level of a particular transform coefficient. If the latter falls exactly into the value interval including absolute values 2 to 5, all the flags sig_flag 92, gt1_flag 96, par_flag 94 and gt3_flag 98 are encoded to represent the particular residual level of a particular transform coefficient; and in addition, the rest 99 is encoded for the residual level of a transform coefficient whose absolute value falls into the rest interval outside the initial value domain 90.
[0364] 5.2.5.1 Context selection for significant value flags
[0365] Figure 9 An embodiment of context-adaptive binary arithmetic coding (CABAC) for the significant value bin 92 is shown. According to an embodiment, the decoder 20 (shown in Figure 8 ) is configured to use context-adaptive binary arithmetic decoding 240 of a binarization 300 of the current residual level to decode 210 the current residual level from the data stream 14, the binarization 300 comprising the significant value bin 92 indicating whether the current residual level is zero or non-zero. The decoder 20 can be configured to select a first context 242 for decoding the first bin (i.e. the significant value bin 92) of the current residual level depending on a value 230 obtained by mapping a further predetermined number of bits (e.g. LSBs) of the current transform state 220 using a second mapping 202. The second mapping 202 can be equal regardless of the quantization mode information 18, e.g. regardless of which of the default transform tables 106a, b, c is selected as the transform table 106 to be used. The further predetermined number of bits represents one or more bits of the current transform state 220, e.g. located at one or more predetermined bit positions (e.g. LSBs or bits located at the second but not the last significant bit position).
[0366] According to an embodiment, the further predetermined number of bits is two and represents, e.g., the LSB (last significant bit) and the bit located at the second but not the last significant bit position. The value 230 can be derived from the second mapping 202 which is max(state-1, 0) or max(0, (state & 3) - 1) as shown in section 5.1. Alternative mapping approaches are possible.
[0367] For the significance flag (i.e., the significance binarization 92), the selected probability model (i.e., the context model or context 242) depends, for example, on:
[0368] • whether the current transform block 112 is a luma or chroma transform block;
[0369] • the conversion state of dependent quantization;
[0370] • the x and y coordinates of the current transform coefficient (the current residual level represents the residual level of the current transform coefficient);
[0371] • the absolute residual values of the partial reconstruction in the local neighboring region (after the first pass 601).
[0372] The variable state represents the state for the current transform coefficient (i.e., the current conversion state 220). As explained above, the state 220 for the current transform coefficient is given by the state for the previous coefficient in the encoding order 102 and a parity check (or, more generally, a binary function) of the previous transform coefficient level.
[0373] In case x and y are the coordinates of the current transform coefficient within the transform block 112, let diag = x + y be the diagonal position of the current coefficient. Given the diagonal position, then the diagonal class index dsig is derived as follows:
[0374] dsig = (diag < 2? 2 : (diag < 5? 1 : 0)) for luma transform blocks
[0375] and
[0376] dsig = (diag < 2? 1 : 0) for chroma transform blocks
[0377] The ternary operator (c? a : b) represents an if-then-else statement. If the condition c is true, then the value a is used, otherwise the value b is used (c is false).
[0378] The context 242 can depend on the absolute values of the partial reconstruction in the local neighboring region. As in VTM-7, the local neighboring region can be given by the template T 132, e.g., as shown in Figure 7 But other templates 132 are possible. The used samples 132 can also depend on whether a luma or chroma block is coded. Let sumAbs be the sum of the absolute values of the partial reconstruction in the template T 132 (after the first pass 601):
[0379]
[0380] where abs1[k] denotes the absolute level of the partial reconstruction of scan index k after the first pass. In case of the binarization of VTM-7, it is given as follows:
[0381] abs1[k] = sig_flag[k] + gt1_flag[k] + par_flag[k] + 2*gt3_flag[k].
[0382] It is noted that abs1[k] is equal to the value coeff[k] obtained after the first scan pass 601 (see above).
[0383] It is assumed that the possible probability model (i.e. the context 242) for sig_flag 92 is a structure of a one-dimensional array. And let ctxIdSig be an index identifying the used probability model 242. The context index ctxIdSig (e.g. an index representing the context 242) can be derived 110 as follows:
[0384] If the transform block is a luma block,
[0385] ctxIdSig = min((sumAbs + 1) » 1, 3) + 4*dsig + 12*state2CtxSet(state & 3)
[0386] If the transform block is a chroma block,
[0387] ctxIdSig = 36 + min((sumAbs + 1) » 1, 3) + 4*dsig + 8*state2CtxSet(state & 3)
[0388] Here, the operator "»" denotes a bit shift to the right (two's complement arithmetic). This operation is the same as dividing by two and rounding the result down to the next integer.
[0389] It is noted that first the state (i.e. the current conversion state 220) is mapped to a value between 0 and 3, inclusive. This can be done by applying a bitwise "and" with 3 (i.e. state & 3). Then, the function state2CtxSet(..) maps this value to one of the values 230 (0, 1, 2). In a preferred embodiment, the function is given as follows:
[0390]
[0391] According to this embodiment, the second mapping 202 can be understood as two consecutive mappings, where first the current conversion state 220 is mapped to a first value, and then this first value is mapped to the value 230 (the selection of the context 242 depends on this value 230).
[0392] It is noted that the context index depends on the state 220, but it does not depend on the used quantization method 180.
[0393] It is noted that other structures of the context model (i.e. different contexts 242) are possible. But in any case, the chosen probability model (i.e. the context 242) for encoding the sig_flag 92 depends on:
[0394] • whether a luma or chroma block is encoded;
[0395] • the state variable 220 (but the same function is used for all supported quantization variants, for example);
[0396] • the diagonal class dsig;
[0397] • the sum of absolute values of the partial reconstruction within the local template 132 (after the first pass 601), in particular min((sumAbs + 1) » 1, 3).
[0398] 5.2.5.2 Mapping parameters for transform coefficient levels encoded by-pass
[0399] According to an implementation, the media signal is a video, and the decoder 20 (shown in Figure 8 ) is configured to decode 210 the residual levels 142 from the data stream 14 by decoding the residual levels 142 for one image block 84 (e.g. representing a transform coefficient block 112) before the residual levels 142 for another image block 84, as shown in Figure 6 The decoder 20 can be configured to decode 210 the binarized 300 binary codes for the residual levels 142 for the one image block 84 sequentially among one or more scans 60 (i.e. along the scan order 102), and to use context-adaptive binary arithmetic decoding for a predetermined number of leading binary codes of the binarized 300 binary codes for the residual levels 142, and to use the equal-probability by-pass mode for the following binary codes of the binarized 300 binary codes for the residual levels 142. As shown in Figure 9 For example, the significant value binary codes 92, the parity binary codes 94, the greater than one binary codes 96, and the greater than three binary codes 98 are decoded in the first scan 601 (i.e. the first pass), and the remaining binary codes 99 are encoded in the second scan 602. Alternatively, the residual levels 142 can be decoded in one scan 60 or three or more scans 60. Figure 9An exemplary transform coefficient block 112 is shown, where the transform coefficients are scanned along the encoding order 102 and start with the first non-zero residual level 120. Because the residual level 304 represents the last residual level that does not exceed the predetermined number of leading bins, the binarization of the residual level 146 is decoded using CABAC up to the residual level 304, and the binarization of the residual level 144 is encoded using the equiprobable bypass mode. The decoder 20 can be configured to determine the binarization 300 of the predetermined residual level 145 by using the third mapping 118 to map a further predetermined number of bits (e.g., LSBs) of the conversion state 220 (i.e., the conversion state related to the predetermined residual level 145) to the binarization parameter 302, and performing the determination of the binarization 300 of the predetermined residual level 145 in dependence on the binarization parameter 302, to determine the binarization 300 of the predetermined residual level 145 (where the predetermined number 304 of bins is exhausted). According to an embodiment, the determination of the binarization can be performed by the decoder 20 for all residual levels 144 (where the predetermined number 304 of bins is exhausted). The bits of the conversion state 220 are, for example, one or more bits located at one or more predetermined bit positions of the conversion state 220 (e.g., LSBs, or bits located at the second but not the last significant bit position).
[0400] According to an embodiment, the decoder 20 is configured to determine the binarization 300 for the predetermined residual level 145 in dependence on the binarization parameter 302, such that the binarization 300 is a default binarization (e.g., Golomb-Rice code) modified by re-directing one binarization code associated with the predetermined level (which depends on the binarization parameter 302, e.g., posO) to be associated with level zero, and re-directing one or more binarization codes associated with one or more levels (from zero to the predetermined level minus one) to be associated with one or more levels (from one to the predetermined level). The modification of the default binarization can be implemented by the equation |q| = (abs_level == posO? 0 : (abs_level < posO? abs_level + 1 : abs_level)), as described below.
[0401] According to an embodiment, the further predetermined number of bits is one.
[0402] Similar to VTM-7, the absolute levels (i.e., residual levels 144) for which no data is encoded in the first scan pass 601 are entirely encoded in bypass mode (i.e., the equiprobable bypass mode). It is encoded using the same class of codes as for the remaining part rem99. Based on the encoded syntax elements (see below), the Rice parameter and the variable posO (i.e., the binarization parameter 302) can be determined. These absolute levels |q| 144 are, for example, not directly encoded, but first mapped to a syntax element abs_level, which is then encoded using a parameter code.
[0403] Depending on the variable posO 302, the absolute values |q| 144 can be mapped to the syntax element abs_level. This can be expressed as:
[0404] abs_level = (|q| == 0? posO : (|q| <= posO? |q| - 1 : |q|)
[0405] At the decoder side, mapping the syntax element abs_level to the absolute value |q| can be expressed as follows:
[0406] |q| = (abs_level == posO? 0 : (abs_level < posO? abs_level + 1 : abs_level)
[0407] The Rice parameter RPabs for encoding abs_level is derived as follows:
[0408] RPabs = tabRPabs[min(sumAbs, 31)],
[0409] where sumAbs denotes the sum of the absolute quantization indices within the local template 132; i.e., the sum of the absolute levels of the residual levels (see above).
[0410] The variable posO 302 can depend on both: the sum of the reconstructed absolute values in the local template 132 and the state variable 220. The variable posO 302 is derived 118 from the quantization state 220 and the Rice parameter Rpabs according to:
[0411] posO = (1 + (state & 1)) « RPabs.
[0412] It should be noted that the variable posO 302 does indeed depend on the state 220, but not on the used quantization method 180.
[0413] The equation posO = (1 + (state & 1)) « RPabs can represent the third mapping 118, which maps a further predetermined number of bits of the transition state 220 to the binarization parameter posO 302, where (state & 1) indicates that the LSB is mapped to the binarization parameter posO. The further predetermined number of bits is one.
[0414] Alternatively, the equation posO = (state < 2? 1 : 2) « RPabs can represent the third mapping 118, which maps a further predetermined number of bits of the transition state 220 to the binarization parameter posO 302, where (state < 2? 1 : 2) indicates that the second but not the least significant bit position of the transition state 220 is mapped to the binarization parameter posO (see last item of section 5.1). The further predetermined number of bits is one.
[0415] 5.3 Aspects of other embodiments
[0416] Several aspects of the above-described embodiments can be modified. In the following, some of these aspects will be indicated:
[0417] • The number of supported quantization variants and / or the actual state transition tables 106 for these variants can be modified;
[0418] • Different mechanisms to indicate the selected quantization variant can be used;
[0419] • Different binarization functions of the quantization index can be used for the state transition procedure. This means that the state transition can be represented as:
[0420] nextState = stateTransTable[currState] [binFun(currLevel)],
[0421] where binFunc() represents any binarization function (i.e. a function with possible function values of 0 and 1).
[0422] • Different binarization procedures can be used for encoding the quantization index. In particular, it is possible to use state-dependent binarization, where a majority binarization is supported and the used binarization depends on the quantization state (and possibly other parameters). In this case, the selection of the binarization can depend on the value of the state variable, but is independent of the selected quantization method. For example, two binarizations can be supported and the selected binarization can depend on the value of (state & 1).
[0423] • Different context models can be used for any binary code. In this case, the context for one or more binary codes can still be dependent on the state variable, but it can be independent of the selected quantization method.
[0424] • Different methods to determine the Rice parameter and / or the quantized index of the bypass coded can be used.
[0425] • Different methods to derive the parameter pos0 that encodes the quantized index of the bypass coded can be used. But in the context of the present application, the determination of the parameter pos0 does not depend on the selected quantization method.
[0426] 5.4 Further implementation
[0427] In the above described implementations, the transform coding is used to encode the prediction residual, and the quantization scheme / mode is selected by the encoder and signaled to the decoder in order to quantize / dequantize the transform coefficients. Entropy coding of the obtained quantized indices (i.e. context adaptive arithmetic coding) can be used. At the decoder side, the set of reconstructed samples is obtained by decoding of the corresponding quantized indices and dequantization of the corresponding transform coefficients, such that the inverse transform yields the prediction residual in the spatial domain. However, the above described implementations can be modified to yield other implementations, e.g. in terms of encoding other residual values (e.g. residual samples in the spatial domain). Instead of using the above described implementations for video coding, the above described concepts can be used to process another media signal, e.g. a sound signal. The main goal of the implementations described below is the lossy coding of blocks of prediction error samples in image and video codecs, but the implementations can also be applied to lossy coding in other fields. In particular, there is no restriction on the set of samples forming a rectangular block.
[0428] While not being limited to video coding using transform-based prediction residual coding, in order to form an example of an encoding architecture that can establish an implementation for encoding and decoding of a transform representation of a block of samples, a description of a video encoder and a video decoder of a block-based predictive codec for image coding of a video is provided in the following. The video encoder and the video decoder are described with reference to Figures 12 to 14 the implementations of the present application as described above can immediately be established in the video encoder and the video decoder of Figure 12 and 13, although the above described implementations can also be used to form a video encoder and a video decoder operating within an encoding architecture not according to Figure 12 and 13.
[0429] Figure 12A device for predictively encoding a video 11 composed of a sequence of pictures 12 into a data stream 14 is shown. For this purpose, block-wise predictive encoding is used. Furthermore, transform-based residual encoding is exemplarily used. The device (or encoder) is denoted by reference sign 10. Figure 13 A corresponding decoder 20 is shown; i.e. a device 20 is configured to predictively decode from the data stream 14 a video 11' composed of pictures 12' in picture blocks, here also exemplarily using transform-based residual decoding, wherein the pictures 12' and the video 11' reconstructed by the decoder 20, respectively, are denoted by apostrophes and deviate from the pictures 12 originally encoded by the device 10 in terms of an encoding loss introduced by quantization of the prediction residual signal. Although Figure 12 and Figure 13 Transform-based prediction residual encoding is exemplarily used, but embodiments of the present application are not limited to this prediction residual encoding. For further details with respect to Figure 12 and 13, it is also true that the following will be described.
[0430] The encoder 10 is configured to spatial-to-spectral transform the prediction residual signal and to encode the thus obtained prediction residual signal into the data stream 14. Likewise, the decoder 20 is configured to decode the prediction residual signal from the data stream 14 and to spectral-to-spatial transform the thus obtained prediction residual signal.
[0431] The encoder 10 can internally comprise a prediction residual signal former 22 which generates a prediction residual 24 to measure a deviation of a prediction signal 26 from an original signal, i.e. the video 11 or the current picture 12. The prediction residual signal former 22 can for example be a subtractor which subtracts the prediction signal from the original signal, i.e. the current picture 12. Then, the encoder 10 further comprises a transformer 28 which spatial-to-spectral transforms the prediction residual signal 24 to obtain a spectral domain prediction residual signal 24', which is then quantized by a quantizer 32 also comprised in the encoder 10. The thus quantized prediction residual signal 24" is encoded into the bit stream 14. For this purpose, the encoder 10 can optionally comprise an entropy encoder 34 which entropy encodes the prediction residual signal into the data stream 14. The prediction residual 26 is generated by a prediction stage 36 of the encoder 10 based on the prediction residual signal 24" decoded into and decodable from the data stream 14. For this purpose, as Figure 12As shown, the prediction stage 36 can internally comprise a dequantizer 38 which dequantizes the prediction residual signal 24" to obtain a spectral domain prediction residual signal 24"' (which corresponds to the signal 24' except for quantization losses); followed by an inverse transformer 40 which inverse transforms (i.e. spectral to spatial transform) the latter prediction residual signal 24"' to obtain a prediction residual signal 24"" (which corresponds to the original prediction residual signal 24 except for quantization losses). A combiner 42 of the prediction stage 36 then recombines (e.g. by additive means) the prediction signal 26 and the prediction residual signal 24"" to obtain a reconstructed signal 46 (i.e. a reconstruction of the original signal 12). The reconstructed signal 46 can correspond to the signal 12'.
[0432] By using, for example, spatial prediction (i.e. intra picture prediction) and / or temporal prediction (i.e. inter picture prediction), the prediction module 44 of the prediction stage 36 then generates the prediction signal 26 on the basis of the signal 46. Details regarding this aspect are described below.
[0433] Likewise, the decoder 20 can internally be composed of means corresponding to the means of the prediction stage 36 (and in a corresponding interconnection). In particular, the entropy decoder 50 of the decoder 20 can entropy decode the quantized spectral domain prediction residual signal 24" from the data stream; in accordance with which, the dequantizer 52, the inverse transformer 54, the combiner 56 and the prediction module 58 of the decoder 20, in the interconnection and cooperation as described above with regard to the modules of the prediction stage 36, restore a reconstructed signal on the basis of the prediction residual signal 24" to cause the output of the combiner 56 to generate the reconstructed signal (i.e. the video 11' or a current picture 12' thereof), as Figure 13 shown.
[0434] Although not specifically described above, it is clear that the encoder 10 can set some encoding parameters (including, for example, prediction modes, motion parameters, and the like) in accordance with some optimization schemes (e.g., optimizing some rate and distortion related criteria (i.e., encoding cost) and / or using some rate control approaches). As described in more detail below, the encoder 10 and the decoder 20 and the corresponding modules 44, 58, respectively, support different prediction modes, for example, intra-picture encoding modes and inter-picture encoding modes, which are based on predictions of blocks of pictures composed in manners described below to form a set or pool of original prediction modes. The granularity at which the encoder and the decoder switch between these prediction compositions can correspond to a subdivision of the pictures 12 and 12', respectively, into blocks. It is noted that some of these blocks can be blocks that are only intra-coded and some blocks can be blocks that are only inter-coded, and, optionally, even some blocks can be blocks that are obtained using both intra- and inter-coding, details of which are presented below. According to an intra-picture encoding mode, the prediction signal of a block is obtained based on spatially, already encoded / decoded neighboring regions of the individual block. A plurality of intra-picture encoding sub-modes can exist, among which a "quasi" indicates an intra-prediction parameter. Directional or angular intra-picture encoding sub-modes can exist according to which the prediction signal of an individual block is filled by extrapolating sample values of the neighboring regions along a specific direction that is specific for the individual directional intra-picture encoding sub-mode. The intra-picture encoding sub-modes can also include, for example, one or more other sub-modes, for example: a DC encoding mode according to which the prediction signal of an individual block specifies a DC value for all samples within the individual block; and / or a planar intra-picture encoding mode according to which the prediction signal of an individual block is approximated or determined to be a spatial distribution of sample values at sample positions of the individual block described by a two-dimensional linear function, the inclination and the offset of the plane defined by the two-dimensional linear function being derived based on neighboring samples. In contrast thereto, according to an inter-picture prediction mode, the prediction signal of a block can be obtained, for example, by temporally predicting the block. For the parameterization of the inter-picture prediction mode, a motion vector can be signaled in the data stream, the motion vector indicating a spatial displacement of a portion of a previously encoded picture of the video 11 from which the prediction signal is here sampled to obtain the prediction signal of an individual block. This means that, in addition to the residual signal encoding comprised in the data stream 14 (e.g., the entropy encoding of the transform coefficient levels representing the quantized spectral domain prediction residual signal 24"), the data stream 14 can have encoded therein prediction related parameters to specify a block prediction mode, prediction parameters for the specified prediction mode (e.g., motion parameters for the inter-picture prediction mode), and, optionally, other parameters controlling the composition of the final prediction signal of a block using the specified prediction mode and prediction parameters, as described in more detail below.Furthermore, the data stream can include control and signaling parameters for the subdivision of the image 12 and 12' into blocks, respectively. The decoder 20 uses these parameters to perform the following operations: subdividing the image in the same way as the encoder, assigning the same prediction mode and parameters to the blocks, and performing the same prediction to produce the same prediction signal.
[0435] Figure 14 One aspect illustrates the relationship between the reconstructed signal (i.e., the reconstructed image 12') and the combination of the signaled prediction residual signal 24" and the prediction signal 26 in the data stream; the other aspect illustrates the combination of the prediction residual signal 24" and the prediction signal 26. As mentioned above, the combination can be an additive manner. The prediction signal 26 is displayed in Figure 14 The subdivision of the image region into blocks 80 of different sizes is shown in Figure 14 The manner of mixing is shown, in which the image region is first subdivided into columns and rows of tree root blocks, and then the tree root blocks are subdivided according to a recursive multi-tree type of subdivision to produce the blocks 80.
[0436] Figure 14 The prediction residual signal 24" in the data stream also shows the subdivision of the image region into blocks 84. These blocks can be referred to as transform blocks, to distinguish from the coding blocks 80. In fact, Figure 14 The encoder 10 and the decoder 20 can use two different subdivisions of the image 12 and 12', respectively, into blocks; i.e., the subdivision into coding blocks 80 and another subdivision into blocks 84. The two subdivisions can be the same (i.e., each block 80 can also form a transform block 84), and vice versa; but Figure 14It is shown that, for example, the subdivision into transform blocks 84 forms an extension of the subdivision into blocks 80 such that any boundary between two blocks 80 coincides with a boundary between two transform blocks 84, or, alternatively, each block 80 coincides with one of the transform blocks 84, or with a cluster of transform blocks 84. However, the determination or selection of the subdivisions can also be independent from each other such that the transform blocks 84 can alternatively span across block boundaries between blocks 80. Similar remarks as made with respect to the subdivision into blocks 80 hold for the subdivision into transform blocks 84; i.e., the blocks 84 can be the result of a regular subdivision of the image area into blocks arranged in columns and rows, the result of a recursive multi-tree subdivision of the image area, or a combination thereof, or any other partitioning. Incidentally, it is noted that the blocks 80 and 84 are not restricted to be square, rectangular, or any other shape. Furthermore, the subdivision of the current image 12 into blocks 80, where the prediction signal is formed, and the subdivision of the current image 12 into blocks 84, where the prediction residual is encoded, are not the only subdivisions available for encoding / decoding. These subdivisions form the granularity at which the prediction signal determination and the residual encoding are performed; but, first, the residual encoding can alternatively be done without subdivision; and, second, for other granularities than these subdivisions, the encoder and the decoder can set certain encoding parameters, which can include some of the aforementioned parameters, such as prediction parameters, prediction signal composition control signals, and the like.
[0437] Figure 14 It is shown that, for example, the subdivision into transform blocks 84 forms an extension of the subdivision into blocks 80 such that any boundary between two blocks 80 coincides with a boundary between two transform blocks 84, or, alternatively, each block 80 coincides with one of the transform blocks 84, or with a cluster of transform blocks 84. However, the determination or selection of the subdivisions can also be independent from each other such that the transform blocks 84 can alternatively span across block boundaries between blocks 80. Similar remarks as made with respect to the subdivision into blocks 80 hold for the subdivision into transform blocks 84; i.e., the blocks 84 can be the result of a regular subdivision of the image area into blocks arranged in columns and rows, the result of a recursive multi-tree subdivision of the image area, or a combination thereof, or any other partitioning. Incidentally, it is noted that the blocks 80 and 84 are not restricted to be square, rectangular, or any other shape. Furthermore, the subdivision of the current image 12 into blocks 80, where the prediction signal is formed, and the subdivision of the current image 12 into blocks 84, where the prediction residual is encoded, are not the only subdivisions available for encoding / decoding. These subdivisions form the granularity at which the prediction signal determination and the residual encoding are performed; but, first, the residual encoding can alternatively be done without subdivision; and, second, for other granularities than these subdivisions, the encoder and the decoder can set certain encoding parameters, which can include some of the aforementioned parameters, such as prediction parameters, prediction signal composition control signals, and the like.
[0438] In Figure 14 Among others, the transform blocks 84 shall have the following meaning. The transformer 28 and the inverse transformer 54 perform their transformations in units of these transform blocks 84. For example, many codecs use some kind of DST or DCT for all transform blocks 84. Some codecs allow to omit the transform such that, for some of the transform blocks 84, the prediction residual signal is directly encoded in the spatial domain. The supported transforms can include one or more of the following:
[0439] o DCT-II (or DCT-III), where DCT denotes Discrete Cosine Transform
[0440] o DST-IV, where DST stands for Discrete Sine Transform.
[0441] o DCT-IV
[0442] o DST-VII
[0443] o Identity Transformation (IT)
[0444] The transform can be a separable transform. A second transform (which can be a non-separable transform) can be applied to the low-frequency range, for example, a subarray extending from the DC component to a certain intermediate frequency component, excluding DC and the highest frequency component. The encoder can decide to use such a second transform, and select one for a specific block, and signal this decision to the decoder.
[0445] Now, returning to the adaptive quantization mode, the decoder and encoder can operate as follows. See example. Figure 8 Decoder 20 can be configured to decode residual level 142 (e.g., as shown in the figure) from data stream 14. Figure 1 The quantization level of the transformation coefficient 100 shown represents the prediction residual 24”), and is sequentially arranged in the following manner (i.e., along as follows). Figure 6 The scan order shown 102) Dequantization 150 Residual level 142: Based on the current conversion state, from the set of default quantizers 151 (e.g., Figure 1 The decoder 20 selects a quantizer from sets 0 and 1; dequantizes the current residual level 154 using the quantizer to obtain a dequantized residual value 104; applies a binary function (e.g., a parity check function) 156 to the current residual level to obtain a characteristic 157 (i.e., one or zero) of the current residual level; and updates the current transition state 160 158 based on the characteristic 157 of the current residual level and using a transition table 106 for the next transition state 162 of the residual level to be used for the next dequantization. The decoder 20 uses the dequantized residual value 104 to reconstruct a media signal 170, such as an image or video. Based on the quantization mode information 180 included in data stream 14 (e.g., the example described in Section 5.1.3), decoder 20 can select a default transformation table 106 from a set 159 of default transformation tables 106a, b, c as the transformation table to be used. Each of the default transformation tables 106a, b, c represents a domain surjective mapping from a combination 163 of a set of one or more transformation states having the current residual level (see, for example...). Figure 10) to a set of one or more transition states 161, wherein the difference between the default transition tables 106a, b, c consists in the cardinality / , m, n of the set of one or more transition states. Irrespective of which of the default transition tables 106a, b, c is selected as the transition table 106 to be used, the decoder 20 is configured to perform the selection 152 of the quantizer by mapping (using a first mapping 108, e.g. see section 5.2.4) a predetermined number of LSBs of the current transition state (e.g. one, i.e. only the LSB) to a default quantizer; the predetermined number as well as the first mapping are equal (i.e. independent of the selected mode 180 or transition table 106a, b, c, respectively) irrespective of which of the default transition tables 106a, b, c is selected as the transition table 106 to be used. Decoding 210 the current residual level 142 from the data stream 14 can be done with a context adaptive binary arithmetic decoding of the binarization of the current residual level. The binarization can comprise a significant value bin (sig_flag) which indicates whether the current residual level is zero or non-zero, and a first context for decoding the bin of the current residual level can be selected by mapping (e.g. using a second mapping 110, as described in section 5.2.5.1) a further predetermined number (e.g. 2, which is taken by " & 3") of LSBs of the current transition state to a first context. Additionally or alternatively, before decoding 210 residual levels with respect to another image block, residual levels 142 with respect to one image block are decoded 210 (e.g. Figure 10 in block 84, the residual of which is determined by decoding 210 Figure 6The transform coefficient block 112 shown is subjected to inverse transform to obtain), whereas the residual levels 142 are decoded from the data stream 14, and the binarized bin codes for the residual levels 142 of one image block are decoded sequentially among one or more scans (see pass 114 in section 2.2.2), and the leading bin codes of a predetermined number (e.g. remRegBins) of binarized bin codes for the residual levels are decoded using context-adaptive binary arithmetic decoding, and the following bin codes (k >= startIdxBypass) of the binarized bin codes for the residual levels are decoded using the equal-probability bypass mode; for the bin codes of the predetermined number used up (k >= startIdxBypass), the binarization of the predetermined residual levels 145 is determined by mapping a further predetermined number (e.g. 1) of LSBs of the transform state 220 that inherently appears for the predetermined residual levels 145 using the third mapping 118 to a binarization parameter 302 (e.g. posO), and depending on the binarization parameter 302, the determination of the binarization of the predetermined residual levels is performed; the third mapping 118 is equal, no matter which default transform table 106a, b, c is selected as the transform table 106 to be used. Depending on the binarization parameter 302, the binarization of the predetermined residual levels 145 can be determined such that the binarization is a default binarization (e.g. Rice code) modified by redirecting one binarization code associated with the predetermined level (which depends on the binarization parameter) to be associated with level zero, and redirecting one or more binarization codes associated with one or more levels (from zero to the predetermined level minus one) to be associated with one or more levels (from one to the predetermined level).
[0446] Next, the decoder 20 is configured to reconstruct 170 the media signal using the dequantized residual values 104 by performing a prediction to obtain a predicted version of the media signal, and correcting the predicted version using the dequantized residual values 104; or reconstruct the media signal using the dequantized residual values 104 by performing a prediction to obtain a predicted version of the media signal, and inverse transforming the dequantized residual values 104 to obtain a media residual signal, and correcting the predicted version using the media residual signal.
[0447] In the following, additional embodiments and aspects of the present application will be described, which can be used alone or in combination with any of the features and functions and details described herein.
[0448] According to a first aspect, a decoder (20) can be configured to:
[0449] Decoding (210) from a data stream (14) residual levels (142) representing prediction residuals, and sequentially dequantizing (150) the residual levels (142) by:
[0450] Depending on a current conversion state (220), selecting (152) a quantizer (153) from a set of default quantizers (151);
[0451] Dequantizing (154) a current residual level using the quantizer (153) to obtain a dequantized residual value;
[0452] Applying (156) a binary function to the current residual level to obtain a characteristic (157) of the current residual level;
[0453] and
[0454] Depending on the characteristic (157) of the current residual level, updating (158) the conversion table (106) to be used with
[0455] the current conversion state (220),
[0456] Using the dequantized residual value (104) to reconstruct (170) a media signal,
[0457] Depending on quantization mode information (180) contained in the data stream (14), selecting a default conversion table from a set of default conversion tables (159) as the conversion table (106) to be used, wherein each of the default conversion tables represents a surjective mapping from a domain of a combination (163) of sets (161) of one or more conversion states to the sets (161) of one or more conversion states having the characteristic (157) of the current residual level, wherein the default conversion tables differ in a cardinality of the sets (161) of one or more conversion states, and
[0458] Irrespective of which default conversion table is selected as the conversion table (106) to be used, performing the selection (152) of the quantizer (153) by using a first mapping (200) mapping a predetermined number of bits of the current conversion state (220) to the default quantizers, wherein the predetermined number and the first mapping (200) are equal irrespective of which default conversion table is selected as the conversion table (106) to be used.
[0459] Depending on the second aspect (when referring back to the first aspect), in the decoder (20), the predetermined number of bits can be one.
[0460] According to a third aspect (when referring back to the first or second aspect), the decoder (20) can be configured to select, depending on the quantization mode information (180), a conversion table (106) to be used from a set (159) of default conversion tables (106a, b, c), and to manage the set (159) of default conversion tables as different parts of a unified conversion table, wherein for each default conversion table, each of the set (161) of one or more conversion states is indexed using a state index that is different from the state indices used for indexing any other default conversion table in the unified conversion table (e.g. the default conversion tables are disjoint from each other for the conversion states of the different parts of the unified conversion table).
[0461] According to a fourth aspect (when referring back to the third aspect), the decoder (20) can be configured to perform the updating (158) of the current conversion state (220) by looking up an entry in the unified conversion table, the entry corresponding to the combination of the characteristic (157) of the current residual level and the current conversion state (220).
[0462] According to a fifth aspect (when referring back to the third or fourth aspect), in the decoder (20), the quantization mode information (180) can comprise a syntax element in the data stream (14) indicating the starting conversion state used for dequantizing the first residual level in a reverse quantization order (e.g. performing the dequantization in sequence along this order), such that the syntax element indicates the conversion table (106) to be used from among the default conversion tables by fitting to the state index in the part of the unified conversion table only for the conversion table (106) to be used.
[0463] According to a sixth aspect (when referring back to any one of the first to fifth aspects), in the decoder (20), the set (151) of default quantizers can consist of:
[0464] a first default quantizer (1511) comprising an even integer multiple of the quantization step size (144) as a reconstruction level; and
[0465] a second default quantizer (1512) comprising an odd integer multiple of the quantization step size (144) and zero as reconstruction levels.
[0466] According to a seventh aspect (when referring back to any one of the first to sixth aspects), the decoder (20) can be configured to:
[0467] use context-adaptive binary arithmetic decoding of a binarization of the current residual level to decode the current residual level from the data stream (14), wherein the binarization comprises a significant value bin (92) representing whether the current residual level is zero or non-zero;
[0468] by using a second mapping (202) to map a further predetermined number of bits of the current conversion state (220) to a value, depending on the obtained value (230) to select a first context (242) of a first binarization code (92) used to decode the current residual level; the second mapping (144) is equal regardless of the quantization mode information (180).
[0469] According to an eighth aspect (when referring back to the seventh aspect), in the decoder (20), the further predetermined number of bits can be two.
[0470] According to a ninth aspect (when referring back to any one of the first to eighth aspects), in the decoder (20), the media signal can be a video, and the decoder (20) can be configured to:
[0471] decode (210) the residual levels (142) from the data stream (14) by:
[0472] decode the residual levels (142) for the image block (84) before the residual levels (142) for another image block (84);
[0473] decode the binarized binarization codes (92) for the image block (84) sequentially in one or more scans, and use context-adaptive binary arithmetic decoding for a predetermined number of leading binarization codes (92) of the binarized binarization codes (92) and use an equal probability bypass mode for subsequent binarization codes (92) of the residual levels (142);
[0474] determine the binarization of the predetermined residual level (142) where the predetermined number of binarization codes (92) is exhausted, by using a third mapping to map a further predetermined number of bits (e.g. one or more bits at one or more predetermined bit positions, e.g. LSBs or bits at the second but not the last significant bit position) of the conversion state (220) to a binarization parameter (302), and depending on the binarization parameter (302) to perform the determination of the binarization of the predetermined residual level (142); the third mapping is equal regardless of the quantization mode information (180).
[0475] According to a tenth aspect (when referring back to the ninth aspect), the decoder (20) can be configured to determine the binarization of the predetermined residual level (142) depending on the binarization parameter (302) such that the binarization is a default binarization modified by: redirecting binarization codes (92) depending on the binarization parameter (302) from being associated with the predetermined level to being associated with level zero, and redirecting one or more binarization codes (92) from being associated with one or more levels from zero to the predetermined level minus one to being associated with one or more levels from one to the predetermined level.
[0476] According to the eleventh aspect (when referring back to the ninth or tenth aspect), in the decoder (20), the further predetermined number of bits can be one.
[0477] According to the twelfth aspect (when referring back to the ninth aspect), in the decoder (20), the surjective mapping and the cardinality of the set of one or more transition states (161) can be selected between two or more of the following, depending on the quantization mode information (180):
[0478] the cardinality of the set of one or more transition states (161) is one;
[0479] the cardinality of the set of one or more transition states (161) is four;
[0480] the cardinality of the set of one or more transition states (161) is eight.
[0481] According to the thirteenth aspect (when referring back to any of the first to twelfth aspects), in the decoder (20), the media signal can be a video, and the decoder (20) can be configured to:
[0482] read the quantization mode information (180), wherein the quantization mode information (180) controls the updating (158) in one of the following ways:
[0483] only once for the video;
[0484] per image block (84);
[0485] on a per-image basis;
[0486] on a per-slice basis; and
[0487] on a per-image sequence basis.
[0488] According to the fourteenth aspect (when referring back to any of the first to thirteenth aspects), in the decoder (20), the media signal can be a video, and the quantization mode information (180) can comprise a first syntax element in the data stream (14) indicating whether the updating (158) for the video or a portion of the video is controlled within the video or the portion of the video by a second syntax element of the quantization mode information (180) in the data stream (14), or which one of a surjective mapping and which one of a cardinality of a set of one or more transition states (161) is to be selected for the updating (158).
[0489] According to the fifteenth aspect (when referring back to the fourteenth aspect), in the decoder (20), the first syntax element can be comprised in one of the following in the data stream (14):
[0490] a video parameter set, wherein the portion is the video and the second syntax element controls updating (158) in units of:
[0491] a sequence of pictures,
[0492] a picture,
[0493] a tile,
[0494] a slice,
[0495] a coding tree block (e.g., a unit in which a picture is uniformly pre- divided before each coding tree block is split into coding blocks by recursive multi- tree subdivision), a coding block (making intra / inter picture mode determinations),
[0496] a coding block (e.g., a unit in which, e.g., intra / inter picture mode determinations are made), or
[0497] a residual transform block;
[0498] a sequence parameter set, wherein the portion is a sequence of pictures and the second syntax element controls updating (158) in units of:
[0499] a picture,
[0500] a tile,
[0501] a slice,
[0502] a coding tree block (e.g., a unit in which a picture is uniformly pre- divided before each coding tree block is split into coding blocks by recursive multi- tree subdivision), a coding block (making intra / inter picture mode determinations),
[0503] a coding block (e.g., a unit in which, e.g., intra / inter picture mode determinations are made), or
[0504] a residual transform block;
[0505] a picture parameter set, wherein the portion is one or more pictures and the second syntax element controls updating (158) in units of:
[0506] a picture,
[0507] a tile,
[0508] a slice,
[0509] a coding tree block (e.g., a unit in which a picture is uniformly pre- divided before each coding tree block is split into coding blocks by recursive multi- tree subdivision), a coding block (making intra / inter picture mode determinations),
[0510] coding blocks (e.g. in units of e.g. making intra / inter picture mode decisions), or
[0511] residual transform blocks; or
[0512] a picture header, wherein the portion is one picture and the second syntax element controls updating (158) in units of:
[0513] a tile,
[0514] a slice,
[0515] coding tree blocks (e.g. units of a uniform pre-subdivision of a picture before each coding tree block is split into coding blocks by means of a recursive multi-tree subdivision), coding blocks (making intra / inter picture mode decisions),
[0516] coding blocks (e.g. in units of e.g. making intra / inter picture mode decisions), or
[0517] residual transform blocks; or
[0518] a parameter set, wherein the portion is one coding tree block and the second syntax element controls updating (158) in units of:
[0519] coding blocks (e.g. in units of e.g. making intra / inter picture mode decisions), or
[0520] residual transform blocks.
[0521] According to a sixteenth aspect (when referring back to any one of the first to fifteenth aspects), in the decoder (20), the media signal can be a color video, and the decoder (20) can be configured to use different versions of the set (151) of default quantizers for different color components.
[0522] According to a seventeenth aspect (when referring back to any one of the first to sixteenth aspects), in the decoder (20), the binary function can be a parity check.
[0523] According to an eighteenth aspect (when referring back to any one of the first to seventeenth aspects), the decoder (20) can be configured to reconstruct (170) the media signal using the dequantized residual values (104) by:
[0524] performing a prediction to obtain a predicted version of the media signal; and
[0525] correcting the predicted version with the dequantized residual values (104).
[0526] According to the nineteenth aspect (when referring to any one of the first to eighteenth aspects), the decoder (20) can be configured to reconstruct (170) the media signal using the dequantized residual values (104) by:
[0527] performing a prediction to obtain a predicted version of the media signal;
[0528] inverse transforming the dequantized residual values (104) to obtain a media residual signal; and
[0529] correcting the predicted version with the media residual signal.
[0530] According to the twentieth aspect, an encoder can be configured to:
[0531] predictively encode a media signal to obtain a residual signal;
[0532] sequentially quantize residual values representing the residual signal to obtain residual levels (142) by:
[0533] select (152) a quantizer (153) from a set (151) of default quantizers in dependence on a current conversion state (220);
[0534] quantize a current residual value using the quantizer (153) to obtain a current residual level;
[0535] update (158) the current conversion state (220) in dependence on a feature (157) of the current residual level obtained by applying (156) a binary function to the current residual level (e.g. parity check) and in dependence on quantization mode information (180) transmitted in a data stream (14); and
[0536] encode the residual levels (142) into the data stream (14);
[0537] wherein the current conversion state (220) is converted from a domain of a combination (163) of one or more conversion states' sets (161) having the feature (157) of the current residual level to the one or more conversion states' sets (161) according to an onto mapping depending on the quantization mode information (180), wherein a cardinality of the one or more conversion states' sets (161) differs in dependence on the quantization mode information (180); and
[0538] independent of the quantization mode information (180) (e.g. independent of or independent from the quantization mode information (180)), the current conversion state (220) is converted using the first mapping (200) by using a predetermined number of bits of the current conversion state (220) (e.g.
[0539] One or more bits located at one or more predetermined bit positions, e.g. LSBs or bits located at the second but not the last significant bit position, map to a default quantizer to perform the selection (152) of the quantizer (153); the predetermined number and the first mapping (200) are equal regardless of the quantization mode information (180).
[0540] According to a twenty-first aspect (when referring back to the twentieth aspect), in the encoder, the predetermined number of bits can be one.
[0541] According to a twenty-second aspect (when referring back to the twentieth or twenty-first aspect), the encoder can be configured to select, depending on the quantization mode information (180), a conversion table (106) to be used from the set (159) of default conversion tables, and to manage the set (159) of default conversion tables as different parts of a uniform conversion table, wherein for each default conversion table, each of the set (161) of one or more conversion states is indexed using a state index that is different from the state indices used to index any other default conversion table in the uniform conversion table (e.g. the default conversion tables are disjoint from each other with respect to the conversion states of the different parts of the uniform conversion table).
[0542] According to a twenty-third aspect (when referring back to the twenty-second aspect), the encoder can be configured to perform the updating (158) of the current conversion state (220) by looking up an entry in the uniform conversion table, the entry corresponding to the combination of the characteristics (157) of the current residual level and the current conversion state (220).
[0543] According to a twenty-fourth aspect (when referring back to the twenty-second or twenty-third aspect), the encoder can be configured to encode into the data stream (14) as the quantization mode information (180) a syntax element representing a starting conversion state used for quantizing the first residual value in a quantization order (e.g. in which order the inverse quantization is performed in sequence), the syntax element indicating from the default conversion tables the conversion table (106) to be used by fitting to the state index in the part of the uniform conversion table corresponding to the conversion table (106) to be used.
[0544] According to a twenty-fifth aspect (when referring back to any one of the twentieth to twenty-fourth aspects), in the encoder, the set (151) of default quantizers consists of:
[0545] a first default quantizer (1511) comprising as coding levels even integer multiples of the quantization step size (144); and
[0546] a second default quantizer (1512) comprising as coding levels odd integer multiples of the quantization step size (144) and zero.
[0547] According to a twenty-sixth aspect (when referring back to any one of the twentieth to twenty-fifth aspects), the encoder can be configured to:
[0548] encode the current residual level into the data stream (14) using context-adaptive binary arithmetic coding of a binarization of the current residual level, wherein the binarization comprises a significant-value bin (92) representing whether the current residual level is zero or non-zero;
[0549] determine a first context (242) for encoding the first bin (92) of the current residual level by using a second mapping (144) to map a further predetermined number of bits of the current conversion state (220) to a value (230), wherein the second mapping (144) is equal for any quantization mode information (180), and select the first context (242) in dependence on the obtained value (230).
[0550] According to a twenty-seventh aspect (when referring back to the twenty-sixth aspect), the further predetermined number of bits can be two in the encoder.
[0551] According to a twenty-eighth aspect (when referring back to any one of the twentieth to twenty-seventh aspects), the media signal can be a video in the encoder, and the encoder can be configured to:
[0552] encode the residual levels (142) into the data stream (14) by:
[0553] encode the residual levels (142) for the image block (84) before the residual levels for another image block (84);
[0554] encode the binarized bins for the residual levels for the image block (84) sequentially in one or more scans, and use context-adaptive binary arithmetic coding for a predetermined number of leading bins of the binarized bins for the residual levels, and use an equal-probability bypass mode for subsequent bins of the binarized bins for the residual levels;
[0555] determine the binarization of the predetermined residual level, wherein the predetermined number of bins is exhausted, by using a third mapping to map a further predetermined number of bits of the conversion state (e.g. one or more bits at one or more predetermined bit positions, such as the LSBs or the bits at the second but not the last significant bit position) to a binarization parameter (302), and determine the binarization of the predetermined residual level in dependence on the binarization parameter (302); the third mapping is equal for any quantization mode information (180).
[0556] According to a twenty-ninth aspect (when referring back to the twenty-eighth aspect), the encoder can be configured to determine binarization of the predetermined residual levels according to a binarization parameter (302) such that the binarization is a default binarization modified by redirecting binarization codes according to the binarization parameter (302) from being associated with the predetermined level to being associated with level zero and redirecting one or more binarization codes from being associated with one or more levels from zero to the predetermined level minus one to being associated with one or more levels from one to the predetermined level.
[0557] According to a thirtieth aspect (when referring back to the twenty-eighth or twenty-ninth aspect), the further predetermined number of bits can be one in the encoder.
[0558] According to a thirty-first aspect (when referring back to the twenty-eighth aspect), in the encoder, the surjective mapping and the cardinality of the set of one or more transition states (161) can be selected between two or more of the following according to quantization mode information (180):
[0559] the cardinality of the set of one or more transition states (161) is one;
[0560] the cardinality of the set of one or more transition states (161) is four;
[0561] the cardinality of the set of one or more transition states (161) is eight.
[0562] According to a thirty-second aspect (when referring back to any one of the twenty- first to thirty-first aspects), in the encoder, the media signal can be video and the encoder can be configured to:
[0563] the quantization mode information (180) is written in one of the following ways to control the updating (158):
[0564] only once for the video;
[0565] per image block (84);
[0566] on a per-image basis;
[0567] on a per-slice basis; and
[0568] on a per-image sequence basis.
[0569] According to a thirty-third aspect (when dependent on any of the twentieth to thirty-second aspects), in the encoder, the media signal can be a video, and the quantization mode information (180) can comprise a first syntax element in the data stream (14) indicating whether the video or a portion of the video is to have an update (158) controlled by a second syntax element of the quantization mode information (180) in the data stream (14) within the video or the portion of the video, or which one of the surjective mappings and which one of the cardinalities of the set (161) of one or more transition states is to be selected for the update (158).
[0570] According to a thirty-fourth aspect (when dependent on the thirty-third aspect), in the encoder, the first syntax element can be comprised in the data stream (14) within one of:
[0571] a video parameter set, wherein the portion is the video and the second syntax element controls the update (158) in units of:
[0572] a picture sequence,
[0573] a picture,
[0574] a tile,
[0575] a slice,
[0576] a coding tree block (e.g. a unit in which a picture is uniformly pre- subdivided before being split into coding blocks by means of recursive multi- tree subdivision), a coding block (making an intra / inter picture mode determination),
[0577] a coding block (e.g. in units making an intra / inter picture mode determination), or
[0578] a residual transform block;
[0579] a sequence parameter set, wherein the portion is a picture sequence and the second syntax element controls the update (158) in units of:
[0580] a picture,
[0581] a tile,
[0582] a slice,
[0583] a coding tree block (e.g. a unit in which a picture is uniformly pre- subdivided before being split into coding blocks by means of recursive multi- tree subdivision), a coding block (making an intra / inter picture mode determination),
[0584] a coding block (e.g. in units making an intra / inter picture mode determination), or
[0585] residual transform block;
[0586] a picture parameter set, wherein the portion is one or more pictures and the second syntax element controls updating (158) in units of:
[0587] a picture,
[0588] a tile,
[0589] a slice,
[0590] a coding tree block (e.g. a unit in which a picture is uniformly pre- divided before each coding tree block is split into coding blocks by means of recursive multi-type subdivision), a coding block (making intra / inter picture mode determination),
[0591] a coding block (e.g. a unit in which e.g. intra / inter picture mode determination is made), or
[0592] a residual transform block; or
[0593] a picture header, wherein the portion is one picture and the second syntax element controls updating (158) in units of:
[0594] a tile,
[0595] a slice,
[0596] a coding tree block (e.g. a unit in which a picture is uniformly pre- divided before each coding tree block is split into coding blocks by means of recursive multi-type subdivision), a coding block (making intra / inter picture mode determination),
[0597] a coding block (e.g. a unit in which e.g. intra / inter picture mode determination is made), or
[0598] a residual transform block; or
[0599] a parameter set, wherein the portion is one coding tree block and the second syntax element controls updating (158) in units of:
[0600] a coding block (e.g. a unit in which e.g. intra / inter picture mode determination is made), or
[0601] a residual transform block.
[0602] According to a thirty-fifth aspect (when referring back to any one of the twentieth to thirty-fourth aspects), in the encoder, the media signal can be a color video, and the encoder can be configured to use different versions of the set (151) of default quantizers for different color components.
[0603] According to a thirty-sixth aspect (when referring back to any one of the twentieth to thirty-fifth aspects), in the encoder, the binary function can be a parity check.
[0604] According to a thirty-seventh aspect (when referring back to any one of the twentieth to thirty-sixth aspects), the encoder can be configured to predictively encode the media signal to obtain a residual signal by:
[0605] performing the prediction to obtain a predicted version of the media signal and the residual signal.
[0606] According to a thirty-eighth aspect (when referring back to any one of the twentieth to thirty-seventh aspects), the encoder can be configured to predictively encode the media signal to obtain a residual signal by:
[0607] performing the prediction to obtain a predicted version of the media signal and the residual signal;
[0608] forward transforming the residual signal.
[0609] According to a thirty-ninth aspect, the method can have the following steps:
[0610] decoding a residual level from the data stream, wherein the residual level represents a prediction residual, and sequentially dequantizing the residual level by:
[0611] selecting a quantizer from a set of default quantizers, according to a current conversion state;
[0612] dequantizing the current residual level using the quantizer to obtain a dequantized residual value;
[0613] updating the current conversion state according to a property of the current residual level obtained by applying a binary function to the current residual level (e.g. a parity check), and according to quantization mode information contained in the data stream;
[0614] and
[0615] reconstructing the media signal using the dequantized residual value;
[0616] converting the current conversion state from a domain of a combination of a set of one or more conversion states having the property of the current residual level to the set of one or more conversion states according to a surjective mapping depending on the quantization mode information, wherein a cardinality of the set of one or more conversion states differs according to the quantization mode information; and
[0617] Regardless of the quantization mode information (e.g. independent of the quantization mode information or independent of the quantization mode information), the selection of the quantizer is performed by using a first mapping of a predetermined number of bits of the current conversion state (e.g. one or more bits at one or more predetermined bit positions, e.g. the LSBs or the bits at the second but not the last significant bit position) to the default quantizers; the predetermined number and the first mapping are equal regardless of the quantization mode information.
[0618] According to a fortieth aspect, a method can have the following steps:
[0619] predictively encoding a media signal to obtain a residual signal;
[0620] sequentially quantizing residual values representing the residual signal to obtain residual levels by:
[0621] selecting a quantizer from a set of default quantizers depending on a current conversion state;
[0622] quantizing a current residual value using the quantizer to obtain a current residual level;
[0623] updating the current conversion state depending on a property of the current residual level obtained by applying a binary function to the current residual level (e.g. a parity check) and depending on quantization mode information transmitted in a data stream;
[0624] and
[0625] encoding the residual levels into the data stream;
[0626] wherein the current conversion state is converted from a domain of a combination of a set of one or more conversion states having the property of the current residual level to the set of one or more conversion states according to a surjective mapping depending on the quantization mode information, wherein the cardinality of the set of one or more conversion states differs depending on the quantization mode information; and
[0627] Regardless of the quantization mode information (e.g. independent of the quantization mode information or independent of the quantization mode information), the selection of the quantizer is performed by using a first mapping of a predetermined number of bits of the current conversion state (e.g. one or more bits at one or more predetermined bit positions, e.g. the LSBs or the bits at the second but not the last significant bit position) to the default quantizers; the predetermined number and the first mapping are equal regardless of the quantization mode information.
[0628] According to a forty-first aspect, a data stream can be encoded by a method according to the fortieth aspect.
[0629] According to a forty-second aspect, there can be a computer program having a program code which, when executed on one or more computers, performs the method according to the thirty-ninth or the fortieth aspect.
Claims
1. A decoder (20) configured to: decode (210) from a data stream (14) a residual level (142) representing a prediction residual, and sequentially dequantize (150) the residual level (142) by: selecting (152) a quantizer (153) from a set of default quantizers (151) depending on a current conversion state (220); dequantizing (154) a current residual level using the quantizer (153) to obtain a dequantized residual value; applying (156) a binary function to the current residual level to obtain a characteristic (157) of the current residual level; and updating (158) the current conversion state (220) depending on the characteristic (157) of the current residual level with a conversion table (106) to be used, reconstructing (170) a media signal using the dequantized residual value (104), selecting a default conversion table from a set of default conversion tables (159) as the conversion table (106) to be used depending on quantization mode information (180) contained in the data stream (14), wherein each of the default conversion tables represents a surjective mapping from a domain of combinations (163) of sets of conversion states (161) having the characteristic (157) of the current residual level to sets of conversion states (161), wherein the default conversion tables differ in a cardinality of the sets of conversion states (161), wherein the cardinality is four or more, and performing the selection (152) of the quantizer (153) by mapping a predetermined number of bits of the current conversion state (220) to the default quantizers using a first mapping (200) regardless of which default conversion table is selected as the conversion table (106) to be used, wherein the predetermined number and the first mapping (200) are invariant regardless of which default conversion table is selected as the conversion table (106) to be used.
2. The decoder of claim 1, wherein the predetermined number of bits is one bit.
3. The decoder of claim 1, wherein the decoder is configured to manage the set of default conversion tables (159) as different parts of a uniform conversion table, wherein for each default conversion table, each conversion state of the sets of conversion states (161) is indexed using a state index that is different from a state index used to index any other default conversion table in the uniform conversion table.
4. The decoder of claim 3, wherein the decoder is configured to perform updating (158) the current conversion state (220) by looking up in the uniform conversion table an entry corresponding to a combination of the characteristic (157) of the current residual level and the current conversion state (220). 5. The decoder of claim 3, wherein the quantization mode information (180) comprises a syntax element in the data stream (14) representing a start transition state used in dequantizing the first residual level in the reverse quantization order such that the syntax element indicates from among the default transition tables the transition table (106) to be used by fitting to state indices used only in a portion of a uniform transition table corresponding to the transition table (106) to be used.
6. The decoder of claim 1, wherein the set (151) of default quantizers consists of: a first default quantizer (151 1) comprising even integer multiples of a quantization step size (144) as reconstruction levels; and a second default quantizer (1512) comprising odd integer multiples of the quantization step size (144) and zero as reconstruction levels.
7. The decoder of claim 1, configured to: decode the current residual level from the data stream (14) using a context-adaptive binary arithmetic decoding of a binarization of the current residual level into binary codes, the binary codes comprising significant value binary codes (92) representing whether the current residual level is zero or non-zero; select a first context (242) for decoding a first binary code (92) of the current residual level in dependence on a value (230) obtained by mapping a further predetermined number of bits of the current transition state (220) to the value (230) using a second mapping (144), the second mapping (144) being equal regardless of which default transition table is selected as the transition table (106) to be used.
8. The decoder of claim 7, wherein the further predetermined number of bits is two bits.
9. The decoder of claim 1, wherein the media signal is a video and the decoder is configured to: decode (210) the residual levels (142) from the data stream (14) by: decoding residual levels (142) with respect to a picture block (84) before residual levels (142) with respect to another picture block (84); decoding in sequence binary codes of a binarization of the residual levels (142) with respect to the picture block (84) in one or more scans and using context-adaptive binary arithmetic decoding for a predetermined number of leading binary codes of the binary codes of the binarization of the residual levels and using an equal probability bypass mode for subsequent binary codes of the binarization of the residual levels; determining the binarization of a predetermined residual level in which the predetermined number of binary codes is exhausted by mapping a further predetermined number of bits of a transition state to a binarization parameter (302) using a third mapping and in dependence on the binarization parameter (302) to perform the determination of the binarization of the predetermined residual level.
10. The decoder of claim 9, configured to determine the binarization of the predetermined residual levels in dependence on the binarization parameter (302) such that the binarization is a default binarization modified by redirecting one binarization code in dependence on the binarization parameter (302) from being associated with a predetermined level to being associated with level zero and by redirecting one or more binarization codes from being associated with one or more levels from zero to the predetermined level minus one to being associated with one or more levels from one to the predetermined level.
11. The decoder of claim 9, wherein the further predetermined number of bits is one bit.
12. The decoder of claim 1, wherein the set of default transform tables (159) comprises: a first default transform table, wherein the base number of the set of transform states (161) is four; and a second default transform table, wherein the base number of the set of transform states (161) is eight.
13. The decoder of claim 1, wherein the media signal is video and the decoder is configured to: read the quantization mode information (180) from the data stream (14) in one of the following ways and to perform the selection of the one default transform table from among the set of default transform tables (159): only once for the video; per image block (84); on a per image basis; on a per slice basis; and on a per image sequence basis.
14. The decoder of claim 1, wherein the media signal is video and the quantization mode information (180) comprises a first syntax element in the data stream (14) indicating whether the selection of the one default transform table from among the set of default transform tables (159) is controlled for the video or a portion of the video via a second syntax element of the quantization mode information (180) in the data stream (14) within the video or the portion of the video or which default transform table of the set of default transform tables (159) is to be selected as the one default transform table for the portion of the video.
15. The decoder of claim 14, wherein the first syntax element is contained in one of the following in the data stream (14): a video parameter set, wherein the portion is the video and the second syntax element controls the selection of the one default transform table from among the set of default transform tables (159) in units of: an image sequence, an image, a tile, a slice, a coding tree block, a coding block, or a residual transform block; a sequence parameter set, wherein the portion is an image sequence and the second syntax element controls the selection of the one default transform table from among the set of default transform tables (159) in units of: an image, a tile, a slice, a coding tree block, a coding block, or a residual transform block. a picture parameter set, wherein the portion is one or more pictures, and the second syntax element controls selection of the one default transform table from among the set (159) of default transform tables in units of: a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block; or a picture header, wherein the portion is one picture, and the second syntax element controls selection of the one default transform table from among the set (159) of default transform tables in units of: a tile, a slice, a coding tree block, a coding block, or a residual transform block; or a parameter set, wherein the portion is one coding tree block, and the second syntax element controls selection of the one default transform table from among the set (159) of default transform tables in units of: a coding block; or a residual transform block.
16. The decoder of claim 1, wherein the media signal is a color video, and the decoder is configured to use different versions of the set (151) of default quantizers for different color components.
17. The decoder of claim 1, wherein the binary function is a parity check.
18. The decoder of claim 1, configured to reconstruct (170) the media signal using the dequantized residual values (104) by: performing a prediction to obtain a predicted version of the media signal; correcting the predicted version with the dequantized residual values (104).
19. The decoder of claim 1, configured to reconstruct (170) the media signal using the dequantized residual values (104) by: performing a prediction to obtain a predicted version of the media signal; inverse transforming the dequantized residual values (104) to obtain a media residual signal; and correcting the predicted version with the media residual signal.
20. An encoder configured to: predictively encode a media signal to obtain a residual signal; sequentially quantize residual values representing the residual signal to obtain residual levels (142) by: selecting (152) a quantizer (153) from a set (151) of default quantizers in dependence on a current transform state (220); quantizing (154) a current residual value using the quantizer (153) to obtain a current residual level; applying (156) a binary function to the current residual level to obtain a property (157) of the current residual level; and updating (158) the current transform state (220) in dependence on the property (157) of the current residual level using a transform table (106) to be used; encode the residual levels (142) into a data stream (14). selecting a default conversion table from a set of default conversion tables (159) as the conversion table (106) to be used in dependence on quantization mode information (180) transmitted in the data stream (14), wherein each of the default conversion tables represents a surjective mapping from a domain of combinations (163) of sets of conversion states (161) having a characteristic (157) of the current residual level to a set of conversion states (161), wherein the default conversion tables differ in a cardinality of the set of conversion states (161), wherein the cardinality is four or more; and performing the selection (152) of the quantizer (153) by mapping a predetermined number of bits of the current conversion state (220) to the default quantizer using a first mapping (200), wherein the predetermined number and the first mapping (200) are invariant regardless of which default conversion table is selected as the conversion table (106) to be used.
21. The encoder of claim 20, wherein the predetermined number of bits is one bit.
22. The encoder of claim 20, wherein the encoder is configured to manage the set of default conversion tables (159) as different parts of a uniform conversion table, wherein for each default conversion table, each conversion state of the set of conversion states (161) is indexed using a state index that is different from a state index used to index any other default conversion table in the uniform conversion table.
23. The encoder of claim 22, wherein the encoder is configured to perform updating (158) the current conversion state (220) by looking up an entry in the uniform conversion table that corresponds to a combination of the characteristic (157) of the current residual level and the current conversion state (220).
24. The encoder of claim 22, configured to encode a syntax element as the quantization mode information (180) into the data stream (14), the syntax element representing a starting conversion state used in quantizing a first residual value in quantization order, such that the syntax element indicates the conversion table (106) to be used from among the default conversion tables by fitting to state indices used only in a part of the uniform conversion table corresponding to the conversion table to be used.
25. The encoder of claim 20, wherein the set of default quantizers (151) consists of: a first default quantizer (1511) comprising even integer multiples of a quantization step size (144) as encoding levels; and a second default quantizer (1512) comprising odd integer multiples of the quantization step size (144) and zero as encoding levels.
26. The encoder of claim 20, configured to: encoding the current residual level into the data stream (14) using context-adaptive binary arithmetic coding of a binarization of the current residual level, the binarization comprising a significant-value bin (92) that indicates whether the current residual level is zero or non-zero; selecting a first context (242) for encoding a first bin (92) of the current residual level in dependence on a value (230) obtained by mapping a further predetermined number of bits of the current transform state (220) to the value using a second mapping (144), wherein the second mapping (144) is identical irrespective of which default transform table is selected as the transform table to be used.
27. The encoder of claim 26, wherein the further predetermined number of bits is two bits.
28. The encoder of claim 20, wherein the media signal is video and the encoder is configured to: encode the residual levels into the data stream (14) by: encoding residual levels for one image block (84) before residual levels for another image block (84); encoding binarized bins for the residual levels for the one image block (84) sequentially in one or more scans, and using context-adaptive binary arithmetic coding for a predetermined number of leading bins of the binarized bins for the residual levels, and using an equal-probability bypass mode for subsequent bins of the binarized bins for the residual levels; determining the binarization of a predetermined residual level in which the predetermined number of bins is exhausted by mapping a further predetermined number of bits of a transform state to a binarization parameter (302) using a third mapping, and the determination of the binarization of the predetermined residual level is made in dependence on the binarization parameter (302).
29. The encoder of claim 28, configured to determine the binarization of the predetermined residual level in dependence on the binarization parameter (302) such that the binarization is a default binarization modified by: redirecting one binarization code from being associated with a predetermined level to being associated with level zero, and redirecting one or more binarization codes from being associated with one or more levels from zero to the predetermined level minus one to being associated with one or more levels from one to the predetermined level in dependence on the binarization parameter (302).
30. The encoder of claim 28, wherein the further predetermined number of bits is one bit.
31. The encoder of claim 20, wherein the set of default transform tables (159) comprises: a first default transform table in which the base of the set of transform states (161) is four; and a second default transform table in which the base of the set of transform states (161) is eight.
32. The encoder of claim 20, wherein the media signal is video and the encoder is configured to: encode the residual levels into the data stream (14) by: encoding residual levels for one image block (84) before residual levels for another image block (84); encoding binarized bins for the residual levels for the one image block (84) sequentially in one or more scans, and using context-adaptive binary arithmetic coding for a predetermined number of leading bins of the binarized bins for the residual levels, and using an equal-probability bypass mode for subsequent bins of the binarized bins for the residual levels; The quantization mode information (180) is written into the data stream (14) in one of the following ways, and the one default transform table is selected from among the set of default transform tables (159): only once for the video; per image block (84); per picture; per slice; and per picture sequence.
33. The encoder of claim 20, wherein the media signal is a video, and the quantization mode information (180) comprises a first syntax element in the data stream (14) indicating whether selection of the one default transform table from among the set of default transform tables (159) is controlled within the video or a portion of the video via a second syntax element of the quantization mode information (180) in the data stream (14), or which default transform table of the set of default transform tables (159) is to be selected as the one default transform table for the portion of the video.
34. The encoder of claim 33, wherein the first syntax element is contained in one of the following in the data stream (14): a video parameter set, wherein the portion is the video, and the second syntax element controls selection of the one default transform table from among the set of default transform tables (159) in units of: a picture sequence, a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block; a sequence parameter set, wherein the portion is a picture sequence, and the second syntax element controls selection of the one default transform table from among the set of default transform tables (159) in units of: a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block; a picture parameter set, wherein the portion is one or more pictures, and the second syntax element controls selection of the one default transform table from among the set of default transform tables (159) in units of: a picture, a tile, a slice, a coding tree block, a coding block, or a residual transform block; or a picture header, wherein the portion is one picture, and the second syntax element controls selection of the one default transform table from among the set of default transform tables (159) in units of: a tile, a slice, a coding tree block, a coding block, or a residual transform block; or a parameter set, wherein the portion is one coding tree block, and the second syntax element controls selection of the one default transform table from among the set of default transform tables (159) in units of: a coding block, or a residual transform block.
35. The encoder of claim 20, wherein the media signal is a color video, and the encoder is configured to use different versions of the set of default quantizers (151) for different color components.
36. The encoder of any of the preceding claims 20 to 35, wherein the binary function is a parity check.
37. The encoder of claim 20, configured to predictively encode a media signal to obtain a residual signal by: performing a prediction to obtain a predicted version of the media signal and the residual signal.
38. The encoder of claim 20, configured to predictively encode a media signal to obtain a residual signal by: performing a prediction to obtain a predicted version of the media signal and the residual signal; forward transforming the residual signal.
39. A decoding method for supporting adaptive dependent quantization of transform coefficient levels, comprising: decoding from a data stream residual levels representing a prediction residual and sequentially dequantizing the residual levels by: selecting a quantizer from a set of default quantizers in dependence on a current transition state; dequantizing a current residual level using the quantizer to obtain a dequantized residual value; applying a binary function to the current residual level to obtain a property of the current residual level; and reconstructing a media signal using the dequantized residual value in dependence on the property of the current residual level using a transition table to be used to update the current transition state, selecting a default transition table from a set of default transition tables as the transition table to be used in dependence on quantization mode information contained in the data stream, wherein each of the default transition tables represents a surjective mapping from a domain of combinations of sets of transition states having the property of the current residual level to a set of transition states, wherein the default transition tables differ in a cardinality of the set of transition states, wherein the cardinality is four or more, and regardless of which default transition table is selected as the transition table to be used, performing the selection of the quantizer by mapping a predetermined number of bits of the current transition state to the default quantizers using a first mapping, wherein the predetermined number and the first mapping are invariant regardless of which default transition table is selected as the transition table to be used.
40. An encoding method for supporting adaptive dependent quantization of transform coefficient levels, comprising: predictively encoding a media signal to obtain a residual signal; sequentially quantizing residual values representing the residual signal to obtain residual levels by: selecting a quantizer from a set of default quantizers in dependence on a current transition state; quantizing a current residual value using the quantizer to obtain a current residual level; applying a binary function to the current residual level to obtain a property of the current residual level; and updating the current transition state in dependence on the property of the current residual level using a transition table to be used; encoding the residual levels into a data stream; selecting, from among a set of default conversion tables, a default conversion table to be used as the conversion table, in dependence on quantization mode information transmitted in the data stream, wherein each of the default conversion tables represents a surjective mapping from a domain of combinations of conversion states having the characteristic of the current residual level to a set of conversion states, wherein the default conversion tables differ in a cardinality of the set of conversion states, wherein the cardinality is four or more; and performing the selection of the quantizer by mapping a predetermined number of bits of a current conversion state to the default quantizer using a first mapping, regardless of which default conversion table is selected as the conversion table to be used, wherein the predetermined number and the first mapping are invariant regardless of which default conversion table is selected as the conversion table to be used.
41. A computer program product having program code which, when the program code runs on one or more computers, performs the decoding method according to claim 39 or the encoding method according to claim 40.
Citation Information
Patent Citations
Coding of significance maps and transform coefficient blocks
CN102939755A
Context initialization in entropy coding
CN103733622A