Dependent quantization
Dependent scalar quantization improves video coding efficiency by determining reconstruction levels based on quantization index parity, reducing average quantization error and enhancing coding efficiency.
Patent Information
- Application Number
- JP2025194256
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-03-29
- Filing Date
- 2025-11-13
- Publication Date
- 2026-02-10
AI Technical Summary
Existing video coding technologies face a trade-off between bitrate and quantization distortion due to the use of independent scalar quantization, where coarser quantization reduces bitrate but increases distortion, and finer quantization reduces distortion but increases bitrate.
Implementing dependent scalar quantization that determines the signs from quantization indexes and transform coefficients, using a state variable to select reconstruction level sets based on the parity of preceding indices, thereby improving encoding and decoding efficiency.
This approach enhances coding efficiency by reducing average quantization error and allowing effective control of reconstruction points, combining with weighting matrices for transform coefficients.
Smart Images

Figure 2026021588000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to media signal coding, such as video coding, and in particular to lossy codecs that use quantization in the encoding of media signals, for example for the quantization of prediction residuals, as is done, for example, in HEVC. [Background technology]
[0002] When setting the quantization parameters, the encoder must make a compromise: coarser quantization reduces the bitrate but increases quantization distortion, and finer quantization reduces distortion but increases the bitrate. It would be advantageous to have a near-term concept that increases the coding efficiency for a given range of available quantization levels. Summary of the Invention [Problem to be solved by the invention]
[0003] The object of the present invention is to provide a concept based on dependent scalar quantization, determining the signs from the quantization indexes and thus the transform coefficients, thereby improving the efficiency in both encoding and decoding.
[0004] This object is achieved by the subject matter of the independent claims of the present application. [Means for solving the problem]
[0005] In this application, a derived value of the current sample is calculated based on the sign of the quantization index, a state variable updated based on the index, and a magnitude corresponding to the index (e.g., 2 × index), and a transform coefficient is determined as Δ × derived value. The state variable transitions based on the parity of the preceding index, and is used to select two reconstruction level sets. In this application, in sequential coding of a sample sequence, dependent scalar quantization is used, which selects a reconstruction level set (two sets) of the current sample depending on the quantization index sequence of the preceding sample, and updates the state variable (e.g., 0 to 3) according to the parity of the index. This allows for effective use of the grid of quantization points formed along a time series, increases the effective density of allowable reconstruction points, and reduces the average quantization error. Hereinafter, the signal to be processed is referred to as a media signal.
[0006] According to an embodiment, the media signal is a two-dimensional signal, e.g., an image, and the sequence of samples is obtained by using some scanning pattern that transforms the two-dimensional spatial arrangement of the samples into a one-dimensional sequence along the configuration of the quantization point grid that then occurs.
[0007] According to an embodiment, the sequence of samples representing a media signal represents an image or a part thereof, in other words it is obtained by transforming an image block or a block of spatial samples, such as a transform block of the image, i.e., predictor residual samples in a transform coefficient block, in which certain transform coefficients are scanned from the sequence of samples according to a certain coefficient scan. The transform may be a linear transform or any other transform, and for reconstruction purposes, an inverse transform or some other inverse transform approximating the inverse transform is used. Additionally or alternatively, the sequence of samples may represent a prediction residual.
[0008] According to an embodiment, the selection of the reconstruction level set to be applied to the current sample may be determined based on the least significant bit (parity) of the quantization index of the immediately preceding sample. The number of sets to be selected is two. In this method, a state machine updates the state for each sample, and the state transitions according to the parity of the quantization index of the previous sample. The state variable takes four values, for example, from 0 to 3, and each state uniquely defines the reconstruction level set to be used. The initial state is set in common by the encoder and decoder. This dependent scalar quantization allows for effective control of the placement of reconstruction points, resulting in improved coding efficiency (reduced average quantization error).
[0009] Preferably, and in accordance with an embodiment of the present application, the multiple quantization level settings are parameterized by a predetermined quantization step size, and information about the predetermined quantization step size is signaled in the data stream. In the case of a sequence of samples representing transform coefficients of a transform block, a specific quantization step size for parameterizing the multiple quantization level settings may be determined for each transform coefficient (sample). For example, the quantization step sizes for the transform coefficients of the transform block may be related in a predetermined manner to a single signaled quantization step size signaled in the data stream. For example, a single quantization step size may be signaled in the data stream for the entire transform block and scaled individually for each transform coefficient relative to a default or a scaling factor set by also being coded in the data stream. Here, the transform coefficient may be determined as the product of the quantization step size Δ and the derived value.
[0010] Therefore, advantageously, the dependent scalar quantization used here allows the combination of the dependent scalar quantization with the concept of using a weighting matrix to weight the overall scaling factors of the transform coefficients of a transform block. [Brief explanation of the drawings]
[0011] The setting of the selected quantization level may take into account the entropy encoding the absolute value of the quantization index. Furthermore, the context dependency may include a dependency on the quantization level for previous samples of the sequence of samples, such as for a region nearby the current sample. Advantageously, embodiments are the subject of the dependent claims. Preferred embodiments of the present application are described below with reference to the following figures: [Figure 1] FIG. 1 shows a block diagram of a typical video encoder as an example of an image encoder that may be implemented to operate in accordance with the embodiments described below. [Figure 2] FIG. 2 shows (a) a transform encoder, (b) a block diagram of a transform encoder to describe the basic approach of block-based transform coding. [Figure 3] FIG. 3 shows a histogram of the distribution describing a uniform reconstruction quantizer. [Figure 4] Figure 4 shows (a) a transform block subdivided into sub-blocks, (b) a schematic diagram of sub-blocks to illustrate an example for scanning transform coefficient levels, where a typical one is used in H.265 | MPEG-H HEVC; in particular, (a) shows the conversion from a 16x16 transform block to a 4x4 sub-block and the coding order of the sub-blocks, and (b) shows the coding order of transform coefficient levels within a 4x4 sub-block. [Figure 5] FIG. 5 shows a schematic diagram of a multidimensional output space spanned by one axis per transform coefficient and arrangements of allowable reconstruction vectors for the simple case of two transform coefficients for examples of (a) independent scalar quantization and (b) dependent scalar quantization. [Figure 6a] Figure 6a shows a block diagram of a transform decoder using dependent scalar quantization, thereby forming a media decoder according to the present application. The modifications to conventional transform coding (which uses an independent scalar quantizer) can be derived by comparing with Figure 2b. [Figure 6b]Figure 6b shows a block diagram of a transform decoder using dependent scalar quantization, thereby forming a media decoder according to the present application. The modifications to conventional transform coding (which uses an independent scalar quantizer) can be derived by comparison with Figure 2a. [Figure 7a] FIG. 7a shows a schematic diagram of a concept for quantization performed within an encoder for coding transform coefficients according to an embodiment, such as by the quantization stage of FIG. 6b. [Figure 7b] FIG. 7b shows a schematic diagram of a concept for dequantization performed within a decoder for decoded transform coefficients according to an embodiment, such as by the dequantization stage of FIG. 6a. [Figure 8a] FIG. 8a shows a schematic diagram of the collection of available quantization settings during switching according to the previous level; in particular, an example of dependent quantization using two settings of reconstruction levels completely determined by a single quantization step size Δ is shown. The two available settings of reconstruction levels are highlighted with different colors (blue for set 0 and red for set 1). Examples of quantization indices indicating the reconstruction levels within the settings are given by the numbers below the circles. The hollow black circles indicate two different subsets within the set of reconstruction levels; the subsets can be used to determine the reconstruction level setting for the next transform coefficient in the reconstruction order. The diagram shows three configurations with two settings of reconstruction levels: (a) the two settings are disjoint and symmetric about zero; (b) both settings include a reconstruction level equal to zero but are otherwise disjoint; the settings are asymmetric around zero; (c) both settings include a reconstruction level equal to zero but are otherwise disjoint; both settings are symmetric around zero. [Figure 8b]FIG. 8b shows a schematic diagram of the collection of available quantization settings during switching according to the previous level; in particular, an example of dependent quantization using two settings of reconstruction levels completely determined by a single quantization step size Δ is shown. The two available settings of reconstruction levels are highlighted with different colors (blue for set 0 and red for set 1). Examples of quantization indices indicating the reconstruction levels within a setting are given by the numbers below the circles. The hollow black circles indicate two different subsets within the set of reconstruction levels; the subsets can be used to determine the reconstruction level setting for the next transform coefficient in the reconstruction order. The diagram shows three configurations with two settings of reconstruction levels: (a) the two settings are disjoint and symmetric about zero; (b) both settings include a reconstruction level equal to zero but are otherwise disjoint; the settings are asymmetric around zero; (c) both settings include a reconstruction level equal to zero but are otherwise disjoint; both settings are symmetric around zero. [Figure 8c] FIG. 8c shows a schematic diagram of the collection of available quantization settings during switching according to the previous level; in particular, an example of dependent quantization using two settings of reconstruction levels completely determined by a single quantization step size Δ is shown. The two available settings of reconstruction levels are highlighted with different colors (blue for set 0 and red for set 1). Examples of quantization indices indicating the reconstruction levels within a setting are given by the numbers below the circles. The hollow black circles indicate two different subsets within the set of reconstruction levels; the subsets can be used to determine the reconstruction level setting for the next transform coefficient in the reconstruction order. The diagram shows three configurations with two settings of reconstruction levels: (a) the two settings are disjoint and symmetric about zero; (b) both settings include a reconstruction level equal to zero but are otherwise disjoint; the settings are asymmetric around zero; (c) both settings include a reconstruction level equal to zero but are otherwise disjoint; both settings are symmetric around zero. [Figure 9a] FIG. 9a shows pseudocode illustrating an example of a method for reconstructing transform coefficients. k denotes an index that identifies the reconstruction order of the current transform coefficient, the quantization index for the current transform coefficient is indicated by level[k], the step size Δk to apply to the current transform coefficient is indicated by quant_step_size[k], and trec[k] represents the value of the reconstructed transform coefficient t_k̂. The variable setId[k] specifies the set of reconstruction levels to apply to the current transform coefficient. It is determined based on the previous transform coefficient in the reconstruction order; possible values of setId[k] are 0 and 1. The variable n specifies an integer factor of the quantization step size; it is given by the reconstruction level (i.e., the value of setId[k]) and the selected set of transmitted quantization index levels[k]. [Figure 9b] Figure 9b shows pseudocode illustrating another embodiment of the pseudocode in Figure 9a. The main change is that the multiplication by the quantization step is expressed using an integer implementation using scale and shift parameters. Typically, the shift parameter (represented by shift) is constant for the transform block, and only the scale parameter (given by scale[k]) may depend on the position of the transform coefficient. The variable add represents a rounding offset, and add = (1 << (shift-1)) is typically set equal. Using Δ_k is the fractional quantization step for the transform coefficient, and the parameters shift and scale[k] are chosen in such a way that we have Δk ≈ scale[k] · 2-shift. [Figure 10a] FIG. 10a shows an example schematic diagram for dividing the reconstruction level settings into two subsets. The two illustrated quantization settings are the example quantization settings of FIG. 8c. The two subsets of quantization set 0 are labeled with "A" and "B," and the two subsets of quantization set 1 are labeled with "C" and "D." [Figure 10b]FIG. 10b shows a table of an example of determining a quantization set (set of available reconstruction levels) to be used for a next transform coefficient based on a subset with the two last quantization indexes. The subsets are shown in the left table column; they are measured independently by the used quantization set (for the two last quantization indexes) and the so-called path (which may be measured by the parity of the quantization indexes). The quantization settings and the path for the subset, in parentheses, are listed to the left in the second column. The third column identifies the associated quantization setting. In the last column, values of so-called state variables are shown, which can be used to simplify the method of determining the quantization setting. [Figure 10c] FIG. 10c shows a further exemplary state transition table for how to switch between available quantization settings, here in a configuration with four states. [Figure 11] 11 shows pseudocode illustrating an example of the reconstruction process of transform coefficients for a transform block. The array level represents the transmitted transform coefficient levels (quantization indices) for the transform block, and the array trec represents the corresponding reconstructed transform coefficients. The 2d table state_trans_table specifies the state transition table, and the table setld specifies the quantization sets associated with the states. [Figure 12] Figure 12 shows an example of a state transition table, state_trans_table, and a table, setId, which specifies the quantization set associated with a state. C-style syntax represents the table specified in the table of Figure 10c. [Figure 13] FIG. 13 shows pseudocode illustrating another reconstruction method for transform coefficient levels, where quantization indices equal to 0 are excluded from state transitions and dependent scalar quantization. [Figure 14] FIG. 14 shows a schematic diagram illustrating state transitions in dependent scalar quantization as a lattice structure. The horizontal positions represent different transform coefficients in the reconstruction order. The vertical axis represents different possible states in the dependent quantization and reconstruction process. The connections shown identify available paths between states for different transform coefficients. [Figure 15] FIG. 15 shows an example of a basic trellis cell. [Figure 16] 16 shows a schematic diagram of an example trellis for dependent scalar quantization of eight transform coefficients. The first state (left side) represents the initial state, which is set equal to 0 in this example. [Figure 17a] FIG. 17a shows a conceptual schematic diagram for entropy decoding quantization levels performed within an encoder for coding transform coefficients according to an embodiment, such as by the entropy decoder in FIG. 6b. [Figure 17b] FIG. 17b shows a conceptual schematic diagram for entropy coding quantization levels performed within an encoder for coding transform coefficients according to an embodiment, such as by the entropy decoder in FIG. 6a. [Figure 18] Figure 18 shows a table with examples for binarization of absolute values of quantization indexes. From left to right: (a) unary binarization; (b) Exponential-Golomb (Exp-Golomb) binarization; (c) concatenated binarization consisting of unary prefix (first two binaries marked in blue) and suffix; (d) concatenated binarization consisting of unary prefix (first two binaries marked in blue) binaries indicating path / parity (in red) and Exponential-Golomb suffix (individual codes for both paths / parities). [Figure 19] Figure 19 shows a schematic diagram of a transform block for the explanation of the concept for entropy coding of transform coefficient levels: (a) Signaling of the position of the first non-zero quantization index in the coding order (black samples). In addition to the position of the first non-zero transform coefficient, only the bin for coefficients marked in blue is transmitted, while coefficients marked in white are assumed to be equal to 0. (b) Example of a table used to select one or more binary probability models. [Figure 20]FIG. 20 shows a schematic diagram of an example lattice structure that can be utilized to determine a sequence (or block) of quantization indexes that minimizes a cost measure (e.g., a Lagrangian cost measure, D+λR). The lattice structure represents a further example of dependent quantization with four states (see FIG. 16). A trellis is shown for eight transform coefficients (or quantization indexes). The first state (leftmost) represents the initial state, which is assumed to be equal to 0. [Figure 21] FIG. 21 shows a block diagram of a decoder that one may implement to operate in accordance with an embodiment such as that depicted in FIG. 7b and that conforms to the example encoder of FIG. [Figure 22] FIG. 22 shows a schematic diagram of the image subdivision for prediction and residual coding and the relationship between them.
[0012] The following description describes concepts for media signal coding using dependent scalar quantization. However, for ease of understanding, the examples provided below relate to transform coding of transform coefficients using dependent scalar quantization, as described below. However, as described below, embodiments of the present application are not limited to transform coding.
[0013] According to the embodiment described below, transform coding involves transforming a set of samples, a resulting dependent scalar quantization of the transform coefficients, and entropy coding of the resulting quantization indices. At the decoder side, a set of reconstructed samples is obtained by entropy decoding the quantization indices, a dependent reconstruction of the transform coefficients, and an inverse transform. In contrast to conventional transform coding, which consists of a transform, an independent scalar quantization, and an entropy coding, the set of allowed reconstruction levels for a transform coefficient depends on the transmitted quantization indices, also called the transform coefficient level preceding the current transform coefficient in the reconstruction order. Furthermore, the entropy coding of the quantization indices, i.e., the transform coefficients specifying the reconstruction levels used for dependent scalar quantization, is described. Furthermore, an adaptive selection between conventional transform coding and transform coding with dependent scalar quantization, and the dependent scalar quantization, is also possible. A concept for adapting the quantization step size used for quantization is described. The following description is primarily targeted at lossy coding of blocks of predictions of blocks of prediction error samples in image and video coding, but the embodiments described below may also be applied to other areas of lossy coding, such as audio coding, or the like. That is, the embodiments described below are not limited to sets of samples forming rectangular blocks, but to prediction error samples, i.e., sets of samples representing the difference between the original and predicted signal. Rather, the embodiments described below may be readily transferred to other scenarios, such as audio signal coding, coding without prediction, coding in the spatial domain rather than the transform domain, etc.
[0014] All modern video codecs, such as the international video coding standards H.264 | MPEG-4 AVC [1] and H.265 | MPEG-H HEVC [2], follow the basic approach of hybrid video coding: a video image is divided into blocks, samples of the blocks are predicted using intra-picture or inter-prediction, and the resulting prediction error signal (the difference between the original samples and the samples of the predicted signal) is coded using transform coding.
[0015] FIG. 1 shows a simplified block diagram of a typical modern video encoder. The video images of a video sequence are coded in a particular order called the coding order. The encoding order of the images can be different from the capture and display order. For the actual encoding, each video image is divided into blocks. A block comprises samples of a rectangular area of a particular color component. The entity of all color component blocks that correspond to the same rectangular area is often called a unit. In H.265 | MPEG-H HEVC, for the purpose of block distribution, it distinguishes between coding tree blocks (CTBs), coding blocks (CBs), prediction blocks (PBs), and transform blocks (TBs). The related units are called coding tree units (CTUs), coding units (CUs), prediction units (PUs), and transform units (TUs).
[0016] Typically, a video image is first divided into fixed-size units (i.e., fixed-size blocks aligned with respect to all color components). In H.265|MPEG-H HEVC, these fixed-size units are called coding tree units (CTUs). Each CTU can be further divided into multiple coding units (CUs). A coding unit is the entity for which a coding mode (e.g., intra- or intra-picture coding) is selected. In H.265|MPEG-H HEVC|, the split of a CTU into one or more CUs is specified by a quadtree (QT) syntax and transmitted as part of the bitstream. The CUs of a CTU are processed in a so-called z-scan order. That is, the four blocks resulting from the split are processed in raster scan order; and if some of the blocks are further split, the corresponding four blocks (including any contained smaller blocks) are processed before the next block at the higher division level is processed.
[0017] If a CU is coded in an intra coding mode, an intra prediction mode for the luma signal is transmitted, and if the video signal includes a chroma component, another intra prediction mode for the chroma signal is transmitted. In ITU TH.265, if the CU size is equal to the smallest CU size (indicated in the sequence parameter settings), the luma block can be divided into four equally sized blocks, and a separate luma prediction mode for each of these blocks is transmitted. The actual intra prediction and coding is performed based on the transform blocks. For each transform block of a coded CU in an image, a prediction signal is extracted using already reconstructed samples of the same color component. The algorithm used to generate the prediction signal for the transform block is determined by the transmitted intra prediction mode.
[0018] CUs coded in intra-picture coding mode can be further split into multiple prediction units (PUs). A prediction unit is a luma entity, and for color video, two related chroma blocks (covering the same image area) are used to signal a prediction parameter setting. A CU can be coded as a single prediction unit, or it can be split into two non-quadratic (symmetric and asymmetric splits are supported) or four quadratic prediction units. For each PU, an individual setting of motion parameters is transmitted. Each setting of motion parameters includes the number of motion hypotheses (one or two in H.265 | MPEG-H HEVC), and for each motion hypothesis, the reference picture (indicated by the reference picture index in the reference picture list) and the associated motion vector. In addition, in H.265 | MPEG-H HEVC, motion parameters are not explicitly transmitted but are derived based on the motion parameters of spatial or temporal neighboring blocks, which is what is provided in the so-called joint mode. If a CU or PU is coded in joint mode, only an index to a list of motion parameter candidates (this list is derived using motion data of spatial or temporal neighboring blocks) is transmitted. The index completely determines the motion parameter settings used. The prediction signal for an inter-coded PU is formed by motion-compensated prediction. For each hypothesis (specified by a reference picture and a motion vector), a prediction signal is formed of a displaced block in a specific reference picture, where the permutation relative to the current PU is specified by the motion vector. The permutation is typically specified with subsample accuracy (in H.265 | MPEG H | HEVC, motion vectors have a precision of one-quarter luma sample). For non-integer motion vectors, the prediction signal is generated by interpolating a reconstructed reference picture (typically using a separable FIR filter). Using multi-hypothesis prediction, the final prediction signal for a PU is formed by a weighted sum of the prediction signals for each motion hypothesis. Typically, the same setting of motion parameters is used for the luma and chroma blocks of the PU.Even though state-of-the-art video coding standards use translational transpose vectors to identify the motion of a current region (block of samples) compared to a reference image, higher-order motion models (e.g., affine motion models) can also be used, in which case additional motion parameters must be transmitted for the motion hypothesis.
[0019] For intra- and inter-picture coded CUs, the prediction error signal (also called the residual signal) is typically transmitted by transform coding. In H.265|MPEG-H HEVC, the block of luma residual samples of a CU and the block of chroma residual samples (if present) are partitioned into transform blocks (TBs). The partitioning of a CU into transform blocks is indicated by a quadtree syntax, also called a residual quadtree (RQT). The resulting transform blocks are coded using transform coding: a 2D transform is applied to the block of residual samples, the resulting transform coefficients are quantized using independent scalar quantization, and the resulting transform coefficient levels (quantization indices) are the entropy coded. For P and B slices, a skip_flag is transmitted at the beginning of the CU syntax. If this flag is equal to 1, it indicates that the corresponding CU consists of a single prediction unit coded in merge mode (i.e., merge_flag is inferred to be equal to 1), and all transform coefficients are equal to zero (i.e., the reconstructed signal is equal to the predicted signal). In that case, only merge_idx is transmitted in addition to skip_flag. If skip_flag is equal to 0, the prediction mode (inter or intra) is indicated following the syntax function described above.
[0020] Since an already coded image can be used for motion-compensated prediction of blocks in a subsequent image, the image must be completely reconstructed at the encoder. For a block (obtained by reconstructing transform coefficients with a given quantization index and inverse transform), the reconstructed prediction error signal is added to the corresponding prediction signal, and the result is stored in a buffer for the current image. After all blocks of the image have been reconstructed, one or more loop filters can be applied (e.g., a deblocking filter and a sample-adaptive offset filter). The final reconstructed image is then stored in a decoded image buffer.
[0021] In the following, a new concept for transform coding of prediction error signals is described. The concept is applicable to blocks coded within and between images. It is also applicable to transform coding of non-rectangular sample regions. In contrast to conventional transform coding, transform coefficients are not independently quantized. Instead, the set of available reconstruction levels for a particular transform coefficient depends on the quantization index selected for other transform coefficients. Furthermore, the modification of the quantization index for coding entropy increases the coding efficiency when combined with the described independent scalar quantization.
[0022] All major video coding standards, including the latest standards H.265|MPEG-H HEVC mentioned above, utilize the concept of transform coding to code a coding block of prediction error samples. The prediction error samples of a block represent the error between the original signal samples and the prediction signal samples for the block. The prediction signal is obtained either by intra-picture prediction (in which case the prediction signal samples for the current block are extracted based on already reconstructed samples of neighboring blocks in the same picture) or by inter-picture prediction (in which case the prediction signal samples for the current block are extracted based on already reconstructed samples). The original prediction error signal samples are obtained by subtracting the prediction signal sample values from the original signal sample values for the current block.
[0023] The transform coding of a block of samples may consist of a linear transformation of the quantization indexes, scalar quantization, and entropy coding. At the encoder side (see Figure 2a), an NxM block of original samples undergoes a linear analytical transform
number
[0024] The result is an NxM block of transform coefficients. k represents the original prediction error samples in a different signal space (or a different coordinate system). The N × M transform coefficients are quantized using N × M independent scalar quantizers. Each transform coefficient t k are the quantization indices q, also called transform coefficient levels k The resulting quantization index q k is the entropy that is coded and written to the bitstream.
[0025] At the decoder side (see Figure 2b), the transform coefficient level q k is decoded from the received bitstream.k is the reconstructed transform coefficient t ‘ k The reconstructed N × M blocks of samples are then mapped to the linear composite transform
number
[0026] Even if a video coding standard only specifies a composite transform B, the inverse of composite transform B may be specified by the encoder, i.e.,
number
number
[0027] For a typical prediction error signal, the transform has the effect of concentrating the signal energy in very few transform coefficients, reducing the statistical dependence between the resulting transform coefficients compared to the original prediction error samples.
[0028] In most modern video coding standards, the separable discrete cosine transform (Type II) or its integer approximation is used. The transform can, however, be easily replaced without modifying other aspects of the transform coding system. Examples for proposed improvements included in literature or standardization documents: Use of Discrete Sine Transform (DST) for image prediction blocks within a picture (possibly dependent on intra prediction mode and block size). H.265|MPEG-H HEVC already includes DST for predicted 4x4 transform blocks within a picture. Switched transforms: The encoder selects which transform among a set of predefined transforms to actually use. The set of available transforms is known to both the encoder and the decoder, so it can be efficiently represented by an index into the list of available transforms. The set of available transforms and their order in the list can depend on other coding parameters for the block, such as the selected intra-prediction mode. In special cases, the transforms used are determined entirely by coding parameters, such as the intra-prediction mode and / or the shape of the block, so that no syntax elements need to be transmitted to specify the transforms. Non-separable transform: The transform used in the encoder and decoder refers to a non-separable transform. Note that the concept of switched transform may include one or more non-separable transforms. For complex reasons, the use of non-separable transforms may be limited to specific block sizes. Multilevel transform: A practical transform consists of two or more transform stages. The first transform stage can consist of a computationally low-complexity separable transform. Then, in the second stage, a subset of the resulting transform coefficients is further transformed using a non-separable transform. Compared to a non-separable transform for the entire transform block, the two-stage approach has the advantage that the more complex non-separable transform is applied to a smaller number of samples. The multilevel transform concept can be efficiently combined with the switched transform concept.
[0029] Transform coefficients are quantized using a scalar quantizer. As a result of quantization, the set of allowable values for the transform coefficients is reduced. That is, the transform coefficients are mapped to a countable set (in practice, a finite set) of so-called reconstruction levels. The set of reconstruction levels represents a proper subset of the set of possible transform coefficient values. To simplify the entropy coding below, the allowable reconstruction levels are represented by quantization indices (also called transform coefficient levels), which are transmitted as part of the bitstream. At the decoder side, the quantization indices (transform coefficient levels) are mapped to reconstructed transform coefficients. The possible values for the reconstructed transform coefficients correspond to a set of reconstruction levels. At the encoder side, the result of scalar quantization is a block of transform coefficient levels (quantization indices).
[0030] In state-of-the-art video coding standards, isomorphic reconstruction quantizers (URQs) are used. Their basic design is illustrated in Figure 3. URQs have the property that the reconstruction levels are equally spaced. The distance Δ between two neighboring reconstruction levels is called the quantization step size. One of the reconstruction levels is equal to 0. Therefore, the complete set of available reconstruction levels is individually specified by the quantization step size Δ. The decoder mapping of quantization index q of the reconstructed transform coefficients t' is, in principle, given by a simple formula:
number
[0031] Video decoders typically use integer arithmetic with standard precision (e.g., 32 bits), so the actual method used in the standard may differ slightly from simple multiplication. When ignoring clipping to the supported dynamic range for the transform coefficients, the reconstructed transform coefficients in H.265|MPEG-H HEVC are
number
number
[0032] Older video standards, such as H.262|MPEG-2 video, also use modulated U for the distance between reconstruction level zero and the first non-zero reconstruction level that is increased compared to the normal quantization step size (e.g., two-thirds of the normal quantization step size). Identify the RQs.
[0033] The quantization step size (or scaling and shifting parameters) for the transform coefficients is determined by two factors: Quantization parameter QP: The quantization step size can generally be changed on a block-by-block basis. To that end, video coding standards provide a predetermined set of quantization step sizes. The quantization step size used (or, equivalently, the parameters "scale" and "shift" mentioned above) is indicated using an index in a predefined list of quantization step sizes. The index is called the quantization parameter (QP). In H.265|MPEG-H HEVC, the relationship between QP and quantization step size is:
number
number
number
number
number
number
number
number
number
[0034] The main purpose of a quantization weight matrix is to provide the possibility to introduce quantization noise in a perceptually meaningful direction. With an appropriate weight matrix, the spatial contrast sensitivity of human vision can be exploited to achieve a better compromise between bit rate and subjective reconstruction quality. Nevertheless, many encoders use so-called flat quantization matrices, which can be efficiently transmitted using high-level syntax elements. In this case, the same quantization step size Δ is used for all transform coefficients in a block. The quantization step size is then completely specified by the quantization parameter QP.
[0035] A block of transform coefficient levels (quantization indices for the transform coefficients) is entropy coded (i.e., it is transmitted losslessly as part of the bitstream). Because a linear transform can only reduce linear dependencies, the entropy coding for the transform coefficient levels is generally designed in such a way that the remaining nonlinear dependencies between the transform coefficient levels in the block can be exploited for efficient coding. Well-known examples are motion level coding in MPEG-2 Video, motion level final coding in H.263 and MPEG-4 Visual, environment adaptive variable length coding (CAVLC) in H.264|MPEG-4 AVC, and environment-based adaptive bilevel arithmetic coding (CABAC) in H.264|MPEG-4 AVC and H.265|MPEG-HHEVC.
[0036] CABAC, specified in the state-of-the-art video coding standards H.265|MPEG-H HEVC, follows a general concept that can be applied to a wide variety of transform block sizes. Transform blocks larger than 4x4 samples, such as 10 in FIG. 4a, are divided into 4x4 sub-blocks 12. The division is illustrated in FIG. 4a for the example of a 16x16 transform block 10. The coding order of the 4x4 sub-blocks, as well as the coding order of the transform coefficient levels inside the sub-blocks, is generally specified by the reverse diagonal scan 14 shown in FIG. 4. For predicted blocks within a particular image, a horizontal or vertical scan pattern is used (depending on the actual intra-prediction mode). The coding order always starts with the high-frequency region.
[0037] In H.265|MPEG-H HEVC, transform coefficient levels are transmitted based on 4x4 sub-blocks. Lossless coding of transform coefficient levels includes the following steps: 1. The syntax element coded_block_flag is transmitted and indicates whether there are any non-zero transform coefficient levels in the transform block. If coded_block_flag is equal to 0, no further data is coded for the transform block. 2. The x and y coordinates of the first non-zero transform coefficient level in the coding order (e.g., the block-wise reverse diagonal scan order illustrated in Figure 4) are transmitted. The transmission of the coordinates is split into a prefix part and a suffix part. The standard uses the syntax elements last_sig_coeff_x_prefix, last_sig_coeff_y_prefix, last_sig_coeff_x_suffix, and last_sig_coeff_x_suffix. 3. Starting with the 4x4 sub-block in coding order that contains the first non-zero transform coefficient level, the 4x4 sub-blocks are processed in coding order, where encoding of the sub-blocks includes the following main steps: a. The syntax element coded_sub_block_flag is sent, which indicates whether the sub-block contains non-zero transform coefficient levels. For the first and last 4x4 sub-blocks (i.e., the sub-blocks containing the first non-zero transform coefficient level or DC level), this flag is not sent and is inferred to be equal to 1. For all transform coefficient levels inside a sub-block with coded_sub_block_flag equal to b.1, the syntax element significant_coeff_flag indicates whether the corresponding transform coefficient level is not equal to zero. The syntax element significant_coeff_flag indicates whether the corresponding transform coefficient level is not equal to zero. This flag is only transmitted if its value cannot be inferred based on already transmitted data. In particular, if a flag is not transmitted for the first significant scan position (identified by the transmitted x and y coordinates) and the DC coefficient is located in a different sub-block than the first non-zero coefficient (in coding order), and all other significant_coeff_flags for the last sub-block are equal to zero, then it is not transmitted for the DC coefficient. The flag coeff_abs_level_greater1_flag is sent if the first 8 transform coefficient levels according to c.significant_coeff_flag (if any) are equal to 1. It indicates whether the absolute value of the transform coefficient block is greater than or equal to 1. d. If the first transform coefficient level according to coeff_abs_level_greater1_flag is equal to 1 (if any), the flag coeff_abs_level_greater2_flag is transmitted, which indicates whether the absolute value of the transform coefficient level is greater than or equal to 2. For all levels with significant_coeff_flag equal to e.1 (exceptions are described below), the syntax element coeff_sign_flag is sent, which specifies the sign of the transform coefficient level. f. For all transform coefficient levels whose absolute value is not already fully specified by the values of significant_coeff_flag, coeff_abs_level_greater1_flag, and coeff_abs_level_greater2_flag (the absolute value is fully specified if any of the sent flags is equal to zero), the remainder of the absolute value is transmitted using the multi-level syntax element coeff_abs_level_remaining.
[0038] In H.265 | MPEG-H HEVC, all syntax elements are coded using contexts based on adaptive binary arithmetic coding (CABAC). All non-binary syntax elements are first mapped to successive binary decisions, called binaries. The resulting sequence of bins is coded using binary arithmetic coding. For that purpose, each binary value is associated with a probability model (a binary probability mass function), also called a context. For most binaries, context implies an adaptive probability model, which means that the binary probability mass function is updated based on the actual coded binary value. Conditional probabilities can be exploited by changing the context for a particular binary value based on already transmitted data. CABAC also includes a so-called bypass mode, in which a constant probability mass function (0.5, 0.5) is used.
[0039] The context selected for encoding coded_sub_block_flag depends on the value of coded_sub_block_flag for neighboring sub-blocks that have already been coded. The context for significant_coeff_flag is selected based on the scan position (x and y coordinates) within the sub-block, the dimensions of the transform block, and the values of coded_sub_block_flag of neighboring sub-blocks.
[0040] For the flags coeff_abs_level_greater1_flag and coeff_abs_level_greater2_flagn, the context selection depends on whether the current sub-block contains a DC coefficient and whether coeff_abs_level_greater1_flag equal to 1 has been sent for neighboring sub-blocks. For coeff_abs_level_greater1_flag, it depends on the number and value of already coded coeff_abs_level_greater1_flag for the sub-block.
[0041] The sign coeff_sign_flag and the remaining absolute value coeff_abs_level_remaining are coded in the bypass mode of the binary arithmetic coder. An adaptive binarization scheme is used to map coeff_abs_level_remaining to a sequence of binary values (binary decisions). The binarization is controlled by a single parameter (called the Rice parameter) that is applied based on the values already coded for the subblock.
[0042] H.265|MPEG-H HEVC also includes so-called code data hiding modes, in which (under certain circumstances) the transmission of the code for the last non-zero level inside a sub-block is omitted. Instead, the code for this level is incorporated into the parity of the sum of absolute values for the level of the corresponding sub-block. It should be noted that the encoder must take this aspect into account when determining the appropriate transform coefficient level.
[0043] Video coding standards only specify the bitstream syntax and the reconstruction process. If we consider transform coding for the original prediction error samples and a given quantization step size, the coder has many degrees of freedom: a given quantization index for the transform block
number
number
number
number
number
[0044] The simplest method of quantization is to quantize the original transform coefficients
number
number
number
number
number
number
number
[0045] The quantization process is the Lagrangian function
number
number
[0046] The relationship between QP and quantization step size (H.264|MPEG-4 AVC or H.265|MPEG-H HEVC)
number
number
number
number
number
number
[0047] The transform coefficient index k then specifies the coding order (or scanning order) of the transform coefficient levels.
number
number
number
number
number
number
[0048] The exact calculation of rate terms for H.265|MPEG-H HEVC transform coefficient coding is very complex because most binary decisions are coded using adaptive probability models. However, if we neglect some aspects of probability model selection and ignore that probability models are applied inside transform blocks, we can design an RDOQ algorithm with reasonable complexity.
[0049] The RDOQ algorithm implemented in the reference software for H.265|MPEG-H HEVC consists of the following basic processing steps: 1. For each scanning position k, the conversion coefficient level
number
number
number
number
number
number
[0050] The following describes a modified concept for transform coding. In modulation associated with conventional transform coding, transform coefficients are not independently quantized and reconstructed. Instead, the allowed reconstruction levels for a transform coefficient depend on the quantization index chosen for the previous transform coefficient in the reconstruction order. The concept of dependent scalar quantization is combined with modified entropy coding, in which the probability model selection (or, alternatively, coding name table selection) for a transform coefficient depends on the set of allowed reconstruction levels.
[0051] The advantage of dependent quantization of transform coefficients is that the allowable reconstruction vectors are packed into an N-dimensional signal space (where N refers to the number of samples or transform coefficients in the transform block), making it denser. The reconstruction vectors for a transform block refer to the ordered reconstructed transform coefficients (or alternatively, the ordered reconstructed samples) of the transform block. The effect of dependent scalar quantization is illustrated in Figure 5 for the simplest case of two transform coefficients. Figure 5 shows the allowable reconstruction vectors (here representing points in the 2D plane) for independent scalar quantization. Thus, the second transform coefficient
number
number
number
number
number
number
number
number
number
[0052] Dependent scalar quantization of transform coefficients has the effect that, for a given average number of reconstructed vectors per N-dimensional unit capacity 20, the expected distance between a given input vector of transform coefficients and the nearest available reconstructed vector is reduced. As a result, the average distortion between the input vector of transform coefficients and the vector's reconstructed transform coefficients can be reduced for a given average number of bits. In vector quantization, this effect is called space-filling gain. Using dependent scalar quantization for transform blocks, most of the potential space-filling gain for high-dimensional vector quantization can be utilized. And, in contrast to vector quantization, the implementation complexity of the reconstruction process (or decoding process) is similar to that of conventional transform coding with independent scalar quantizers. A block diagram of a transform decoder 30 with dependent scalar quantization is shown in FIG. 6a and includes an entropy decoder 32, a quantization stage 34 for dependent scalar inverse quantization, and a synthesis transformer 36. The main change (highlighted in red) is the dependent quantization. The reconstruction order index, as indicated by the vertical arrow
number
number
number
number
number
[0053] A transform coder 40 suitable for the decoder of FIG. 6a is shown in FIG. 6b and implements the analysis transform 4 6, a quantization module 44, and an entropy coder 42. An analytical transform 46, which is the inverse of the analytical transform (or its inverse approximation) 36, is used, and entropy coding 42 is typically assigned to the specific entropy decoding process 32. However, similar to conventional transform coding, there is a lot of freedom to choose the transform indices given the original transform coefficients.
[0054] Embodiments of the present invention are not limited to block-based transform coding. They are also applicable to transform coding of some finite collection of samples. Any type of transform that maps a set of samples onto a set of transform coefficients can be used. This includes linear and nonlinear transforms. The used transform may represent a fully perfect (the number of transform coefficients is greater than the number of samples) or a nearly perfect (the number of transform coefficients is less than the number of samples) transform. In embodiments, dependent quantization is combined with a linear transform of an orthogonal set of basis functions. For integer implementations, the transform may deviate from a linear transform (due to rounding of the transform steps). Furthermore, the transform can refer to an integer approximation using orthogonal basis functions (only if the basis functions are nearly orthogonal). In special cases, the transform can refer to an identity transform, in which case the samples or the remaining samples are directly quantized.
[0055] In an embodiment, the concept of dependent quantization discussed below is used for transform coding of sample blocks (e.g., blocks of original samples or blocks of prediction error samples). In this context, transform can refer to separable transforms, non-separable transforms, and combinations of a separable first transform and a non-separable second transform (where the second transform can be applied to either all transforms or a subset of coefficients obtained after the first transform). If a non-separable second transform is applied to a subset, the subset can consist of a sub-block of the obtained matrix of transform coefficients, or it can represent any other sub-block of transform coefficients obtained after the first transform (e.g., an arbitrarily shaped region inside a block of transform coefficients). Transform may also refer to a multi-level transform having two or more transform levels.
[0056] Details and embodiments regarding dependent quantization of transform coefficients are described below. Furthermore, various methods for entropy coding of quantization indexes that identify reconstructed transform coefficients for dependent scalar quantization are described. Furthermore, optional methods for block-adaptive selection between dependent and independent scalar quantization, as well as methods for adapting the quantization step size, are presented. Finally, approaches for determining quantization indexes for dependent quantization at the encoder are discussed.
[0057] Dependent quantization of a transform coefficient, which refers to the concept of setting available reconstruction levels for a transform coefficient, depends on the quantization index chosen for the previous transform coefficient in the reconstruction order (generally limited to quantization indexes within the same transform block).
[0058] In embodiments, multiple sets of reconstruction levels are predefined and based on the quantization index for the previous transform coefficient in the coding order, and one of the predefined set is selected to reconstruct the current transform coefficient. Some embodiments for defining the set of reconstruction levels are described first. Identification and signaling of the selected reconstruction level are described below. Yet further embodiments for selecting one of the predefined set of reconstruction levels for the current transform coefficient (based on the selected quantization index for the previous transform coefficient in the reconstruction order) are described.
[0059] In an embodiment, and as shown in FIGS. 7a and 7b, a set 48 of allowable reconstruction levels for a current transform coefficient 13′ is selected 54 from among a collection 50 (two or more sets) of predefined sets 52 of reconstruction levels (based on a set 58 of quantization indices 56 for previous transform coefficients in the coding order 14). Therefore, collection 50 may be used for all coefficients 13, but reference will be made following a description of possible parameterizable characteristics of set 52. In a particular embodiment, and as shown in the following example, the values of the reconstruction levels in the set of reconstruction levels are parameterized 60 by a block-based quantization parameter. That is, the video codec supports modulation of the quantization parameter (QP) based on a block (which can correspond to a single transform block 10 or multiple transform blocks), and the value of a reconstruction level in the used set of reconstruction levels is determined by the selected quantization parameter. For example, if the quantization parameter is increased, the distance between neighboring reconstruction levels is also increased, and if the quantization parameter is decreased, the distance between neighboring reconstruction levels is also decreased (or vice versa).
[0060] In a typical version, a block-based quantization parameter (QP) determines the quantization step size Δ (or corresponding magnitude and shift parameters, as described above), and all reconstruction levels (all sets of reconstruction levels) represent integer multiples of the quantization step size Δ. However, it should be noted that each set of reconstruction levels includes only a subset of integer multiples of the quantization step size Δ. In such a configuration for dependent quantization, all possible reconstruction levels for all sets of reconstruction levels that represent integer multiples of the quantization step size can be thought of as an extension of isomorphic reconstruction quantizers (URQs). Its fundamental advantage is that the reconstructed transform coefficients can be calculated by an algorithm with very low computational complexity (described in more detail below).
[0061] The sets 52 of reconstruction levels may be completely disjoint; however, one or more reconstruction levels may be included in multiple sets (while the sets may still differ at other reconstruction levels).
number
number
number
number
number
number
number
number
number
number
[0062] The actual calculation, i.e., the reconstructed transform coefficients
number
number
number
number
number
number
number
[0063] Due to limitations to integer implementations, the reconstructed transform coefficients
number
number
[0064] In an embodiment, the dependent scalar quantization for the transform coefficients uses two distinct sets of reconstruction levels 52.
number
number
number
number
number
[0065] Three configurations for the two sets of reconstruction levels are illustrated in Figure 8. Note that all reconstruction levels lie on a grid given by integer multiples of the quantization step size Δ. Note further that a particular reconstruction level can be included in both sets.
[0066] The two sets shown in FIG. 8a are relatively prime. Each integer multiple of the quantization step size Δ is only included in one of the sets. The first set (Set 0) includes all even integer multiples of the quantization step size, while the second set (Set 1) includes all odd integer multiples of the quantization step size. In both sets, any distance between two neighboring reconstruction levels is twice the quantization step size. These two sets are always suitable for high-rate quantization, i.e., the standard deviation of the transform coefficients is much larger than the quantization step size. However, in video coding, quantizers generally operate in the low-rate range. In general, the absolute values of many original transform coefficients are closer to zero than any non-zero multiple of the quantization step size. In that case, it is generally preferable if zero is included in both quantization sets (sets of reconstruction levels).
[0067] The two quantization sets shown in FIG. 8b both contain zero. In set 0, the distance between the reconstruction level equal to zero and the first reconstruction level greater than zero is equal to the quantization step size, and all other distances between two neighboring reconstruction levels are equal to twice the quantization step size. Similarly, in set 1, the distance between the reconstruction level equal to zero and the first reconstruction level greater than zero is equal to the quantization step size, and all other distances between two neighboring reconstruction levels are equal to twice the quantization step size. Note that both reconstruction sets are asymmetric around zero. This can be inefficient because it makes it difficult to accurately estimate code probabilities.
[0068] Another configuration for two sets of reconstruction levels is shown in FIG. 8c. The reconstruction levels included in the first quantization set (designated as number set 0) represent even integer multiples of the quantization step size (note that this set is actually the same as set 0 in FIG. 8a). The second quantization set (designated as number set 1) includes all odd integer multiples of the quantization step size, and furthermore, the reconstruction level is equal to zero. Note that both reconstruction sets are symmetric about zero. The reconstruction level equal to zero is included in both reconstruction sets; otherwise, the reconstruction sets are disjoint. The combination of both reconstruction sets includes all integer multiples of the quantization step size.
[0069] 1 is not limited to the configuration shown in FIG. 8. Any other two different sets of reconstruction levels can be used. Multiple reconstruction levels may be included in both sets. Or, the combination of both quantization sets may not include all possible integer multiples of the quantization step size. Furthermore, more than two sets of reconstruction levels can be used for dependent scalar quantization of transform coefficients.
[0070] Regarding the signaling of the chosen reconstruction level, the following should be noted: The encoder's selection in step 64 among the allowable reconstruction levels must be indicated in the bitstream 14. As with conventional independent scalar quantization, this can be achieved using so-called quantization indices 56, also referred to as transform coefficient levels. A quantization index (or transform coefficient level) is an integer number that individually identifies an available reconstruction level within a quantization set 48 (i.e., within a set of reconstruction levels). The quantization index is sent to the decoder (e.g., using any entropy coding technique) as part of the bitstream 14. At the decoder side, the reconstructed transform coefficients can be individually calculated based on the reconstruction level (which is determined 54 by the previous quantization index 58 in the coding / reconstruction order) and the current set 48 of transmitted quantization indices 56 for the current transform coefficient.
[0071] In an embodiment, the assignment of quantization indices of reconstruction levels within the set of reconstruction levels 48 (or quantization set) follows the following rules: For illustrative purposes, the reconstruction levels in FIG. 8 are assigned using their associated quantization indexes (the quantization indexes are given by the numbers under the circles representing the reconstruction levels). If the set of reconstruction levels includes a reconstruction level equal to 0, then a quantization index equal to 0 is assigned to the reconstruction level equal to 0. A quantization index equal to 1 is assigned to the smallest reconstruction level equal to or greater than 0, a quantization index equal to 2 is assigned to the next reconstruction level equal to or greater than 0 (i.e., the second smallest reconstruction level equal to or greater than 0), and so on. Or, in other words, reconstruction levels equal to or greater than 0 are assigned integer numbers equal to or greater than 0 (i.e., using 1, 2, 3, etc.) in increasing order of their values. Similarly, a quantization index −1 is assigned to the largest reconstruction level less than 0, a quantization index −2 is assigned to the next (i.e., second largest) reconstruction level less than 0, and so on. Or, in other words, reconstruction levels less than 0 are assigned integer numbers less than 0 (i.e., −1, −2, −3, etc.) in decreasing order. For the example in FIG. 8, the described assignment of quantization indexes is illustrated for all quantization sets, except for set 1 in FIG. 8a (which does not include a reconstruction level equal to 0).
[0072] For quantization sets that do not include a reconstruction level equal to 0, one way to assign quantization indices to reconstruction levels is as follows: all reconstruction levels greater than or equal to 0 are assigned with quantization indices greater than or equal to 0 (by increasing their value order), and all reconstruction levels less than 0 are assigned with quantization indices less than 0 (by decreasing their value order). Therefore, the assignment of quantization indices essentially follows the same concept as for quantization sets that include a reconstruction level equal to 0, with the difference that there are no quantization indices equal to 0 (see the labels for quantization set 1 in Figure 8a). This aspect should be taken into account in the entropy coding of the quantization indexes. For example, a quantization index is often transmitted by encoding its absolute value (ranging from 0 to the maximum supported value), and for absolute values different from 0, the sign of the quantization index is also encoded. If a quantization index equal to 0 is not available, the entropy coding can be modified to transmit the absolute level minus 1 (the value for the corresponding syntax element ranging from 0 to the maximum supported value), and the sign is always transmitted. Alternatively, the assignment rule for assigning quantization indices to reconstruction levels could be modified. For example, one of the reconstruction levels close to zero could be assigned with a quantization index equal to 0. The remaining reconstruction levels would then be labeled according to the following rule: quantization indices equal to or greater than 0 are assigned to reconstruction levels equal to or greater than the reconstruction level with a quantization index equal to 0 (the quantization index increases with the value of the reconstruction level), and quantization indices less than 0 are assigned to reconstruction levels less than the reconstruction level with a quantization index equal to 0 (the quantization index decreases with the value of the reconstruction level). One possibility for such assignment is illustrated in FIG. 8a by the numbers in parentheses (if no number in parentheses is given, other numbers apply).
[0073] As mentioned above, in an embodiment, two different sets 52 of reconstruction levels (which we also call quantization sets) are used, and the reconstruction levels within both sets 52 represent integer multiples of the quantization step size Δ, including the case where different quantization step sizes are used for different transform coefficients within a transform block (e.g., by specifying a quantization weight matrix), and the case where the quantization step size is modified on a block-by-block basis (e.g., by sending a block quantization parameter within the bitstream).
[0074] The use of reconstruction levels representing integer multiples of the quantization step size allows for a computationally less complex algorithm for the reconstruction of the transform coefficients at the decoder side. This is illustrated below based on the example of FIG. 8c (similar simple algorithms exist for other configurations, in particular the sets shown in FIGS. 8a and 8b). In the configuration shown in FIG. 8, a first quantization set, set 0, includes all even integer multiples of the quantization step size, and a second quantization set, set 1, includes all odd integer multiples of the quantization step size plus a reconstruction level equal to 0 (and which is included in both quantization sets). The reconstruction process 62 for the transform coefficients 13′ could be implemented similarly to the algorithm specified in the pseudocode of FIG. 9.
[0075] In the pseudocode of Figure 9a, level[k] refers to the quantization index 56 sent in the data stream 14 for the transform coefficient 13', and setId[k] (equal to 0 or 1) specifies the identifier of the current set 48 of reconstruction levels (which is determined 54 based on the previous quantization index 58 in the reconstruction order 14, described in more detail below). The variable n represents the integer multiple of the quantization step size given by the quantization index level[k] and the set identifier setId[k]. If the transform coefficient 13' is coded using the first set of reconstruction levels (setId[k] == 0) that includes an even integer multiple of the quantization step size, then the variable n is twice the transmitted quantization index. If the transform coefficients 13′ are coded using a second set of reconstruction levels (setId[k]==1), we have three cases: (a) if level[k] is equal to 0, then n is also equal to 0; (b) if level[k] is greater than 0, then n is equal to 2 times the quantization index level[k] minus 1; (c) if level[k] is less than 0, then n is equal to 2 times the quantization index level[k] plus 1. This can be specified using the sign function:
number
[0076] Then, if the second quantization set is used, the variable n is equal to twice the quantization index level[k] minus the sign function sign(level[k]).
[0077] Once the variable n (which specifies the integer element of the quantization step size) is determined, the reconstructed transform coefficients
number
number
[0078] As mentioned above, the quantization step size
number
number
number
number
number
[0079] Another modulation in Figure 9b, related to Figure 9a, is that the switching between the two sets of reconstruction levels is implemented using a triplet if-then-else operator (a?b:c), as known from programming languages such as C, and the modulation could be reversed but also applied to Figure 9a.
[0080] Besides the selection of the set of reconstruction levels 52 described above, another important design aspect of dependent scalar quantization is the algorithm 54 used to switch between the defined quantization sets (sets of reconstruction levels). The algorithm 54 used determines the "packing density" that can be achieved in the N-dimensional space 20 of transform coefficients (and thus also in the n-dimensional space 20 of reconstructed samples). Ultimately, higher packing density results in increased coding efficiency. The selection process 54 may use a fixed selection rule for each transform coefficient, where the chosen set of reconstruction levels depends on the specific number of the immediately preceding reconstruction level in the coding order (this number is typically 2), as described below. This rule may be implemented using a state transition table or lattice structure, as further outlined below.
[0081] A particularly exemplary method for determining 54 the reconstruction set 48 for the next transform coefficient is based on partitioning a quantizer set 52, as illustrated in FIG. 10a for a particular example. Note that the quantizer set shown in FIG. 10a is the same quantizer set as in FIG. 8c. Each of the two (or more) quantizer sets 52 is partitioned into two subsets. For the example in FIG. 10a, the first quantizer set (labeled Set 0) is partitioned into two subsets (labeled A and B), and the second quantizer set (labeled Set 1) is also partitioned into two subsets (labeled C and D). Although it is not the only possibility (options are discussed below), the division for each quantization set is typically done by associating immediately neighboring reconstruction levels (and thus neighboring quantization indices) with different subsets. In an embodiment, each quantization set is divided into two subsets. In Figures 8 and 10a, the division of a quantization set into subsets is indicated by open and filled circles.
[0082] For the embodiment illustrated in Figures 10a and 8c, the following division rules apply: Subset A consists of all even quantization indices of quantization set 0; Subset B consists of all odd quantization indices of quantization set 0; · Subset C consists of all even quantization indices of quantization set 1; Subset D consists of all odd quantization indices of quantization set 1. Note that the used subset is typically not explicitly stated inside the bitstream. Instead, it can be extracted based on the quantization set used (e.g., set 0 or set 1) and the quantization indexes actually transmitted. For the division shown in FIG. 10a, the subsets can be extracted by the bitwise "and" transform index level operation, and1. Subset A consists of all quantization indexes in set 0 for (level &1) equal to 0, subset B consists of all quantization indexes in set 0 for (level &1) equal to 1, subset C consists of all quantization indexes in set 1 for (level &1) equal to 0, and subset D consists of all quantization indexes in set 1 for (level &1) equal to 1.
[0083] In an embodiment, the quantization set (set of allowable reconstruction levels) used to reconstruct the current transform coefficient is determined based on the subset associated with the last two or more quantization indexes. As an example, the two last subsets used (given by the last two quantization indexes) are shown in the table of FIG. 10b. The table should be read as follows: the first subset given in the first table column represents the subset for the immediately preceding coefficient, and the second subset given in the first table column represents the subset for the coefficient preceding the immediately preceding coefficient. The determination 54 of the quantization set 48 specified in this table refers to a specific embodiment. In another embodiment, the quantization set 48 for the current transform coefficient 13′ is determined by the subset with the last three or more quantization indexes 58. For the first transform coefficient of a transform block, we do not have any data about the subset of the preceding transform coefficients. In an embodiment, a predefined value is used in these cases. In a further embodiment, we estimate the subset A for all unavailable transform coefficients. That is, if we reconstruct the first transform coefficient, the two previous subsets are estimated as "AA", and thus, consistent with the table in Figure 10b, quantization set 0 is used. For the second transform coefficient, the subset of the immediately previous quantization index is determined by its value (because set 0 is used for the first transform coefficient, and the subset is either A or B), however, the second last quantization index (which does not exist) is estimated to be equal to A. Of course, any other rule can be used to estimate the default value for the absent quantization indexes. Other syntax elements for extracting the default subset for the absent quantization indexes can also be used. As yet another alternative, the last quantization index of the previous transform block for initialization can also be used.For example, the concept of dependent quantization of transform coefficients may be used by means of applying dependence across transform block boundaries (but only to boundaries of larger blocks such as CTUs or slices, images, etc.).
[0084] It should be noted that the subset of quantization indices (A, B, C, or D) is determined by using a quantization set (set 0 or set 1) and a quantization set (e.g., A or B for set 0 and C or D for set 1). The selected subset in a quantization set is also called a path (because it specifies a path if we represent the dependent quantization process as a lattice structure, as will be explained later). In our convention, a path equals 0 or 1. Then, subset A corresponds to path 0 in set 0, subset B corresponds to path 1 in set 0, subset C corresponds to path 0 in set 1, and subset D corresponds to path 1 in set 1. Therefore, the quantization set 48 for the next transform coefficient following coefficient 13' is also determined individually by the quantization set (set 0 or set 1) and pass (pass 0 or pass 1) involving the two (or more) last quantization indices 58 of the current coefficient 13', and the set 48 for coefficient 13' is determined by the indices 58 for the two previous coefficients 13'. In the table of FIG. 10b, the associated quantization sets and passes are specified in the second column. For the first column, the first entry represents the set and pass for the immediately previous coefficient, and the second entry represents the set and pass for the coefficient preceding the immediately previous coefficient.
[0085] It should be noted that paths are often determined by simple arithmetic operations. For example, for the configuration shown in FIG. 10a, the path is
number
[0086] A variable pass can be defined as any binary function of the previous quantization index (transform coefficient level) level[k]: path=binFunction(level[k]). For example, as an alternative function, we mean that if the last transform coefficient level is equal to 0, then the variable pass is equal to zero, and if the last transform coefficient level is not equal to 0, then the variable pass is equal to 1.
number
[0087] The transformation between quantization sets (set 0 and set 1) can be fully represented by a state function. An example of such a state variable is shown in the last column of the table in FIG. 10b. In this example, the state variable has four possible values (0, 1, 2, 3). On the one hand, the state variable specifies the quantization set to be used for the current transform coefficient. In the example table in FIG. 10b, quantization set 0 is used only if the state variable is equal to 0 or 1, and quantization set 1 is used only if the state variable is equal to 2 or 3. On the other hand, the state variable also specifies the possible transformations between quantization sets. By using state variables, the rules in the table in FIG. 10b can be described with a small state transformation table. As an example, the table in FIG. 10c specifies a state transformation table for the rule given in the table in FIG. 10b. Given the current state, it specifies the quantization set for the current transform coefficient (second column). It further specifies the state transformation based on the pass with the chosen quantization index (given a quantization set, the pass specifies the subset A, B, C, D used). Note that by using the concept of state variables, it is not necessary to actually retain the chosen subset. When reconstructing the transform coefficients for a block, it is sufficient to upgrade the state variables and determine the pass with the used quantization index.
[0088] In an embodiment, the paths are given by the parity of the quantization indexes. Becomes the current quantization index
number
number
[0089] In one embodiment, a state variable with four possible values is used. In other embodiments, state variables with a different number of possible values are used. Of particular interest are state variables where the possible values for the state variable are integer powers of two, i.e., 4, 8, 16, 32, 64, etc. It should be noted that in certain configurations (such as those given in the tables of FIGS. 10b and 10c), a state variable with four possible values is equivalent to an approach in which the current quantization set is determined by a subset of the last two quantization indexes. A state variable with eight possible values corresponds to a similar approach in which the current quantization set is determined by a subset of the last three quantization indexes. A state variable with 16 possible values corresponds to an approach in which the current quantization set is determined by a subset of the last four quantization indexes, etc. Although it is generally preferable to use state variables with possible values equal to integer powers of two, there is no limitation to this configuration.
[0090] Using the concept of a state variable, the current state, such a current quantization set 54 is determined by the previous state (in reconstruction order) and the previous quantization index individually. However, for the first transform coefficient of a transform block, there is no previous state and no previous quantization index. Therefore, it is necessary that the state for the first transform coefficient of the block be defined individually. Different possibilities exist. It should be noted that the state permutation settings are from 0 to 3, where, in this example, there are in fact no two previous coefficients before the first coefficient. Possible choices are: The first state for the transformation block 10 is always set equal to a constant predefined value. In an embodiment, the first state is set equal to 0. The initial state value is sent explicitly as part of the bitstream. This implies an approach where only a subset of the possible state values can be indicated by the corresponding syntax element. The values of the first states are derived based on other syntax elements for the transform block 10. That is, even if the corresponding syntax element (or elements) are used to convey other features to the decoder, they are also used to derive the first states for the dependent scalar quantization. For example, the following approach can be used: in entropy coding of quantization (cp above), the position of the first non-zero quantization index in the coding order 140 can be transmitted (e.g., using x and y coordinates) before the actual value of the quantization index is transmitted (all quantization indexes preceding this first non-zero quantization index in the coding order 14 are not transmitted and are presumed to be equal to 0). The position of the first non-zero quantization index can also be used to extract initial values for state variables. As a simple example, if M represents the number of quantization indexes that are not specified as equal to 0 by the position of the first non-zero quantization index in the coding order, then the initial state can be set equal to s = M%4, where the operator % denotes modulo operation. · The concept of state migration can be applied across transformation block boundaries. That means that the first state for a transform block is set equal to the last state of the previous transform block in the coding order (after the final state upgrade). That is, for the first quantization index, the previous state is set equal to the last state of the last transform block, and the previous quantization index is set equal to the last quantization index of the last transform block. In other words, the variable coefficients of multiple variable blocks (or a subset of variable coefficients that are not estimated to be equal to 0 depending on the position of the first non-zero quantization index transmitted in block 10 as the coding order continues) are treated as a combination for dependent scalar quantization. Only if certain conditions are met, multiple transform blocks are not treated as a combination. For example, multiple transform blocks may be considered as a combination only if they represent blocks obtained by the latest block dividing operation (i.e., they represent blocks obtained by the last division level) and / or they represent coded blocks in an image.
[0091] The concept of state transitions for dependent scalar quantization allows for a low-complexity implementation for the reconstruction of transform coefficients at the decoder. An example of the reconstruction process of transform coefficients for a single transform block is shown in Figure 11 using C-style pseudocode.
[0092] In the pseudocode of FIG. 11, the index k identifies the reconstruction order 14 of the transform coefficients 13. Note that in the example code, the index k decreases in the reconstruction order 14. The last transform coefficient in the reconstruction order has index k equal to 0. The first index kstart identifies the reconstruction index (or, more precisely, the inverse reconstruction index) of the first reconstructed transform coefficient. The variable kstart may be set equal to the number of transform coefficients in the transform block minus 1, or it may be set equal to the index of the first non-zero quantization index in the encoding / reconstruction order (e.g., if the position of the first non-zero quantization index is transmitted in the applied entropy coding method). In the latter case, all previous transform coefficients (with index k>kstart) are assumed to be equal to 0. The reconstruction process for each single transform coefficient is the same as in the example of FIG. 9b. For the example in FIG. 9b, the quantization index is represented by level[k], and the associated reconstructed transform coefficient is represented by trec[k].
[0093] The state variable is represented by "state". Note that in the example of Figure 11, the state is set equal to 0 at the beginning of the transform block. However, as mentioned above, other initializations (e.g., based on the values of some syntax elements) are possible. The 1D table setId[] specifies the quantization sets with different values of the state variables, and the 2D table state_trans_table[][] specifies the state transitions given by the current state (first argument) and (second argument). In the example, the paths are given by the parity of the quantization index (bitwise and using the operator &), but as mentioned above, other concepts (in particular, other binary functions of level[k]) are possible. Examples in C-style syntax for the tables setId[] and state_trans_table[][] are given in Figure 12 (these tables are identical to the tables in Figure 10c).
[0094] Instead of using the table state_trans_table[][] to determine the next state, one can use arithmetic operations giving the same result. Similarly, the table setId[] can also be implemented using arithmetic operations (for the given example, the value of the table can be represented by shifting one bit to the right, i.e., "state>>1"). Or, a combination of the 1d table setId[] and a table lookup with a sign function could be implemented using arithmetic operations.
[0095] In another embodiment, all quantization indices equal to 0 are excluded from the state transitions and the dependent reconstruction process. The information whether a quantization index is equal to or not equal to 0 is primarily used to divide the transform coefficients into zero and non-zero transform coefficients (possibly related to the position of the first non-zero quantization index). The reconstruction process for dependent scalar quantization is only applied to the ordered set of non-zero quantization indices. Transform coefficients with quantization indices equal to 0 are simply set equal to 0. The corresponding pseudocode is shown in FIG. 13.
[0096] In another embodiment, a block of transform coefficients is decomposed into sub-blocks 12 (see FIG. 4), and the syntax includes a flag indicating whether the sub-block contains any non-zero transform coefficient levels. Then, in one embodiment, quantization coefficients of sub-blocks that only contain quantization indices equal to 0 (as indicated by the corresponding flag) are included in the dependent reconstruction and state transition process. In another embodiment, the quantization indices of these sub-blocks are excluded from the dependent reconstruction / dependent quantization and state transition process of FIGS. 7a and 7b.
[0097] The state transitions of the dependent quantization performed at 54 can also be represented using a lattice structure, as shown in FIG. 14. The lattice shown in FIG. 14 corresponds to the state transitions specified in the table of FIG. 10c. For each state, there are two paths connecting the state for the current transform coefficient with the two possible states for the next transform coefficient in the reconstruction order. The paths are labeled Path 0 and Path 1, and this number corresponds to the path variable introduced above (which, for the embodiment, is equal to the parity of the quantization index). Note that each path individually specifies a subset (A, B, C, or D) for the quantization indexes. In FIG. 14, the subsets are specified in parentheses. Given an initial state (e.g., State 0), the paths through the lattice are individually specified by the transmitted quantization indexes.
[0098] In the example of Figure 14, the states (0, 1, 2 and 3) for scan index k have the following properties (cp. also Figure 15): State 0: The previous quantization index level [k-1] specifies the reconstruction level of set 0, and the current quantization index level [k] specifies the reconstruction level of set 0. State 1: The previous quantization index level [k-1] specifies the reconstruction level of set 1, and the current quantization index level [k] specifies the reconstruction level of set 0. State 2: The previous quantization index level [k-1] specifies the reconstruction level of set 0, and the current quantization index level [k] specifies the reconstruction level of set 1. State 3: The previous quantization index level [k-1] specifies the reconstruction level of set 1, and the current quantization index level [k] specifies the reconstruction level of set 1.
[0099] A lattice consists of a concatenation of so-called basic trellis cells. An example of such a basic lattice cell is shown in FIG. 15. Note that one is not limited to lattices with four states. In other embodiments, the lattice can have more states. In particular, a number of states representing an integer power of two is appropriate. Even if the lattice has more than two states, each node for a current transform coefficient is generally associated with two states for a previous transform coefficient and two states for a next variable coefficient. However, it is also possible for a node to be associated with more than two states for a previous transform coefficient or more than two states for a next transform coefficient. Note that a fully associated lattice (each state is associated with all previous states and all states for a next variable coefficient) corresponds to independent scalar quantization.
[0100] In an embodiment, the initial state cannot be freely chosen (because it would require some side information rate to transmit this information to the decoder). Instead, the initial state is either set to a predefined value, or its value is derived based on other syntax elements. In this case, not all paths and states are available for the first transform coefficient (or the first transform coefficient starting with the first non-zero transform coefficient). As an example for a four-state trellis, Figure 16 shows the trellis structure for the case where the initial state is set equal to 0.
[0101] In one embodiment of dependent scalar quantization, there are different sets 52 of allowable reconstruction levels (also called quantization sets) for transform coefficients 13. The quantization set 48 for the current transform coefficient is determined 54 based on the value 58 of the preceding quantization index. If we consider the example in FIG. 10a and compare the two quantization sets, it is clear that the distance between a reconstruction level equal to zero and a neighboring reconstruction level is greater for set 0 than for set 1. Therefore, if set 0 is used, the probability that a quantization index is equal to 0 is greater, and if set 1 is used, it is less. In an embodiment, this effect is exploited in the entropy coding of quantization indexes 56 within bitstream 14 by switching a code table or probability model based on the quantization set (also called state) used for the current quantization index. This concept is illustrated in FIGS. 17a and 17b, typically using binary arithmetic coding.
[0102] In accordance with the codename table or probability model, the appropriate switching of all preceding quantization index paths (with a subset of the used quantization set) must be known during entropy decoding of the current quantization index (or binarization decision corresponding to the current quantization index). Therefore, the transform coefficients 13 need to be coded in a reconstruction order 14. Therefore, in the embodiment, the coding order 14 of the transform coefficients 13 is equal to their reconstruction order 14. In contrast to the embodiment, any coding / reconstruction order of the quantization indexes allows for any other individually defined order, such as the zigzag scan order specified in the video coding standards H.262|MPEG-2 Video or H.264|MPEG-4 AVC, the orthogonal scan specified in H.265|MPEG-H HEVC, or the horizontal or vertical scan additionally specified in H.265|MPEG-H HEVC.
[0103] In an embodiment, the quantization index 56 is coded using binary arithmetic coding. To that end, the non-binary quantization index 56 is first mapped 80 onto a sequence 82 of binarization decisions 84 (commonly referred to as binary values). The quantization index 56 is often transmitted as an absolute value 88 and, for absolute values greater than 0, a code 86. While the code 86 is transmitted as a single binary value, there are many possibilities for mapping the absolute value 88 onto the sequence 82 of binarization decisions 84. Four examples of binarization methods are shown in the table of FIG. 18. The first is so-called unary binarization. Here, a first binary value 84a indicates whether the absolute value is greater than 0. A second binary value 84b (if present) indicates whether the absolute value is greater than or equal to 1, and a third binary value 84c (if present) indicates whether the absolute value is greater than or equal to 2, etc. The second example is exponential-Golomb binarization, which consists of a unary portion 90 and a fixed-length portion 92. As a third example, a binarization scheme is shown, which consists of two unary coded binary values 94 and a suffix portion 96, represented using exponential-Golomb codes. Such binarization schemes are often used in practice. The unary binary values 94 are typically coded using an adaptive probability model, while the suffix portion is coded using a constant probability model with pmf(0.5, 0.5). The number of unary binary values is not limited to be equal to 2; it can be more than 2, and it can actually vary for transform coefficients within a transform block. Instead of exponential-Golomb codes, any other uniquely decodable code can be used to represent the suffix portion 96. For example, an adaptive Rice code (or a similar parameterized code) may be used, where the actual binarization depends on the value of the quantization index already coded for the current transform block. The final binarization scheme example in FIG. 18 also begins with a unary portion 98 (where the first binary value specifies whether the absolute value is greater than 0, and the second binary value specifies whether the absolute value is greater than or equal to 1).However, the next binary value 100 specifies the parity of the absolute value, and finally, the remainder is represented using an Exponential-Golomb code 102 (note that a separate Exponential-Golomb code is used for each of the two values of the parity binary). As should be noted for the previous binarization scheme, the Exponential-Golomb code 102 can be replaced with any other uniquely decodable code (e.g., a parameterized Rice code, a unary code, or any combination of different codes), and the number of 98 unary encoded binaries can be increased or decreased, or even to accommodate different transform coefficients in the transform block.
[0104] At least the binary portion of the magnitude is typically coded using an applied probability model (also called a context). In an embodiment, one or more binary probability models are selected 103 based on the quantizer set 48 (or, more generally, the corresponding state variable) for the corresponding transform coefficient to which the quantization index 56 belongs. The selected probability model or context can depend on multiple parameters or characteristics of the quantization indexes already transmitted, but one of the parameters is the quantizer set 48 or state to be applied to the coded quantization index.
[0105] In an embodiment, the syntax for transmitting the quantization index of a transform block includes a binary value such as 84a, or a first binary value among 90, 94, 98 that specifies whether the quantization index 56 is equal to zero or whether it is not equal to zero. The probability model used to encode this binary value is selected from a set of two or more probability models, or context, or context models. The selection of the probability model used depends on the quantization set 48 (i.e., the set of reconstruction levels) applied to the corresponding quantization index. In another embodiment, the probability model depends on the current state variable (the state variable containing the quantization index used), or in other words, the quantization set 48, or the parity 104 of the two (or more) immediately previous coefficients.
[0106] In a further embodiment, the syntax for transmitting the quantization index includes a binary value specifying whether the absolute value of the quantization index (transform coefficient level) is greater than or equal to 1, such as 84b, or a second binary value 90, 94, or 98. The probability model used to code this binary value is selected 103 from a set of two or more probability models. The selection of the probability model used depends on the quantizer set 48 (i.e., the set of reconstruction levels) applied to the corresponding quantization index. In another embodiment, the probability model used depends on the current state variables (the state variables include the quantizer set used), or in other words, the quantizer set 84, and the parity 104 of the immediately previous coefficient.
[0107] In another embodiment, the syntax for transmitting the quantization indexes of a transform block includes a binary value specifying whether the quantization index is equal to or not equal to zero, and a binary value specifying whether the absolute value of the quantization index (variable coefficient level) is greater than or equal to 1. The probability model selection 103 for both binary values depends on the quantization set or state variable for the current quantization index. Instead, only the probability model selection 103 for the first binary value (i.e., the binary value indicating whether the quantization index is equal to or not equal to zero) depends on the quantization set or state variable for the current quantization index.
[0108] The state variables determining quantization set 48, or 48 and 104, can only be used to select a probability model if (at least) the pass variables (given by parity) for the preceding quantization index in the coding order are known. This is similar to HEVC, for example, in the case where quantization indexes are coded based on 4x4 subblocks for each subblock, but transmitting quantization indexes using corresponding multiple passes on a 4x4 arrangement of transform coefficients. In HEVC, in the first pass on a 4x4 subblock, a flag sig_coeff_flag is transmitted, which indicates whether the corresponding quantization index is different from zero. In the second pass, for coefficients with sig_coeff_flag equal to 1, a flag coeff_abs_level_greater1_flag is transmitted, which indicates whether the absolute value of the corresponding quantization index is greater than or equal to 1. At most eight of these flags are transmitted for a subblock. Next, the flag coeff_abs_level_greater2_flag is transmitted for the first quantization index (if any), with coeff_abs_level_greater1_flag equal to 1. Then, a sign flag is transmitted for the entire sub-block. And finally, the non-binarized syntax element coeff_abs_level_remaining is transmitted, specifying the residual for the absolute value of the quantization index (variable coefficient level). Adaptive Rice coding may be used to transmit the non-binarized syntax element coeff_abs_level_remaining. In this case, there is not enough information to determine 48 and 104 in the first pass.
[0109] The binary coding order must be changed to use a probabilistic model selection that depends on the current set of reconstruction levels (or state variables).
[0110] It has two basic properties: In an embodiment, all binary values 82 specifying the absolute value 88 of a quantization index are coded consecutively, i.e., the binary values 82 (relative to the absolute value 88) of all preceding quantization indexes 56 in the coding / reconstruction order 14 are coded before the first binary value of the current quantization index is coded. The sign binary value 86 may or may not be coded (indeed, it may be determined using the quantization set), i.e., in a second pass over the transform coefficients 13 along the pass 14. The sign binary value for a subblock may be coded after the binary value for the subblock's absolute value, but before any binary value for the next subblock. In this case, the least significant bit portion of the binary representation of the previous absolute quantization index, such as the LSB, may be used to control the selection 54. In another embodiment, only a subset of the binary values specifying the absolute values of the quantization indexes is coded consecutively in the first pass on the transform coefficients. However, these binary values individually specify the quantization sets (or state variables). The remaining binary values are coded in one or more further passes on the transform coefficients. Such an approach is possible based on a binarization scheme similar to the one shown in the right column of the table in FIG. 18. This includes, in particular, an approach in which the binarization includes an explicit parity flag (as one example illustrated in the right column of FIG. 18). If we assume that the pass is specified by the parity of the quantization index, it is sufficient to send the unary part 98, the parity binary value 100, in the first pass on the transform coefficients. However, it is possible for the first pass to include one or more further binary values. The remaining binary values included in 102 can be transmitted in one or more further passes. In other embodiments, a similar concept is used. For example, the unary binary number can be modified or even constructed based on already transmitted symbols, i.e., along pass 14. And, a different adaptive or non-adaptive code (e.g., an adaptive Rice code as in H.265|MPEG-H HEVC) can be used instead of an Exponential-Golomb code. Different passes can be used based on sub-block, where the binary values for a sub-block are coded in multiple passes, but all binary values of a sub-block are transmitted before any binary values of the next sub-block are transmitted.
[0111] Figure 17b shows a method for the decoding process of the quantization indexes 56, thus showing the inverse of Figure 17a. The entropy decoding 85' of the binary values 82 and the mapping 80 of the binary values 82 to the levels 56 (the inverse of binarization) perform the inverse, as does the context extraction 103 in the same way.
[0112] One aspect to note is that dependent quantization of transform coefficients may be advantageously combined with entropy coding when the choice of a probability model for one or more binary values of the binarized representation of quantization indices (also called quantization levels) depends on the quantizer set 48, i.e., the set of applicable reconstruction levels along path 14, or on the corresponding state variables for the current quantization index, i.e., 48 and 104. The quantizer set 48 (or state variables) is given by the quantization index (or a subset of binary values representing the quantization index) for the previous transform coefficient in the coding and reconstruction order.
[0113] According to some embodiments, the described selection of probability models is combined with one or more of the following entropy coding embodiments: The sending of a flag for a transform block specifies whether any quantization index for the variable block is not equal to zero or whether all quantization indexes for the variable block are equal to zero. Dividing the coefficients of a transform block into multiple sub-blocks (at least for transform blocks 10 that exceed a predetermined defined size, given by the size of the block or the number of included samples). If a transform block is divided into multiple sub-blocks, a flag is transmitted for one or more of the sub-blocks 12 (unless it is inferred based on already transmitted syntax elements) specifying whether the sub-block contains non-zero quantization indices. Sub-blocks may also be used to specify the binary coding order. For example, binary coding can be divided into sub-blocks 12 so that all binary values of a sub-block are coded before any binary values of the next sub-block are transmitted. However, binary values for a particular sub-block can be coded in multiple passes on the transform coefficients within this sub-block. For example, all binary values specifying the absolute values of the quantization indices for a sub-block may be coded before any sign binary values are coded. As mentioned above, the binary values for the absolute values can be divided into multiple passes. Transmission of the location of the first non-zero quantization index in the coding order. The location can be transmitted as x and y coordinates specifying a location in a 2D array of transform coefficients 13, it can be transmitted as an index in the scan order, or it can be transmitted by any other means. As shown in FIG. 19a, the transmitted location of the first non-zero quantization index 120 (or transform coefficient 120) in the coding order 14 is assumed to be equal to zero, identifying all variable coefficients preceding the identified coefficient 120 in the coding order 14 (not hatched in FIG. 19). Further data 15 is transmitted only for the coefficient at the identified location 120 and the coefficients following this coefficient 120 in the coding order 14 (marked hatched in FIG. 19). The example in FIG. 19a shows a 16x16 variable block 10 with 4x4 sub-blocks 12; the coding order 14 used is the sub-block-by-subblock orthogonal scan specified in H.265 | MPEG-H HEVC. It should be noted that the coding for the quantization index 56 at the specified location (the first non-zero coefficient in the coding order 120) may be special. For example, if the binarization for the absolute value of the quantization index comprises a binary value specifying whether the quantization index is not equal to 0, this binary value is not transmitted for the quantization index at the specified location (it is already known that the coefficient is not equal to 0); instead, the binary value is assumed to be equal to 1. The absolute values of the quantization indexes are transmitted using a binarization scheme consisting of several binaries coded using an adaptive probability model, and if the adaptively coded binaries do not already completely specify the absolute values, suffixes such as 92, 96, and 102 are coded in bypass mode of the arithmetic coding engine (non-adaptive probability model with pmf(0.5, 0.5) for all binaries). The binarization used for the suffixes may further depend on the values of the quantization indexes already transmitted. The binarization for the absolute value of the quantization index includes an adaptively coded binary value specifying whether the quantization index is not equal to 0, such as 84a or the first of 90, 94, and 98. The probability model (also called context) used to code this binary value is selected from a set of candidate probability (context) models. The selected candidate probability model is determined not only by the quantization set 48 (the set of applicable reconstruction levels) or the state variables (48 and 104) for the current quantization index, but also by the quantization index (or a binary value specifying a part of the quantization index, as described above) already transmitted for the transform block 10 or index 56 (or part thereof) of the coefficient 13 preceding the current quantization index in the coding order 16. In an embodiment, the quantization set 48 (or state variable) determines a subset (also called context set) of available probability models, and the already coded binary value for the quantization index determines the probability model used within this subset (context set). In a first embodiment, the internal probability model of a context set is determined by the value of coded_subblock_flag (which specifies whether a subblock contains a non-zero quantization index) of neighboring subblocks (similar to H.265|MPEG-H HEVC). In another embodiment, the used probability model in the context set is determined based on the values of already coded quantization indices (or already coded binary values for quantization indices) in a neighborhood 122 of the current variable coefficient. An example for such a neighborhood, which may be called table 122, is shown in FIG. 19b. In the figure, the current transform coefficient 13' is marked in black and the neighborhood 122 is shown hatched. Below, listed are some corresponding examples that can be extracted based on the values of quantization indices (or partially reconstructed quantization indices given by binary values transmitted before the current binary value) in the neighborhood, and then used to select a probability model for a predetermined context set: The number of quantization indices is not equal to 0 within the neighborhood region 122. This number can possibly be clamped to a maximum value; and / or The sum of the absolute values of the quantization indices of the neighboring regions 122. This number can be clamped to a maximum value; and / or The difference between the sum of the absolute values of the quantization indices in the neighboring region and the number of quantization indices inside the neighboring region that are not equal to 0. This number can be clamped to a maximum value. For binary coding orders for the absolute values of the quantization indices transmitted in multiple passes (as described above), the above-mentioned measures are calculated using the transmitted binary values. This means that instead of the complete absolute values (which are not known), partially reconstructed absolute values are used. In this context, integer factors of the quantization step size can be used instead of the quantization index (the integer number transmitted inside the bitstream). These factors are the variables n in the pseudocode examples in Figures 9a, 9b, 11, and 13. They represent the integer part of the quotient of the reconstructed transform coefficient and the quantization step size (rounded to the nearest integer). Additionally, other data available for decoding may be can be used to extract probabilistic models (either explicitly or as listed above) (In combination with measured data). Such data include: The position of the current transform coefficient 13' (x coordinate, y coordinate, diagonal number, or any combination thereof) of the block 10. The size of the current block (vertical size, horizontal size, number of samples, or any combination thereof). -Current conversion block 16 aspect ratio. The binarization for the absolute value of the quantization index includes an adaptively coded binary value specifying whether the absolute value of the quantization index is greater than or equal to 1. The probability model (also called a context) used to code this binary value is selected from a set of candidate probability models. The measured probability model is not only measured by the quantization set 48 (acceptable reconstruction levels) or the state variables for the current quantization index, i.e., 48 and 104, but it is also, or exclusively, determined by the quantization indexes already transmitted for the transform block. In one embodiment, the quantization set (or state variables) determines a subset of available probability models (also called a context set), and the data for the previously coded quantization indexes determines the probability model used within this subset (context set). In another embodiment, the same context set is used for all quantization sets / state variables, and the data for the previously coded quantization indexes determines the probability model used. Any of the methods described above (for the binary value specifying whether the quantization index is not equal to 0) can be used to select the probability model.
[0114] An adaptive selection between dependent and independent scalar quantization may be applied. In one embodiment, transform coding with dependent quantization (and potentially applied to entropy coding of the quantization indexes) is applied to all transform blocks. The exceptions are lossless coding modes that do not involve transforms (i.e., no quantization) or coded modes that do not involve transforms.
[0115] The codec may also provide two methods of transform coding: (a) conventional transform coding with independent scalar quantization, and (b) transform coding with dependent quantization. Which transform coding method (independent or transform coding with dependent scalar quantization) is used for a transform block may be explicitly signaled to the decoder using an appropriate syntax element, or may be derived based on syntax elements. Whether transform coding with dependent quantization is possible may also be explicitly signaled first, and then whether transform coding with dependent quantization (if possible) is actually used for the transform block is derived based on other syntax elements.
[0116] Explicit signaling may comprise one or more of the following methods: A high-level syntax element (typically a flag) is transmitted in a high-level syntax structure, such as a sequence parameter set, a picture parameter set, or a slice header, indicating whether transform coding with dependent quantization or conventional transform coding with independent scalar quantization is used. In this context, it is possible to enable / disable the use of dependent quantization only for luma or chroma blocks (or more generally, for blocks of a specific color channel). In another embodiment, the high-level syntax element indicates whether transform coding with dependent quantization is possible for the coded video sequence, image, or slice (it is also possible that transform coding with dependent quantization is only possible for luma or chroma blocks). If transform coding with dependent quantization is possible for the coded video sequence, the actual decision of which block, image, or slice transform coding may be used for may be derived based on specific block parameters (see below). Dedicated syntax elements may also be used to signal different settings for luma and chroma blocks. The syntax may include low-level syntax elements (i.e., block-based syntax elements) that specify whether transform coding with dependent quantization or conventional transform coding is used for the corresponding block. Such syntax elements may be transmitted on a CTU-by-CU basis, a transform block-by-transform block basis, etc. If such a syntax element is coded for an entity containing multiple transform blocks, it applies to all contained transform blocks. It may, however, only apply to transform blocks of a particular color channel (e.g., only luma blocks or only chroma blocks).
[0117] In addition to explicit signaling, the decision whether a transform block is coded using dependent or independent quantization of the transform coefficients could be derived based on parameters for the corresponding block. These parameters have to be coded (or derived based on other already transmitted syntax elements) before the actual quantization indices for the transform coefficients of the block are transmitted. Among other possibilities, the block-adaptive decision could be based on one or more of the following parameters: Block quantization parameter: for example, for blocks with a quantization parameter (QP) below a certain threshold (which may also be indicated in a high-level syntax construct), transform coding with dependent quantization is used, and for blocks with a QP above the threshold, conventional transform coding with independent scalar quantization is used. Position of the first non-zero quantization index in the coding order: Given the position of the first non-zero coefficient, the number of quantization indexes actually transmitted (excluding quantization indexes that are estimated to be equal to 0 based on the transmitted position of the first non-zero quantization index) can be extracted. For example, dependent quantization is used for blocks for which the number of actually transmitted quantization indexes is greater than a threshold (the threshold may or may not depend on the block size or the number of samples within the block), and independent scalar quantization is used for all other blocks. In addition to the position of the first non-zero coefficient, a flag indicating whether the subblock contains a non-zero quantization index can be used to determine the number of actually transmitted quantization indexes. In this case, the so-called coded subblock flag (indicating whether the subblock contains a non-zero coefficient) must be coded before the first quantization index is coded. Transform block size (horizontal size, vertical size or number of samples in a transform block). The transform (or type of transform) used to reconstruct the transform block. Such criteria could be applied to video codecs that support multiple transforms. Any combination of the criteria identified above.
[0118] For embodiments, transform coding with independent scalar quantization and conventional transform coding (with independent scalar quantization) can be used together on an image or slice, preferably using different relationships between quantization step sizes and quantization parameters for the two approaches. For example, a decoder may have one table for mapping block quantization parameters to quantization step sizes used for independent quantization and another table for mapping block quantization parameters to quantization step sizes used for dependent quantization. Alternatively, the decoder may have one table for mapping quantization parameters to quantization step sizes (as in H.265|MPEG-H AVC), but for transform blocks using dependent scalar quantization, the quantization step size is multiplied by a predefined factor before it is used in the transform coefficient reconstruction process. It should be noted that multiplication by the quantization step size to reconstruct the transform coefficients may actually be implemented as a combination of multiplication and bit shifting. The multiplication of the nominal quantization step size with a predefined factor may then be implemented as a bit-shift modulation and a scaling factor modulation. As a further alternative, for transforms, the offset for transform blocks coded using dependent quantization can be added to (or subtracted from) the quantization parameter before it is mapped to the quantization step size used.
[0119] Furthermore, for transform coding with dependent quantization, the quantization step size used can be adapted based on the number of transform coefficients. To that end, any of the techniques mentioned above for modulating the quantization size can be used (lookup tables, multiplication factors, scale modifications, and bit shift parameters). The number of quantization indexes actually transmitted can be determined in one or more of the following ways (or by any other means): The number of transmitted quantization indexes can be derived based on the position of the first non-zero quantization index in the coding order (see above). The number of transmitted quantization indices can be derived based on the position of the first non-zero quantization index in the coding order and the value of the transmitted coded sub-block flag (the coded sub-block flag indicates whether the sub-block has non-zero quantization indices or, alternatively, any non-zero transform coefficients). The number of transmitted quantization indexes (to determine the quantization step size) can be derived based on the number of non-zero quantization indexes in the transform block.
[0120] The quantization step size must be known for entropy coding of the quantization indexes. It should be noted that it is not necessary to use the Therefore, any syntax elements and extracted parameters can be transformed. It can be used to calculate the quantization step size used in the reconstruction of the coefficients.
[0121] As an example method for encoding, the following description is provided.
[0122] To obtain a bitstream that offers a very good trade-off between distortion (reconstruction quality) and bitrate, the quantization index is calculated using a Lagrangian cost measure.
number
number
number
number
[0123] However, as we discussed above, dependencies between transform coefficients can be represented using a lattice structure. For further embodiments, we will use the embodiment in Figure 10a as an example. A lattice structure for an example block of eight transform coefficients is shown in Figure 20. A path through the lattice (from left to right) indicates the possible state transitions for the quantization indexes. Note that each connection between two nodes represents the quantization indexes of a particular subset (A, B, C, D). If we
number
number
number
[0124] An example encoding algorithm for choosing appropriate quantization indices for a transform block could consist of the following main steps: 1. Set the rate-distortion cost for the initial state equal to 0. 2. For every transform coefficient in coding order, do the following: a. For each subset A, B, C, D, determine the quantization index that minimizes distortion for the given original transform coefficients. b. For all lattice nodes (0,1,2,3) for the current transform coefficient, do the following: i. Calculate the rate-distortion cost for two paths connecting the state for the previous transform coefficient with the current state. The cost is calculated based on the previous state and the cost
number
number
[0125] It should be noted that determining quantization indexes based on the Viterbi algorithm is significantly less complex than rate-distortion optimized quantization (RDOQ) for independent scalar quantization. Nevertheless, there are also simpler coding algorithms for dependent quantization. For example, starting from a predefined initial state (or quantization set), quantization indexes could be determined in the coding / reconstruction order by minimizing any cost measure, considering only the influence of the current quantization index. Given a given quantization index for the current coefficient (and all previous quantization indexes), the quantization set for the next transform coefficient is known. Thus, the algorithm can be applied to all transform coefficients in the coding order.
[0126] In summary, the encoders shown in Figures 1, 6b, 7a, and 17a, in any combination thereof, represent examples for the present media signal encoder. At the same time, the encoders shown in Figures 6a, 7b, and 17b, in any combination thereof, represent examples for the present media signal decoder. Nevertheless, just as a precautionary measure, and introduced primarily to illustrate the structure used by an HEVC encoder, it should be noted again that Figure 1 also forms one possible embodiment, i.e., an apparatus for predictively encoding video 211, consisting of a sequence of images 212 in a data stream 14. Block-wise predictive coding is used for this purpose. Furthermore, transform-based residual coding is typically used. The apparatus, or encoder, is indicated with reference numeral 210. Figure 21 depicts a corresponding decoder 220 (i.e., device 220) configured to predictively decode video 211' consisting of image 212' in an image block from data stream 14, and here typically employs transform-based residual decoding, where an apostrophe is used to indicate that image 212' and video 211', respectively, reconstructed by decoder 220 deviate from image 212 originally encoded by device 210 in terms of coding loss introduced by quantization of the prediction residual signal. While Figures 1 and 21 typically employ transform-based predictive residual coding, embodiments of the present application are not limited to this type of predictive residual coding, or even to predictive coding that includes coding and quantization of the prediction residual.
[0127] The encoder 210 is configured to subject the prediction residual signal to a spatial spectral transformation and to encode the prediction residual signal thus obtained in the data stream 14. Similarly, the decoder 220 is configured to decode the prediction residual signal from the data stream 14 and to subject the prediction residual signal thus obtained to a spatial spectral transformation.
[0128] Internally, the encoder 210 may comprise a prediction residual signal former 222, which generates a prediction residual 224 so as to measure the deviation of a prediction signal 226 from the original signal, i.e., the video 211 or the current picture 212. The former prediction residual signal 222 may, for example, be a subtractor which subtracts the prediction signal from the original signal, i.e., the current picture 212. The encoder 210 then further comprises a transform and quantization stage 2, which obtains a transformed prediction residual signal and subjects the prediction residual signal 224 to any of the above-mentioned transformations for subjecting the same to quantization. The quantization concept was described above with reference to FIG. 7a. The thus quantized prediction residual signal 224' is coded in the bitstream 14 and comprises the aforementioned quantization indexes 56. For encoding into the data stream 14, the encoder 210 may optionally comprise an entropy coder 234 that entropy-encodes the quantized and transformed prediction residual signal in the data stream 14. At the decoder side, a reconstructable quantized prediction residual signal 224' is generated based on the internally coded prediction residual signal 224' and subjected to an inverse transform, such as the one described above, to obtain reconstructed transform coefficients 13 and to subject the transform consisting of the coefficients 13 to an inverse transform, to obtain a prediction residual signal 224 that corresponds to the original prediction residual signal 224 except for quantization losses, and can be decoded from the data stream 14 by using an inverse quantization and transform stage 238 that dequantizes the prediction residual signal 224' according to Fig. 7b. A combiner 242 then recombines the prediction signal 226 and the reconstructed prediction residual signal 224', e.g., by addition, to obtain a reconstructed signal 246, i.e., a reconstruction of the original signal 212. The reconstructed signal 246 may be identical to the signal 212′, or may represent a pre-reconstructed signal subjected to in-loop filtering 247 to obtain the reconstructed signal 212′ and / or post-filtering (not shown in FIG. 1).A prediction module 244 then generates a predicted signal 226 on the basis of the signal 246, for example by using spatial prediction 248, i.e. intra-prediction, and / or temporal prediction 249.
[0129] The entropy coder 234 entropy codes not only the prediction residual 224'' in the data stream 14, but also other coded data representative of the image, such as the residual data 224'', prediction mode, prediction parameters, quantization parameters, and / or filter parameters. The entropy coder 234 encodes all this data in a lossless manner into the data stream 14.
[0130] Similarly, the decoder 220 may be composed of internally consistent components and interconnected in a manner consistent with a prediction loop formed by the prediction module 244, the combiner 242, and the inverse of stages 234 and 238. In particular, the entropy decoder 250 of the decoder 220 may decode an entropy-quantized spectral-domain prediction residual signal 224'' from the data stream. Contour extraction may be performed in a synchronous manner with the encoder. The result is, for example, coded data including the prediction residual data 224''. The dequantization and inverse transform module 252 then dequantizes this signal according to FIG. 7b and inverse transforms the same to obtain signal 224''', i.e., the residual signal in the spectral domain, from the prediction loop supplied by signal 224''' and corresponding to the loop in the encoder, prediction module 258, combiner 256 and any in-loop filter 247', the output of which yields a reconstructed signal based on the predicted residual signal 224''', thereby yielding the reconstructed signal, i.e., the video 211' or current picture 212' therefrom, as shown in FIG. 21.
[0131] The encoder 210 may set the same coding parameters, including prediction modes, motion parameters, etc., according to several rate-optimizing and distortion-related criteria, i.e., coding cost, and / or several rate control and optimization schemes, generally shown as encoder control 213 in FIG. 1 . As described above, the encoder 210 and decoder 220, and corresponding modules 244 and 258, respectively, support different prediction modes, such as intra-coding mode and inter-coding mode. The precision with which the encoder and decoder switch between these prediction modes, as determined by several decision modules 243, may correspond to the subdivision of each of the images 212 and 212′ into blocks. Note that some of these blocks may be solely intra-coded, and some may be solely inter-coded, and optionally, even further blocks may be blocks obtained using both intra- and inter-coding. Depending on the intra-coding mode, a prediction signal for a block is obtained based on the spatial, already coded / decoded neighborhood of each block. Several intra-coding sub-modes may exist among the selections, which represent similar types of intra-prediction parameters. For each block's internal directional intra-coding sub-mode, there may be directional or angular intra-coding sub-modes depending on the prediction signal for each block, which depend on estimating neighboring sample values along a specific direction.The intra-coding submode may include one or more further submodes, such as a DC coding mode according to a prediction signal for each block that assigns a DC value to all samples within the respective block, and / or a planar intra-coding mode according to a prediction signal for each block that is approximated or determined as a spatial distribution of sample values described by a system of equations over the sample positions of each block using a gradient and a planar offset defined by the system of equations based on neighboring samples. In comparison, according to an inter-prediction mode, a prediction signal for a block may be obtained, for example, by temporally predicting the block inner. For parameterization of the inter-prediction mode, a motion vector may be signaled in the data stream, and the motion vector may be sampled to obtain a prediction signal for each block, including spatially displacing the position of a previously coded image of video 211 with a previously coded / decoded image. In addition to the residual signal coding provided by the data stream 14, such as entropy-coded transform coefficient levels representing the quantized spectral domain prediction residual signal 224", the data stream 14 may also encode prediction-related parameters for assigning a block prediction mode, prediction parameters for the assigned prediction mode, such as motion parameters for inter-prediction modes, and any further parameters controlling the construction of the final prediction signal for the block using the assigned prediction mode and prediction parameters. Furthermore, the data stream may also comprise transmission parameters controlling the subdivision of the respective images 212 and 212' within the block, such as transform blocks from the transform coefficient block 10 obtained by transforming the individual transform blocks.The decoder 220 subdivides the image in the same way as the encoder did to assign the same prediction modes and parameters to the blocks, and then uses these parameters to perform the same prediction, e.g., an inverse transform on each transform coefficient block 10 to produce one of the transform blocks, resulting in the same prediction signal. That is, the prediction signal may be obtained by intra-image prediction using already reconstructed samples of spatial neighbors of the current image block, or by motion-compensated prediction using samples of an already fully reconstructed image.
[0132] FIG. 22 illustrates the relationship between the reconstructed signal, i.e., the reconstructed image 212′, on the one hand, and the combination of the prediction residual signal 224′″ conveyed by the data stream 14 and the prediction signal 226, on the other hand. As already indicated above, additional combinations may be made. The prediction signal 226 is illustrated in FIG. 22 as a subdivision of the image region within blocks 280 of various sizes, but this is merely an example. The subdivision may be any subdivision, such as a regular subdivision of the image region within rows and columns of blocks, or a multi-tree subdivision of the image 212 within leaf blocks of various sizes, such as a quadtree subdivision, or a mixture thereof, as illustrated in FIG. 22, in which the image region is subdivided first into rows and columns of tree root blocks that are then further subdivided according to a recursive multi-tree subdivision to result in blocks 280.
[0133] The prediction residual signal 224''' in Figure 22 is also illustrated as a subdivision of the image region inside the block 284. These blocks may be called transform blocks to distinguish them from the coding block 280. Essentially, Figure 22 illustrates the encoder 210 and the decoder 220 may use two different subdivisions inside the blocks, image 212 and image 212', respectively, i.e., one subdivision of the interior of the coding block 280 and another subdivision of the interior of the transform block 284. Both subdivisions may be the same, i.e., each block 280 together forms a transform block 284, and vice versa; however, Figure 22 illustrates the case where, for example, a subdivision within a transform block 284 forms an extension of a subdivision within the coding block 280, whereby any boundary between two blocks 280 covers the boundary between two blocks 284, or alternatively, referring to each block 280, coincides with one of the transform blocks 284, or with a group of transform blocks 284. However, the subdivisions may also be determined or selected independently of each other, whereby a transform block 284 could selectively cross block boundaries between blocks 280. Insofar as the subdivision into transform blocks 284 is considered, similar considerations are thus brought forward to the subdivision within blocks 280, i.e. blocks 284 may be the result of a regular subdivision of the image area within blocks arranged in rows and columns, the result of a recursive multi-tree of the image area, or a combination thereof, or any other type of subdivision. Just as an aside, blocks 280 and 284 are not limited to being square, rectangular, or any other shape. Furthermore, the subdivision of the current image 212 within blocks 280 in which the prediction signal is formed and the subdivision of the current image 212 within blocks 280 in which the prediction residual is coded may be the only subdivisions used for coding / decoding.These subdivisions form a determination of the precision with which the prediction signal is and residual coding is performed, but firstly, the residual coding may alternatively be performed without subdivision, and secondly, at other precisions than these subdivisions the encoder and decoder may set certain coding parameters, including some of the parameters mentioned above, such as prediction parameters, etc.
[0134] 22 illustrates the combination of the prediction signal 226 and the prediction residual signal 224''' to directly result in the reconstructed signal 212. However, it should be noted that one or more prediction signals 226 may be combined with the prediction residual signal 224''' to result in the image 212' according to other embodiments, in which prediction signals obtained from other perspectives or from other coding layers are coded / decoded in separate prediction loops, e.g., with separate DPBs.
[0135] In Figure 22, the transform blocks 284 have the following importance: the modules 228 and 252 perform their transforms in units of these transform blocks 284. For example, many codecs use some kind of DST or DCT for all transform blocks 284. Some codecs allow skipping the transform, whereby for some transform blocks 284, the prediction residual signal is directly coded in the spatial domain. However, according to the embodiments described below, the encoder 210 and the decoder 220 are configured to support several transform methods. For example, the encoder 210 and the decoder 220 DCT-II (or DCT-III), where DCT stands for Discrete Cosine Transform DSTIV, where DCT stands for discrete sine transform DCT-IV ·DSTVII Identity Transformation (IT) It can be equipped with:
[0136] Naturally, while module 228 supports all these forward transform versions, decoder 220, or inverse transform 252, supports the corresponding backward, or inverse, versions: Inverse DCT-II (or Inverse DCT-III) ·Inverse DCT-IV ·Inverse DCT-IV Inverted DCT-VII Identity Transformation (IT)
[0137] This allows different transforms to be applied in the horizontal and vertical directions, for example, DCT-II can be used in the horizontal direction and DCT-VII can be used in the vertical direction, or vice versa.
[0138] In any case, it should be noted that the set of supported transformations may comprise just one transformation, such as spatial to spectral or spectral to spatial, and beyond this, reference is made to the above description for further clarification and examples provided.
[0139] Thus, in the above description related to a video decoder, one or more blocks of samples, which may alternatively be referred to as image blocks and may also be referred to as spatial, reconstructed, or decoded samples, are reconstructed by predicting the samples using an intra- or inter-image to obtain a reconstructed block of video samples, analyzing quantization indexes from the bitstream, reconstructing transform coefficients using dependent quantization (the values of the reconstructed transform coefficients are calculated using the current and one or more previous quantization indexes), inverse transforming the reconstructed quantized coefficients to obtain a block of reconstructed prediction error samples, and adding the reconstructed prediction error samples to a prediction signal. Here, the dependent reconstruction of the transform coefficients may be achieved using two distinct (but not necessarily disjoint) sets of allowable reconstruction levels. The transform coefficients may be reconstructed in a predetermined reconstruction order by choosing one of a set of reconstruction levels based on the value of a previous quantization index in the reconstruction order; choosing one reconstruction level of the selected set using the quantization index for the current transform coefficient (the quantization index specifies an integer index in an ordered set of allowable reconstruction values). The selection of the set of allowable reconstruction values may be specified by the parity preceding the quantization index in the reconstruction order. The selection of the set of allowable reconstruction values may be realized using a state transition table (or equivalently, an arithmetic operation) with the following properties: at the beginning of the reconstruction method for a transform block, a state variable is set equal to a predetermined value; the states individually determine the set of allowable reconstruction values for the next transform coefficient; the state for the next transform coefficient is extracted using a lookup table in a 2d table, where the first index is given by the current state and the second index is given by the parity of the current quantization index.Alternatively, or in addition, the first set of allowable reconstruction levels may comprise all even integer multiples of the quantization step size; the second set of allowable reconstruction levels may comprise all odd integer multiples of the quantization step size; and in addition, the reconstruction level is equal to 0. Here, the quantization step size is determined by the position of the transform coefficient within the transform block (e.g., as specified by a quantization weight matrix). Additionally, or alternatively, the quantization step size may be determined by a quantization parameter that can be modified block by block on a block-by-block basis. The entropy coding may include binarization or absolute values of the quantization index values, and one or more binary values may be coded using an adaptive binarization probability model. For one or more binary values coded using the adaptive probability model, the used probability model may be selected from a set of adaptive probability models, and the probability model may be selected based on the set of allowable reconstruction levels (or states). The binarization for the quantization index may comprise a binary value specifying whether the quantization index is equal to or not equal to zero, the binary value being coded using an adaptive probability model chosen from a set of adaptive probability models, and a subset of the set of available probability models may be selected as a set of predetermined allowable reconstruction levels for the current quantization index (or state), and the probability model to be used from the selected subset may be determined using characteristics of the quantization indexes in a neighborhood (in the array of transform coefficients) of the current quantization index. The binary value resulting from the binarization of the quantization index may be entropy decoded in multiple passes over a transform block or sub-blocks of the transform block (representing a subset of the positions of the transform coefficients within the transform block).The binary values coded in the first pass may include any combination of the following binary values: a binary value specifying whether the quantization index is equal to or not equal to zero; a binary value specifying whether the absolute value of the quantization index is greater than or equal to one or not greater than or equal to one; a binary value specifying the parity of the quantization index (or the parity of the absolute value of the quantization index).
[0140] Even if some embodiments are described within an apparatus context, it is clear that the embodiments also refer to corresponding method descriptions, and thus apparatus blocks or structural elements correspond to corresponding method steps or method step features. Similarly, embodiments described within the context of or as method steps also refer to corresponding block or detailed descriptions or corresponding apparatus features. All or some method steps may be performed by (or using) a hardware apparatus, such as a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0141] The encoded data stream of the invention can be stored on a digital storage medium or can be transmitted over a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0142] Depending on the specific implementation conditions, embodiments of the invention may be implemented in hardware or in software. Implementation may be performed using a digital storage medium, such as a floppy disk, DVD, Blu-ray disk, CD, ROM, PROM, EPROM, EEPROM, or FLASH memory, having electronically readable control signals that cooperate with a programmable computer system on which the respective methods are executed. The digital storage medium may therefore be computer-readable.
[0143] Some embodiments according to the invention comprise a data carrier with electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.
[0144] Generally, embodiments of the present invention can be implemented as a computer program having program code which operates to perform one of the methods when the computer program runs on a computer, the program code may for example be stored on a machine readable carrier.
[0145] Other embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0146] In other words, an embodiment of the inventive method is thus a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0147] A further embodiment of the inventive method is thus a data storage medium (or digital storage medium or computer readable medium) having recorded thereon a computer program for performing one of the methods described herein. The data storage medium, digital storage medium or recorded medium is generally non-transitory or non-volatile.
[0148] A further embodiment of the inventive methods is thus a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals may for example be arranged to be transmitted over a data communication link, such as for example the Internet.
[0149] A further embodiment comprises a processing unit, eg a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0150] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0151] Further embodiments comprise an apparatus or system configured to transmit a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, or the like. The apparatus or system may, for example, comprise a file server for transmitting the computer program to the receiver.
[0152] In some embodiments, a programmable logic device (e.g., a field programmable gate array, FPGA) may be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, it is desirable for the methods to be performed by any hardware apparatus.
[0153] For example, the apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0154] The apparatus described herein, or any component of the apparatus described herein, may be implemented at least in part in hardware and / or software (computer programs).
[0155] For example, the methods described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and / or a computer.
[0156] The methods described herein, or any components of the methods described herein, may be implemented at least in part by hardware and / or software.
[0157] The above-described embodiments merely represent illustrative of the principles of this invention. It is understood that those skilled in the art will recognize modifications and variations of the arrangements and details described herein. As such, it is intended that the invention be limited only by the following claims, rather than by the specific details presented herein, and by the description and discussion of the embodiments.
[0158] References [1] ITU-T and ISO|IEC, “Advanced videocoding for audiovisual services,” ITU-T Rec. H.264 and ISO|IEC14406-10 (AVC), 2003. [2]ITU-T and ISO|IEC, “High efficiency video coding,” ITU-T Rec. H.265 and ISO|IEC 23008-10 (HEVC), 2013.
Claims
1. 1. A video decoder for decoding one or more pictures from a video data stream, comprising: decoding quantization indexes from the video data stream; determining a sign corresponding to the decoded quantization index indicating whether its value is positive or negative; determining a state variable that is updated based on the decoded quantization index and indicates whether the value is odd or even; deriving the value based on the sign, the state variables, and a magnitude corresponding to the decoded quantization index; determining a transform coefficient that is a product of a quantization step size and the derived value; Video decoder.
2. the magnitudes corresponding to the decoded quantization indexes include the decoded quantization indexes increased by a factor of two; 2. The video decoder of claim 1.
3. the state variable comprises a value of 0, 1, 2, or 3; 2. The video decoder of claim 1.
4. the derived value is an even number when the state variable contains a value of 0 or 1, and an odd number when the state variable contains a value of 2 or 3.
2. The video decoder of claim 1.
5. When the state variable is 0, if the parity corresponding to the decoded quantization index is 0, the state variable is updated to 0, and if the parity is 1, the state variable is updated to 2; If the state variable is 1, the state variable is updated to 2 if the parity is 0, and updated to 0 if the parity is 1; If the state variable is 2, the state variable is updated to 1 if the parity is 0, and updated to 3 if the parity is 1; When the state variable is 3, if the parity is 0, the state variable is updated to 3, and if the parity is 1, the state variable is updated to 1.
2. The video decoder of claim 1.
6. 1. A method for decoding one or more pictures from a video data stream, comprising: decoding quantization indexes from the video data stream; determining a sign corresponding to the decoded quantization index indicating whether its value is positive or negative; determining a state variable that is updated based on the decoded quantization index and indicates whether the value is odd or even; deriving the value based on the sign, the state variables, and a magnitude corresponding to the decoded quantization index; determining a transform coefficient that is a product of a quantization step size and the derived value; method.
7. the magnitudes corresponding to the decoded quantization indexes include the decoded quantization indexes increased by a factor of two; the state variable comprises a value of 0, 1, 2, or 3; The method of claim 6.
8. the derived value is an even number when the state variable contains a value of 0 or 1, and an odd number when the state variable contains a value of 2 or 3. The method of claim 6.
9. When the state variable is 0, if the parity corresponding to the decoded quantization index is 0, the state variable is updated to 0, and if the parity is 1, the state variable is updated to 2; If the state variable is 1, the state variable is updated to 2 if the parity is 0, and updated to 0 if the parity is 1; If the state variable is 2, the state variable is updated to 1 if the parity is 0, and updated to 3 if the parity is 1; When the state variable is 3, if the parity is 0, the state variable is updated to 3, and if the parity is 1, the state variable is updated to 1. The method of claim 6.
10. 10. A method for implementing a method of claim 6, comprising: storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 6; Non-transitory storage media.
11. 1. A video encoder for encoding one or more pictures into a video data stream, comprising: one or more processors configured to perform operations including encoding quantization indexes into the video data stream; a sign is determined corresponding to the quantization index indicating whether the value is positive or negative; a state variable updated based on the quantization index is determined to indicate whether the value is odd or even; the value is derived based on the sign, the state variable, and a magnitude corresponding to the quantization index; a transform coefficient is determined to be the product of a quantization step size and the derived value; Video encoder.
12. The magnitudes corresponding to the quantization indexes include the quantization indexes increased by a factor of two. The video encoder of claim 11.
13. the state variable comprises a value of 0, 1, 2, or 3; The video encoder of claim 11.
14. the derived value is an even number when the state variable contains a value of 0 or 1, and an odd number when the state variable contains a value of 2 or 3. The video encoder of claim 11.
15. When the state variable is 0, if the parity corresponding to the quantization index is 0, the state variable is updated to 0, and if the parity is 1, the state variable is updated to 2; If the state variable is 1, the state variable is updated to 2 if the parity is 0, and updated to 0 if the parity is 1; If the state variable is 2, the state variable is updated to 1 if the parity is 0, and updated to 3 if the parity is 1; When the state variable is 3, if the parity is 0, the state variable is updated to 3, and if the parity is 1, the state variable is updated to 1. The video encoder of claim 11.
16. 1. A method for encoding one or more pictures into a video data stream, comprising: encoding quantization indexes into the video data stream; a sign is determined corresponding to the quantization index indicating whether the value is positive or negative; a state variable updated based on the quantization index is determined to indicate whether the value is odd or even; the value is derived based on the sign, the state variable, and a magnitude corresponding to the quantization index; a transform coefficient is determined to be the product of a quantization step size and the derived value; method.
17. the magnitudes corresponding to the quantization indexes include the quantization indexes increased by a factor of two; the state variable comprises a value of 0, 1, 2, or 3; 17. The method of claim 16.
18. the derived value is an even number when the state variable contains a value of 0 or 1, and an odd number when the state variable contains a value of 2 or 3.
17. The method of claim 16.
19. When the state variable is 0, if the parity corresponding to the quantization index is 0, the state variable is updated to 0, and if the parity is 1, the state variable is updated to 2; If the state variable is 1, the state variable is updated to 2 if the parity is 0, and updated to 0 if the parity is 1; If the state variable is 2, the state variable is updated to 1 if the parity is 0, and updated to 3 if the parity is 1; When the state variable is 3, if the parity is 0, the state variable is updated to 3, and if the parity is 1, the state variable is updated to 1.
17. The method of claim 16.
20. 17. A method for implementing a method of claim 16, comprising: storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 16; Non-transitory storage media.
Citation Information
Patent Citations
Methods and devices for data compression using adaptive reconstruction levels
US20120008680A1