Conversion coefficient block decoding device and method, and conversion coefficient block coding device and method
By adapting the scanning order and using individually selected context models for significant transform coefficients within large transform coefficient blocks, the encoding efficiency is improved, addressing the inefficiencies in conventional video encoding methods.
Patent Information
- Application Number
- JP2025029652
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2010-04-13
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2031-04-11
AI Technical Summary
Conventional video encoding methods face inefficiencies in encoding large transform coefficient blocks due to increased computational overhead and inaccurate probability estimation in context modeling, leading to suboptimal coding efficiency.
The proposed solution involves an encoding scheme that adapts the scanning order based on the positions of significant transform coefficients within the transform coefficient block, using context models individually selected for each syntax element based on the number of significant coefficients in the vicinity, and dividing large blocks into smaller sub-blocks for more efficient context modeling.
This approach enhances encoding efficiency by reducing the number of syntax elements required and improving the accuracy of probability estimation, particularly for large transform coefficient blocks, thereby achieving better coding performance.
Smart Images

Figure 2025081687000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the encoding of conversion coefficient blocks. Such encoding can be used, for example, in the encoding of images and videos.
Background Art
[0002] In conventional video encoding, usually, the images of a video sequence are divided into blocks. The blocks or the color components of the blocks are predicted either by motion compensation prediction or intra prediction. The blocks may be of different sizes and may be either square or rectangular. All samples of a block or the color components of a block are predicted using the same set of prediction parameters. Such prediction parameters include a reference index (identifying a reference image among a set of already encoded images), a motion parameter (specifying a measure for the movement of a block between a reference image and the current image), a parameter for specifying an interpolation filter, an intra prediction mode, etc. The motion parameter is represented by a displacement vector having horizontal and vertical components, or by a higher-order motion parameter such as an affine motion parameter consisting of six components. Also, it is possible to associate multiple sets of prediction parameters (such as a reference index and a motion parameter) with a single block. In this case, for each set of prediction parameters, a single intermediate prediction signal is generated for the block or the color components of the block, and the final prediction signal is constructed by the weighted sum of the intermediate prediction signals. The weighted parameters, and optionally a constant offset (added to the weighted sum), are fixed to an image, a reference image or a set of reference images, or they are included in the set of prediction parameters for the corresponding block. Similarly, still images may be divided into blocks, and the blocks are predicted by an intra prediction method (such as a spatial intra prediction method or a single intra prediction method for predicting the DC component of a block). Although it rarely occurs, the prediction signal may become zero.
[0003] The difference between the original block or the color component of the original block and the corresponding prediction signal is also referred to as the residual signal. Usually, a two-dimensional transform that is the subject of transformation and quantization is applied to the residual signal, and the resulting transform coefficients are quantized. In this transform coding, a block or the color component of a block for which a specific set of prediction parameters has been used may be further divided before the transform is applied. The transform block is equal to or smaller than the block used for prediction. It is also possible for the transform block to include two or more of the blocks used for prediction. Different transform blocks in a still image or one image of a video sequence can have different sizes, and the transform block can represent a square or rectangular block.
[0004] The resulting quantized transform coefficients, also referred to as transform coefficient levels, are transmitted using entropy coding techniques. Therefore, the blocks of transform coefficient levels are usually mapped, using a scan, to a vector (i.e., an ordered set) of scan-converted coefficient values. Note that different scans can be used for different blocks. A scan such as zigzag scan is often used. For blocks that contain only samples of one field of an interlaced frame (these blocks may be blocks within an encoded field or field blocks within an encoded frame), it is also common to use a different scan that is specially designed for field blocks. The entropy coding algorithm commonly used to code the resulting ordered sequence is run-level coding. Usually, the majority of the transform coefficient levels are zero, and a set of consecutive transform coefficient levels equal to zero is efficiently coded by coding the number of consecutive transform coefficient levels equal to zero (the run). It is represented in terms of rate. For the remaining (non-zero) conversion coefficients, the actual levels are encoded. There are various alternative means for run-level coding. The run before the non-zero coefficient and the level of the non-zero conversion coefficient are encoded together using a single symbol or codeword. Often, a special symbol for end-of-block is included, which is sent after the last non-zero conversion coefficient. It is also possible to first encode the number of non-zero conversion coefficients levels, and according to this number, the levels and runs are encoded.
[0005] Several different approaches are used in the highly efficient CABAC entropy coding in H.246. Here, the coding of the transform coefficient levels is divided into three steps. In the first step, a binary syntax element coded_block_flag indicating whether the transform block contains significant transform coefficient levels (i.e., transform coefficients that are non-zero) (this is referred to as "signaling") is sent for each transform block. If this syntax element indicates the presence of significant transform coefficient levels, the binarized significant map is coded to identify which transform coefficient levels are non-zero values. Then, in reverse scan order, the values of the non-zero transform coefficient levels are coded. The significant map is coded as follows. For each coefficient in scan order, a binary syntax element significant_coeff_flag identifying whether the corresponding transform coefficient level is not equal to zero is coded. If the bin of significant_coeff_flag is equal to 1, i.e., if a non-zero transform coefficient level exists at this scan position, a further binary syntax element last_significant_coeff_flag is coded. This bin indicates whether the current significant transform coefficient level is the last significant transform coefficient level in the block or whether further significant transform coefficient levels follow in scan order. If last_significant_coeff_flag indicates that no further significant transform coefficients follow, no further syntax elements for identifying the significant map for the block are coded. In the next step, the values of the significant transform coefficient levels whose positions in the block have already been determined by the significant map are coded. The values of the significant transform coefficient levels are coded in reverse scan order by using the following three syntax elements. The binary syntax element coeff_abs_greater_one indicates whether the absolute value of the significant transform coefficient level is greater than 1. If the binary syntax element coeff_abs_greater_one indicates that the absolute value is greater than 1, a further syntax element coeff_abs_level_minus_one identifying the absolute value of the transform coefficient level minus 1 is sent.Finally, the binary syntax element coeff_sign_flag that specifies the sign of the transform coefficient value is coded for each significant transform coefficient level. Again, although the syntax elements related to the significance map are coded in scan order, the syntax elements related to the actual values of the transform coefficient levels are transformed in reverse scan order to enable the use of a more appropriate context model.
[0006] In CABAC entropy coding in H.264, all syntax elements for the transform coefficient levels are coded using binary probability modeling. The non-binary syntax element coeff_abs_level_minus_one is first binary-coded, i.e., mapped to a sequence of binary decisions (bins), and these bins are coded sequentially. The binary syntax elements significant_coeff_flag, last_significant_coeff_flag, coeff_abs_greater_one, and coeff_sign_flag are coded directly. Each coded bin (including binary syntax elements) is associated with a context. A context represents a probability model for a class of coded bins. A measure related to the probability of one of two possible bin values is estimated for each context, based on the values of the bins that have already been coded, together with the corresponding context. For some bins related to transform coding, the context used for coding is selected based on the syntax elements that have already been transmitted, or based on the position within the block. selected.
[0007] The significance map identifies information regarding the significance (the transform coefficient level being different from zero) with respect to the scanning position. In the CABAC entropy coding of H.264, for a 4×4 block size, for encoding the binary syntax elements significant_coeff_flag and last_significant_coeff_flag, individual contexts are used for each scanning position, where different contexts are used for the significant_coeff_flag and last_significant_coeff_flag of one scanning position. For an 8×8 block size, the same context model is used for four consecutive scanning positions, which reduces to 16 context models for the significant_coeff_flag and an additional 16 context models for the last_significant_coeff_flag. This method of context-modeling the significant_coeff_flag and last_significant_coeff_flag has several disadvantages due to the large block size. On the other hand, if each scanning position is associated with an individual context model, when blocks larger than 8×8 are encoded, the number of context models increases significantly. Such an increase in the number of context models causes a delay in the adaptability of probability estimation and usually results in inaccuracy of probability estimation. These have an adverse effect on the coding efficiency. On the other hand, the assignment of context models to a number of consecutive scanning positions (as done for 8×8 blocks in H.264) is also not optimal for large block sizes. This is because non-zero transform coefficients usually concentrate in a specific region of the transform block (this region depends on the main structure within the corresponding block of the residual signal).
[0008] After encoding the significant map, the blocks are processed in reverse scan order. If the scan position is significant, i.e., the coefficient is different from zero, the binary syntax element coeff_abs_greater_one is sent. First, the second context model of the corresponding context model set is selected for the coeff_abs_greater_one syntax element. If the encoded value of any coeff_abs_greater_one syntax element within the block is equal to 1 (i.e., the absolute coefficient is greater than 2), the context modeling is switched back to the first context model of that set and this context model is used until the end of the block. Otherwise (all encoded values of coeff_abs_greater_one within the block are zero and the corresponding absolute coefficient level is equal to 1), the context model is selected depending on the number of coeff_abs_greater_one syntax elements equal to zero that have already been encoded or decoded in the reverse scan of the target block. The context model selection for the syntax element coeff_abs_greater_one can be outlined by the following equation. Here, the current context model index Ct+1 is selected based on the previous context model index Ct and the value of the previously encoded syntax element coeff_abs_greater_one represented by bint in the equation. For the first syntax element coeff_abs_greater_one within the block, the context model index is set equal to Ct = 1. [Number]
[0009] When the coeff_abs_greater_one syntax element for the same scan position is equal to 1, a second syntax element c for encoding the absolute transform coefficient level Only oeff_abs_level_minus_one is encoded. The non-binary syntax element coeff_abs_level_minus_one is binary-ized into a sequence of bins, and for the first bin of this binary-ization, a context model index is selected as described below. The remaining bins of the binary-ization are encoded with a fixed context. The context for the first bin of the binary-ization is selected as follows. For the first coeff_abs_level_minus_one syntax element, the first context model of the set of context models for the coeff_abs_level_minus_one syntax element is selected, and the corresponding context model index is set equal to Ct = 0. For each further first bin of the coeff_abs_level_minus_one syntax element, the context modeling is switched to the next context model in the set. Here, the number of context models in the set is limited to 5. The context model selection can be represented by the following formula. Here, the current context model index Ct+1 is selected based on the value of the previous context model index Ct. As described above, for the first syntax element coeff_abs_level_minus_one in the block, the context model index is set equal to Ct = 0. Note that different sets of context models are used for the syntax elements coeff_abs_greater_one and coeff_abs_level_minus_one. [Number]
[0010] For large blocks, this method has several disadvantages. Since the number of significant coefficients is larger than in the case of blocks with a small number of significant coefficients, the selection of the first context model for coeff_abs_greater_one (used when values of coeff_abs_greater_one equal to 1 are encoded for multiple blocks) is usually done overly early and reaches the last context model for coeff_abs_level_minus_one overly early. As a result, most bins of coeff_abs_greater_one and coeff_abs_level_minus_one are encoded with a single context model. However, these bins usually have different probabilities, and for this reason, the use of a single context model for a large number of bins has an adverse effect on the coding efficiency.
[0011] Generally, large blocks increase the computational overhead for performing spectral separation transforms, but better coding efficiency can be achieved when encoding sample arrays such as images or sample arrays representing other spatially sampled information signals such as depth maps if both small and large blocks can be effectively encoded. This is due to the dependency between the spatial resolution and the spectral resolution during the transform of the sample array within the block, and the larger the block, the higher the spectral resolution of the transform. Generally, it is preferable to apply a locally individual transform to the sample array such that the spectral components of the sample array do not vary significantly within the region of such an individual transform. For small blocks, it is guaranteed that the content within the block is relatively consistent. On the other hand, if the block is too small, the spectral resolution is low and the ratio of non-significant transform coefficients to significant transform coefficients is low.
[0012] Therefore, it is preferable to have a coding scheme that enables efficient coding of the transform coefficient block even if the transform coefficient block is large, and to be able to have a significance map of the transform coefficient block. SUMMARY OF THE INVENTION
Problems to be Solved by the Invention
[0013] Therefore, an object of the present invention is to provide an encoding scheme for encoding a transform coefficient block and a significance map indicating the positions of significant transform coefficients in the transform coefficient block, thereby enhancing the encoding efficiency.
Means for Solving the Problems
[0014] This problem is achieved by the subject matter recited in the independent claims.
[0015] According to a first aspect of the present application, the concept underlying the present application is that a high encoding efficiency for encoding a significance map indicating the positions of significant transform coefficients in a transform coefficient block is such that, for the relevant positions in the transform coefficient block, a sequentially extracted syntactic element indicating whether a significant or non-significant transform coefficient is arranged at each position, the scanning order sequentially associated with the position of the transform coefficient block among the positions of the transform coefficient block, can be achieved when it corresponds to the position of the significant transform coefficient indicated by the previously associated syntactic element. In particular, the inventors have found that in standard sample array content such as image, video or depth map content, significant transform coefficients mostly form a set on a particular side of the transform coefficient block corresponding to non-zero frequencies in the vertical direction or low frequencies in the horizontal direction, or vice versa, so that by considering the position of the significant transform coefficient indicated by the previously associated syntactic element, compared to a procedure in which the scanning order is predetermined independently of the position of the significant transform coefficient indicated by the syntactic element associated so far, it is possible to control a further factor of the scanning so that the probability of reaching the last significant transform coefficient in the transform coefficient block earlier is increased. This applies to small blocks as well, but particularly to larger blocks.
[0016] In one embodiment of the present application, the entropy decoder is configured to extract from the data stream information that enables recognition of whether the significant transform coefficient currently indicated by the currently associated syntax element is the last significant transform coefficient independent of the exact position within the transform coefficient block, where the entropy decoder is configured not to expect a further syntax element if the current syntax element is related to such a last significant transform coefficient. This information can include the number of significant transform coefficients within the block. Alternatively, a second syntax element is interleaved with the first syntax element, where the second syntax element indicates whether the last transform coefficient in the transform coefficient block is also the same for the associated position where the significant transform coefficients are arranged.
[0017] In one embodiment, the associator, at a predetermined position within the transform coefficient block, simply applies a scan order according to the positions of the significant transform coefficients shown so far. For example, several sub-paths traversing mutually separated subsets of positions within the transform coefficient block extend substantially diagonally from a pair of sides of the transform coefficient block corresponding to the minimum frequency along the first direction and the maximum frequency along the other direction, to the opposite pair of sides of the transform coefficient block corresponding to the zero frequency along the second direction and the maximum frequency along the first direction. In this case, the associator selects the scan order such that the sub-paths are traversed in an order where the distance from the sub-path to the DC position of the sub-path within the transform coefficient block increases monotonically between sub-paths, each sub-path is traversed without interruption along the direction of travel, and for each sub-path, the direction in which the sub-path is traversed is selected by the associator according to the positions of the significant transform coefficients traversed in the previous sub-path. By this measure, the probability increases that the last sub-path where the last significant transform coefficient is arranged is traversed in a direction such that the probability is high that the last significant transform coefficient is in the first half rather than the second half of this last sub-path, thereby reducing the number of syntax elements indicating whether significant or non-significant transform coefficients are arranged at their respective positions. This can be done. This effect is particularly meaningful in the case of a large transform coefficient block.
[0018] According to a further aspect of the present application, the present application uses a context individually selected for each syntax element according to a number of significant transform coefficients in the vicinity of each syntax element shown to be significant by any of the preceding syntax elements, and indicates whether a significant or non-significant transform coefficient is arranged at each position. When the above-described syntax element is context-adaptively entropy decoded for an associated position within the transform coefficient block, it is based on the insight that a significance map indicating the positions of the significant transform coefficients within the transform coefficient block can be encoded more efficiently. In particular, the inventors have found that for a transform coefficient block of increasing size, the significant transform coefficients are somewhat aggregated in a predetermined area within the transform coefficient block, whereby not only is it sensitive to the number of significant transform coefficients traversed in a predetermined scan order up to that point, but also context adaptation that takes into account the vicinity of the significant transform coefficients results in a better fit of the context, and thus increases the coding efficiency of entropy coding.
[0019] Of course, both aspects outlined above can be combined in a desirable manner.
[0020] Also, according to a further aspect of the present application, the present application is such that a significance map indicating the positions of significant transform coefficients within a transform coefficient block precedes the encoding of the actual values of the significant transform coefficients within the transform coefficient block, and a predetermined scanning order between the positions of the transform coefficients used to sequentially associate the sequence of values of the significant transform coefficients with the positions of the significant transform coefficients, when scanning the positions of the transform coefficients within sub-blocks in the coefficient scanning order and scanning the transform coefficient blocks in sub-blocks using the sub-block scanning order between sub-blocks, and a set selected from a plurality of sets consisting of a number of contexts is selected depending on the values of the significant transform coefficient values, the values of the transform coefficients within the sub-blocks of the transform coefficient blocks already passed in the sub-block scanning order, or the values of the transform coefficients of the sub-blocks located together in the previously decoded transform coefficient blocks, and is used for sequentially context-adaptive entropy decoding, based on the insight that the encoding efficiency for the encoding of the transform coefficient blocks can be improved. In this way, context adaptation becomes very suitable for the above-described characteristics of the significant transform coefficients concentrated in a predetermined area within the transform coefficient block. In other words, the values are scanned within the sub-blocks in a context selected based on sub-block statistics.
[0021] Again, the last aspect can also be combined with any one or both of the aspects identified previously in the present application.
[0022] Preferred embodiments of the present application will be described below with reference to the drawings.
Brief Description of the Drawings
[0023]
Figure 1
Figure 2A
Figure 2B
Figure 2C
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
[0024] In the description of the figures, elements that appear in some of the figures are denoted by the same reference numerals in each of those figures, and in order to avoid unnecessary repetition, repeated descriptions of those elements are avoided as long as their functions are of concern. However, the functions and descriptions given for one drawing shall apply to other drawings as well unless otherwise explicitly stated.
[0025] Figure 1 shows an example of an encoder 10 in which aspects of the present application are implemented. The encoder encodes a series of information samples 20 into a data stream. The series of information samples means any kind of information signal sampled spatially. For example, the sample array 20 can be an image constituting a still image or video. In this case, the information samples are brightness values, color values, luma values, chroma values, etc. The information samples may also be depth values when the sample array 20 is a depth map generated by, for example, an emission time sensor.
[0026] The encoder 10 is a block-based encoder. That is, the encoder 10 encodes the sample array 20 into the data stream 30 in units of blocks 40. Encoding in units of blocks 40 does not necessarily mean that the encoder 10 encodes these blocks 40 completely independently of each other. On the contrary, the encoder 10 may use a reconstructed previously encoded block to extrapolate or intra-predict the remaining blocks, or may use the granularity of the blocks to set encoding parameters, that is, to set the method by which each sample array region corresponding to each block is encoded.
[0027] Also, the encoder 10 is a transform coder. That is, the encoder 10 encodes the block 40 by using a transform for moving the information samples within each block 40 from the spatial domain to the spectral domain. A two-dimensional transform such as the DCT of the FFT is used. Preferably, the block 40 is square or rectangular.
[0028] The subdivision of the specimen array 20 shown in FIG. 1 into blocks 40 serves only for the purpose of explanation. FIG. 1 shows a specimen array 20 subdivided into a regular two-dimensional arrangement of adjacent square or rectangular blocks 40 that do not overlap each other. The size of the blocks 40 can be predetermined. In that case, the encoder 10 does not have to transfer information about the block size of the blocks 40 in the data stream 30 to the decoding side. For example, the decoder can anticipate that predetermined block size.
[0029] On the other hand, several alternatives are possible. For example, the blocks may overlap each other. However, the overlap is limited such that each block has a portion that does not overlap any adjacent block, or such that each specimen of a block overlaps at most one block among the adjacent blocks arranged juxtaposed to the current block along a predetermined direction. In the latter case, the left and right adjacent blocks can overlap so as to completely cover the current block, but this means that the adjacent blocks themselves do not overlap each other. The same applies to adjacency in the vertical and diagonal directions.
[0030] As a further alternative, the subdivision of the specimen array 20 into blocks 40 can be applied by the encoder 10 to the content of the specimen array 20, and the subdivision information used for the subdivision can be transferred to the decoder side via the bit stream 30.
[0031] Figures 2A through 2C show different examples of the subdivision of the specimen array 20 into blocks 40. FIG. 2A shows a quadtree-based subdivision of the specimen array 20 into blocks 40 of different sizes, where representative blocks are shown as 40a, 40b, 40c, and 40d in order of increasing size. According to the subdivision in FIG. 2A, the specimen array 20 is first divided into a regular two-dimensional arrangement of tree blocks 40d. These tree blocks 40d are associated with individual subdivision information regarding whether a certain tree block 40d is further subdivided according to the quadtree structure. The tree block on the left side of block 40d is, by way of example, subdivided into smaller blocks according to the quadtree structure. The encoder 10 can perform one two-dimensional transformation for each of the blocks shown by solid and dashed lines in FIG. 2A. In other words, the encoder 10 can transform the array 20 in units of block subdivision.
[0032] Instead of the quadtree-based subdivision, a more general multi-tree-based subdivision may be used, and the number of child nodes per hierarchical level may vary between different hierarchical levels.
[0033] FIG. 2B shows another example of the subdivision. According to FIG. 2B, the specimen array 20 is first divided into macroblocks 40b in a regular two-dimensional arrangement that are adjacent to each other without overlapping. Here, subdivision information is associated with each microblock 40b, and according to this subdivision information, the microblock is either not subdivided or, if it is subdivided, is subdivided into sub-blocks of the same size in a regular two-dimensional arrangement so that different subdivision granularities are achieved for different microblocks. The result is the subdivision of the specimen array 20 into blocks 40 of different sizes, and representative examples of different sizes are shown as 40a, 40b, and 40a'. As shown in FIG. 2A, the encoder 10 performs a two-dimensional transformation on each of the blocks shown by solid and dashed lines in FIG. 2B. FIG. 2C will be described later.
[0034] Figure 3 shows a decoder 50 that can decode the data stream 30 generated by the encoder 10 and reconstruct a reconstructed version 60 of the sample array 20. The decoder 50 reconstructs the reconstructed version 60 by extracting the transform coefficient block for each of the blocks 40 from the data stream 30 and performing an inverse transform on each of the transform coefficient blocks.
[0035] The encoder 10 and the decoder 50 are each configured to perform entropy encoding / decoding to insert information about the transform coefficient blocks and extract this information from the data stream. Details regarding this will be described later. Note that the data stream 30 does not necessarily have to include information about the transform coefficient blocks for all of the blocks 40 of the sample array 20. Conversely, a subset of the blocks 40 may be encoded in the bitstream 30 by other means. For example, the encoder 10 may determine to stop inserting transform coefficient blocks for certain ones of the blocks 40 and instead insert alternative encoding parameters into the bitstream 30 so that the decoder 50 can predict it, or otherwise fill each block in the reconstructed version 60. For example, the encoder 10 may perform texture analysis for placing blocks in the sample array 20 so that texture synthesis by the decoder on the decoder side fills the sample array 20, and appropriately indicate this in the bitstream. It may be performed and shown as appropriate in the bitstream.
[0036] As will be described in the following drawings, the conversion coefficient block does not necessarily represent the spectral region representation of the original information sample of each block 40 of the sample array 20. Conversely, such a conversion coefficient block may represent the spectral region representation of the prediction residue of each block 40. FIG. 4 shows an embodiment of such an encoder. The encoder of FIG. 4 includes a conversion stage 100, an entropy coder 102, an inverse conversion stage 104, a predictor 106, a subtractor 108, and an adder 110. The subtractor 108, the conversion stage 100, and the entropy coder 102 are connected in series between the input 112 and the output 114 of the encoder of FIG. 4 in this order. The inverse conversion stage 104, the adder 110, and the predictor 106 are connected between the output of the conversion stage 100 and the inverted input of the subtractor 108 in this order, and the output of the predictor 106 is also connected to the other input of the adder 110.
[0037] The coder of FIG. 4 is a prediction-transform-based block coder. That is, the blocks of the sample array 20 entering the input 112 are predicted from the previously encoded and reconstructed portions of the same sample array 20, or from other sample arrays that have been previously encoded and reconstructed and that precede or follow the current sample array 20 in time. The prediction is performed by the predictor 106. The subtractor 108 subtracts the predicted value from such an original block, and the transform stage 100 performs a two-dimensional transform on the prediction residue. The two-dimensional transform itself within the transform stage 100 or subsequent processing leads to quantization of the transform coefficients within the transform coefficient block. The quantized transform coefficient block is losslessly encoded by entropy encoding within the entropy encoder 102 such that, for example, the resulting data stream is output at the output 114. The inverse transform stage 104 reconstructs the quantized residue, and subsequently the adder 100 synthesizes the reconstructed residue with the corresponding prediction in order to obtain the reconstructed information samples on which the predictor 106 bases its prediction of the current encoded prediction block as described above. The predictor 106 can use different prediction modes such as the intra prediction mode and the inter prediction mode to predict the block, and the prediction parameters are transferred to the entropy encoder 102 for insertion into the data stream.
[0038] That is, according to the embodiment of FIG. 4, the transform coefficient block represents the spectral representation of the residue of the sample array rather than the actual information samples of the sample array.
[0039] Note that there are several alternatives in the embodiment of FIG. 4, some of which are described in the introduction of the specification, and that description is incorporated here into the description of FIG. 4. For example, the prediction generated by the predictor 106 may not be entropy encoded. Conversely, side information may be transferred to the decoding side via other coding schemes.
[0040] FIG. 5 shows a decoder capable of decoding the data stream generated by the encoder of FIG. 4. The decoder of FIG. 5 includes an entropy decoder 150, an inverse transform stage 152, an adder 154, and a predictor 156. The entropy decoder 150, the inverse transform stage 152, and the adder 154 are connected in series between the input 158 and the output 160 of the decoder of FIG. 5 in this order. Another output of the entropy decoder 150 is connected to the predictor 156, and then the predictor 156 is connected between the output of the adder 154 and the other input. The entropy decoder 150 extracts the transform coefficient block from the data stream entering the decoder of FIG. 5 at the input 158, where an inverse transform is applied to the transform coefficient block in stage 152 to obtain the residual signal. The residual signal is combined in the adder 154 with the prediction from the predictor 156 to obtain the reconstruction block of the reconstructed version of the sample array at the output 160. The predictor 156 generates a prediction based on the reconstructed version, thereby reconstructing the prediction performed by the predictor 106 on the encoder side. To obtain the same prediction value as used on the encoder side the predictor 156 uses prediction parameters, which are obtained by the entropy decoder 150 from the data stream of the input 158.
[0041] In the above-described embodiments, the spatial granularities in which the remaining prediction and transformation are performed may not be equal to each other. This is shown in FIG. 2C. In the figure, the subdivision of the prediction block of the prediction granularity is shown by a solid line, and the residual granularity is shown by a broken line. As can be seen from the figure, the subdivisions are selected by mutually independent encoders. More precisely, the syntax of the data stream enables the definition of a residual subdivision independent of the prediction subdivision. Alternatively, the residual subdivision may be an extension of the prediction subdivision, and each residual block may be equal to or an appropriate subset of the prediction block. This is shown, for example, in FIGS. 2A and 2B. Here too, the prediction granularity is shown by a solid line and the residual granularity is shown by a broken line. Regarding this, in FIGS. 2A-2C, the large solid block including the broken-line block 40a becomes, for example, a prediction block in which the prediction parameter setting is individually executed, while all the blocks having the associated reference signs become residual blocks in which one two-dimensional transformation is executed.
[0042] The above-described embodiments are common in that the block of the (remaining or original) specimen is converted into a conversion coefficient block on the encoder side, and the conversion coefficient block is inverse-transformed into a specimen reconstruction block on the decoder side. This is shown in FIG. 6. FIG. 6 shows the block of specimen 200. In the case of FIG. 6, this block 200 is, by way of example, a specimen 202 that is two-dimensional and has a size of 4×4. The specimens 202 are regularly arranged along the horizontal direction x and the vertical direction y. By the above-described two-dimensional transformation T, the block 200 is transformed into a block 204 in the spectral domain, that is, a block of conversion coefficients 206. Here, the conversion block 204 has the same size as the block 200. That is, the conversion block 204 has the same number of conversion coefficients 206 as the number of specimens that the block 200 has in both the horizontal and vertical directions. However, since the transformation T is a spectral transformation, the positions of the conversion coefficients 206 in the conversion block 204 correspond to spectral components rather than the spatial positions of the contents of the block 200. In particular, the horizontal axis of the conversion block 204 corresponds to an axis along which the horizontal spectral frequency monotonically increases, the vertical axis corresponds to an axis along which the vertical spatial frequency monotonically increases, and the DC component conversion coefficient is located at the corner of the block 204, here, by way of example, the upper left corner, such that the conversion coefficient 206 corresponding to the highest frequency in both the horizontal and vertical directions is located at the lower right corner. Ignoring the spatial direction, the spatial frequency to which a given conversion coefficient 206 belongs generally increases from the upper left corner to the lower right corner. Inverse transformation T -1 causes the conversion block 204 to be transferred back from the spectral domain to the spatial domain so as to reacquire a replica 208 of the block 200. If no quantization / loss is introduced during the transformation, the reconstruction will be complete.
[0043] As already described above, as the block size of block 200 increases, it can be seen from FIG. 6 that the spectral resolution of the resulting spectral display 204 increases. On the other hand, the quantization noise tends to spread across the entire block 208. For this reason, abrupt and highly localized objects within block 200 tend to introduce deviations in the reconverted block as compared to the original block 200 due to the quantization noise. On the other hand, the main advantageous effect of using a larger block is that, within a larger block compared to a smaller block, the ratio of the number of significant, i.e., non-zero (quantized) transform coefficients to the number of insignificant transform coefficients is reduced, thereby enabling better coding efficiency. In other words, often, the significant transform coefficients, i.e., the transform coefficients that are not quantized to zero, are sparsely distributed across the transform block 204. Due to this, according to the embodiments described in more detail hereinafter, the positions of the significant transform coefficients are signaled into the data stream by means of a significance map. Separately from this, when the transform coefficients are quantized, the values of the significant transform coefficients, i.e., the transform coefficient levels, are transmitted within the data stream.
[0044] Therefore, according to the embodiments of the present application, an apparatus for decoding such a significance map from a data stream or for decoding a significance map along corresponding significant transform coefficient values from a data stream is implemented as shown in FIG. 7, and each of the entropy decoders described above, i.e., decoder 50 and entropy decoder 150, constitutes the apparatus shown in FIG. 7.
[0045] The apparatus of FIG. 7 includes a map / coefficient entropy decoder 250 and an associator 252. The map / coefficient entropy decoder 250 is connected to an input 254, and a syntax element representing a significant map and a significant transform coefficient value is input to this input 254. As described in more detail below, there are different probabilities regarding the order in which the syntax elements describing the significant map and the significant transform coefficient value are input to the map / coefficient entropy decoder 250. The significant map syntax element may precede the corresponding level, and both may be interleaved. However, it is assumed that the syntax element representing the significant map precedes the value (level) of the significant transform coefficient, and the map / coefficient entropy decoder 250 first decodes the significant map and then the transform coefficient levels of the significant transform coefficient.
[0046] As the map / coefficient entropy decoder 250 sequentially decodes the syntax elements representing the significant map and the significant transform coefficient value, the associator 252 is configured to associate these sequentially decoded syntax elements / values with positions within the transform block 256. The scanning order in which the associator 252 associates the sequentially decoded syntax elements representing the significant map and the levels of the significant transform coefficient with each position of the transform block 256 follows the same one-dimensional scanning order as that used on the encoding side to introduce these elements into the data stream among each position of the transform block 256. Also, as outlined in more detail below, the scanning order for the syntax elements of the significant map may or may not be equal to the order used for the significant transform values.
[0047] The map / coefficient entropy decoder 250 can utilize the information about the available transform blocks 256 at that time generated by the associator 252 up to the syntax element / level to be currently decoded, in order to set a probability estimation context for entropy decoding the syntax element / level to be currently decoded, as indicated by the dashed line 258. For example, the associator 252 stores (logs) the information collected so far from the sequentially associated syntax elements such as the level itself, or information about whether significant transform coefficients are arranged at each position, or whether nothing is known about each position of the transform block 256, and the map / coefficient entropy decoder 250 may be allowed to access this memory. Although the memory is not shown in FIG. 7, since there is a memory or log buffer for storing the prior information obtained so far by the associator 252 and the entropy decoder 250, the reference numeral 256 can be considered to also indicate this memory. Therefore, FIG. 7 shows, by an 'x' mark, the significant transform coefficients obtained from the previously decoded syntax elements representing the significant map, and '1' indicates that the significant transform coefficient level of the significant transform coefficient at each position has already been decoded and that it is 1. When the significant map syntax element precedes a significant value in the data stream, an 'x' mark (significant transform coefficient) is recorded at the position of '1' in the memory 256 before decoding each value and recording '1' (this situation would represent the entire significant map).
[0048] The following description focuses on specific embodiments for encoding transform coefficient blocks or significant maps, but those embodiments can be easily migrated to the above-described embodiments. In these embodiments, the binary syntax element coded_block_flag is transmitted to each transform block, and that transform block has a significant transform coefficient level (i.e., non-zero Indicates whether it includes the conversion coefficient (which is ロ). When this syntax element indicates the existence of a significant conversion coefficient level, the significant map is encoded. That is, encoding is performed only when there is a significant conversion coefficient level. The significant map identifies which conversion coefficient levels have non-zero values, as described above. Significant map encoding includes encoding of the binary syntax element significant_coeff_flag, and each binary syntax element significant_coeff_flag identifies whether the corresponding conversion coefficient level is not equal to zero for the associated coefficient position. Encoding is performed in a certain scan order, and this scan order can change depending on the positions of the significant coefficients that have been identified as significant up to that point during significant map encoding. This will be explained in more detail below. Furthermore, significant map encoding includes encoding of the binary syntax element last_significant_coeff_flag. This binary syntax element is distributed at that position together with the sequence of significant_coeff_flag, and significant_coeff_flag signals significant coefficients. When the significant_coeff_flag bin is equal to 1, that is, when there is a non-zero conversion coefficient level at this scan position, an additional binary syntax element last_significant_coeff_flag is encoded. This bin indicates whether the current significant conversion coefficient level is the last significant conversion coefficient level within the block or whether there are additional significant conversion coefficient levels following in the scan order. When last_significant_coeff_flag indicates that there are no further significant conversion coefficients, no further syntax elements are encoded to identify the significant map for that block. Alternatively, the number of significant coefficient positions may be signaled within the data stream before the encoding of the sequence of significant_coeff_flag. In the next step, the values of the significant conversion coefficient levels are encoded. As described above, alternatively, the transmission of the levels may be interleaved with the transmission of the significant map. The values of the significant conversion coefficient levels are encoded in a further scan order exemplified below.The following three syntax elements are used. The binary syntax element coeff_abs_greater_one indicates whether the absolute value of the significant transform coefficient level is greater than 1. When the binary syntax element coeff_abs_greater_one indicates that the absolute value is greater than 1, a further syntax element coeff_abs_level_minus_one that specifies the absolute value of the transform coefficient level minus 1 is sent. Finally, the binary syntax element coeff_sign_flag that specifies the sign of the transform coefficient value is coded for each significant transform coefficient level.
[0049] The embodiments described below further reduce the bit rate and thereby enable higher coding efficiency. To do so, these embodiments use a specific approach to context modeling for syntax elements related to transform coefficients. In particular, a new context model selection for the syntax elements significant_coeff_flag, last_significant_coeff_flag, coeff_abs_greater_one, and coeff_abs_level_minus_one is used. Furthermore, an adaptive switching of the scan during the encoding / decoding of the significance map (which specifies the positions of the non-zero transform coefficient levels) is described. Regarding the meaning of the syntax elements to be mentioned, refer to the above introduction section of this application.
[0050] The coding of the syntax elements significant_coeff_flag and last_significant_coeff_flag that specify the significance map is improved by a new context modeling based on adaptive scanning and a limited neighborhood of the already coded scan positions. These new concepts result in a more efficient coding of the significance map (i.e., a reduction of the corresponding bit rate), especially for large block sizes.
[0051] One aspect of the embodiments outlined below is that the scan order (i.e., the mapping of blocks of transform coefficient values to an ordered set (vector) of transform coefficient values) is adapted during encoding / decoding of the significance map based on the values of syntax elements that have already been encoded / decoded for the significance map.
[0052] In a preferred embodiment, the scan order is adaptively switched between two or more predetermined scan patterns. In a preferred embodiment, the switching is performed only at a certain predetermined scan position. In a more preferred embodiment of the present invention, the scan order is adaptively switched between two predetermined scan patterns. In a preferred embodiment, the switching between the two predetermined scan patterns is performed only at a certain predetermined scan position.
[0053] The advantage of switching the scan pattern is a reduction in bitrate, which results from a decrease in the number of encoded syntax elements. As an intuitive example, referring to FIG. 6, for large transform blocks in particular, it is common for the particularly significant transform coefficient values to concentrate on one of the block boundaries 270, 272. The reason is that the residual blocks mainly contain horizontal or vertical structures. In the most commonly used zigzag scan 274, there is a probability of about 0.5 that the last diagonal sub-scan of the zigzag scan encountering the last significant coefficient starts from the side where the significant coefficients do not concentrate. In that case, a large number of syntax elements for transform coefficient levels equal to zero have to be encoded before reaching the last non-zero transform coefficient value. This can be avoided if the diagonal sub-scan always starts on the side where the significant transform coefficient levels concentrate.
[0054] The preferred embodiments of the present invention will be described in more detail below.
[0055] As described above, in order to enable early adaptation of the context model and achieve high coding efficiency even for large block sizes, it is preferable to keep the number of context models moderately small. Therefore, a particular context should be used for two or more scanning positions. However, since significant transform coefficient levels usually concentrate in a certain area of the transform block (this concentration is usually the result of a certain dominant structure existing in, for example, the residual block), the concept of assigning the same context as that performed for 8×8 blocks in H.264 to a number of consecutive scanning positions is usually not appropriate. In designing context selection, the above consideration that significant transform coefficient levels often concentrate in a predetermined area of the transform block can be used. In the following, a concept in which this consideration can be utilized will be described.
[0056] In a desirable embodiment, a large transform block (for example, larger than 8×8) is divided into a number of rectangular sub-blocks (for example, 16 sub-blocks), and each of these sub-blocks is associated with an individual context model for encoding significant_coeff_flag and last_significant_coeff_flag (here, different context models are used for significant_coeff_flag and last_significant_coeff_flag). The division into sub-blocks may be different for significant_coeff_flag and last_significant_coeff_flag. The same context model may be used for all scanning positions located in a particular sub-block.
[0057] In a further preferred embodiment, large transform blocks (e.g., larger than 8×8) are partitioned into a number of rectangular and / or non-rectangular sub-regions, each of these sub-regions being associated with an individual context model for encoding the significant_coeff_flag and / or the last_significant_coeff_flag. The partitioning into sub-regions may be different for the significant_coe ff_flag and the last_significant_coeff_flag. The same context model is used for all scan positions located within a particular sub-region.
[0058] In a further preferred embodiment, the context model for encoding the significant_coeff_flag and / or the last_significant_coeff_flag is selected based on already encoded symbols in a predetermined spatial neighborhood of the current scan position. The predetermined neighborhood may be different for different scan positions. In a preferred embodiment, the context model is selected based on the number of significant transform coefficient levels in a predetermined spatial neighborhood of the current scan position, where only already encoded significant representations are counted.
[0059] Further details of preferred embodiments of the present invention are described below.
[0060] As described above, for large block sizes, conventional context modeling encodes a large number of bins (usually with different probabilities) with a single context model for the syntactic elements of coeff_abs_greater_one and coeff_abs_level_minus_one. To avoid this drawback for large block sizes, according to an embodiment, a large block is divided into smaller square or rectangular sub-blocks of a specific size, and individual context modeling is applied to each sub-block. Further, multiple sets of context models may be used, where one of these sets of context models is selected for each sub-block based on an analysis of the statistics of previously encoded sub-blocks. In a preferred embodiment of the present invention, the number of transform coefficients greater than 2 (i.e., coeff_abs_level_minus_1>1) in the previously encoded sub-blocks of the same block is used to derive the set of context models for the current sub-block. These extensions to the context modeling of the syntactic elements of coeff_abs_greater_one and coeff_abs_level_minus_one result in more efficient encoding of both syntactic elements, especially for large block sizes. In a preferred embodiment, the block size of the sub-block is 2×2. In other preferred embodiments, the block size of the sub-block is 4×4.
[0061] In a first step, blocks larger than a predetermined size are divided into smaller sub-blocks of a specific size. The encoding process of the absolute transform coefficient levels maps sub-blocks, which are square or rectangular blocks, into an ordered set (vector) of sub-blocks using a scan. Here, different scans can be used for different blocks. In a preferred embodiment, the sub-blocks are processed using a zigzag scan, and the transform coefficient levels within the sub-blocks are processed by a reverse zigzag scan, i.e., a scan reading from the transform coefficients belonging to the highest frequencies in both vertical and horizontal directions to the coefficients related to the lowest frequencies in both directions. In another preferred embodiment of the present invention, a reverse zigzag scan is used for encoding the sub-blocks and for encoding the transform coefficient levels within the sub-blocks. In another preferred embodiment of the present invention, the same adaptive scan (referenced above) used for encoding the significance map is used for processing the entire block of transform coefficient levels.
[0062] The division of the large transform block into sub-blocks avoids the problem of using only one context model for most of the bins of the large transform block. Within the sub-blocks, the latest context modeling (defined in H.264) or fixed contexts are used depending on the actual size of the sub-blocks. Further, the statistics (in terms of probability modeling) for such sub-blocks are for transform blocks of the same size This is different from the lock statistics. This feature can be utilized for the syntax elements of coeff_abs_greater_one and coeff_abs_level_minus_one by extending the set of context models. Prepare multiple sets of context models, and for each sub-block, one of these sets of context models may be selected based on the statistics of the previously encoded sub-blocks in the current transform block or the previously encoded transform blocks. In a preferred embodiment of the present invention, the selected set of context models is derived based on the statistics of the previously encoded sub-blocks in the same block. In another preferred embodiment of the present invention, the selected set of context models is derived based on the statistics of the same sub-blocks of the previously encoded blocks. In a preferred embodiment, the number of sets of context models is set to be equal to 4, and in another preferred embodiment, the number of sets of context models is set to be equal to 16. In a preferred embodiment, the statistics used to obtain the set of context models is the number of absolute transform coefficient levels greater than 2 in the previously encoded sub-blocks. In another preferred embodiment, the statistics used to obtain the set of context models is the difference between the number of significant coefficients and the number of transform coefficient levels with absolute values greater than 2.
[0063] The encoding of the significance map is performed, as outlined below, i.e., by adaptively switching the scan order.
[0064] In a preferred embodiment, the scanning order for encoding the significance map is adapted by switching between two predetermined scanning patterns. The switching between scanning patterns is only performed at a certain predetermined scanning position. The determination of whether the scanning pattern is switched depends on the value of the significance map syntax element that has already been encoded / decoded. In a preferred embodiment, both of the predetermined scanning patterns are scanning patterns with diagonal sub-scans similar to the scanning pattern of zigzag scanning. The scanning patterns are shown in FIG. 8. Both of the scanning patterns 300 and 302 consist of a number of diagonal sub-scans for the diagonal line from the leftmost bottom to the rightmost top, or vice versa. The scanning in the diagonal sub-scan (not shown) is performed from the leftmost top to the rightmost bottom for both of the predetermined scanning patterns. However, the scanning within the diagonal sub-scan is different (as shown in the figure). For the first scanning pattern 300, the diagonal sub-scan is scanned from the leftmost bottom to the rightmost top (left figure in FIG. 8), and for the second scanning pattern 302, the diagonal sub-scan is scanned from the rightmost top to the leftmost bottom (right figure in FIG. 8). In one embodiment, the encoding of the significance map starts with the second scanning pattern. During encoding / decoding of the syntax element, the number of significant transform coefficient values is counted by two counters c 1 and c 2 . The first counter c 1 is , counting the number of significant transform coefficients located in the leftmost bottom part of the transform block. That is, this first counter c 1 is incremented by 1 when the significant transform coefficient levels where the horizontal coordinate x in the transform block is smaller than the vertical coordinate y are encoded / decoded. The second counter c 2 is , counting the number of significant transform coefficients located in the rightmost top part of the transform block. That is, this second counter c 2 is incremented by 1 when the significant transform coefficient levels where the horizontal coordinate x in the transform block is larger than the vertical coordinate y are encoded / decoded. The adaptation of the counters can be performed by the associator 252 in FIG. 7 and can be explained by the following formula. Here, t represents the scanning position index, and both counters are initialized to zero.
Number
Number
[0065] At the end of each diagonal sub-scan, it is determined by the associator 252 which of the first and second predetermined scan patterns 300, 302 is used for the next diagonal sub-scan. This determination is based on the values of the counters c 1 and c 2 . If the count value for the left bottom part of the conversion block is greater than the count value of the left bottom part, the scan pattern for performing the diagonal sub-scan from the left bottom to the upper right is used; otherwise (when the count value for the left bottom part of the conversion block is less than or equal to the count value of the left bottom part), the scan pattern for performing the diagonal sub-scan from the upper right to the left bottom is used. This determination is represented by the following formula.
Number
[0066] The above-described embodiments of the present invention can be easily applied to other scan patterns. As an example, the scan pattern used for field macroblocks in H.264 is decomposed into sub-scans. In a more desirable embodiment, any given scan pattern is divided into sub-scans. For each sub-scan, two scan patterns are defined as those from the left bottom to the upper right (as the basic scan direction) and from the upper right to the left bottom. Further, two counters are introduced to count the number of significant coefficients in the first part (close to the left bottom boundary of the conversion block) and the second part (close to the upper right boundary of the conversion block) within the sub-scan. Finally, at the end of each sub-scan (based on the value of the counter), it is determined whether the next sub-scan is scanned from the left bottom to the upper right or from the upper right to the left bottom.
[0067] Next, an embodiment of how the entropy decoder 250 models the context will be described.
[0068] In one desirable embodiment, context modeling for the significant_coeff_flag is performed as follows. For a 4×4 block, context modeling is performed as defined in H.264. For an 8×8 block, the transform block is separated into 16 2×2 sample sub-blocks, and each of these sub-blocks is associated with an individual context. Note that this concept can be extended to larger block sizes, different numbers of sub-blocks, and non-rectangular sub-regions as described above.
[0069] In a further desirable embodiment, the context model selection for large transform blocks (e.g., for blocks larger than 8×8) is based on the number of already encoded significant transform coefficients in a predetermined neighborhood (within the transform block). An example of the definition of the neighborhood corresponding to a desirable embodiment of the present invention is shown in FIG. 9. Those marked with × surrounded by ○ are available neighborhoods that are always considered for evaluation, and those marked with × and △ are neighborhoods that are evaluated according to the current scan position and the current scan direction): · When the current scan position is within the 2×2 left corner 304, an individual context model is used for each scan position (FIG. 9, left figure), · When the current scan position is not within the 2×2 left corner and is not located in the first row or the first column of the transform block, the neighborhood shown on the right side of FIG. 9 is used to evaluate the number of significant transform coefficients in the neighborhood of the current scan position "x" where there is nothing around it, · When the current scan position "x" where there is nothing around it is in the first row of the transform block, the neighborhood specified in the right figure of FIG. 10 is used, · When the current scan position "x" is in the first column of the block, the neighborhood specified in the left figure of FIG. 10 is used.
[0070] In other words, the decoder 250 adaptively entropy decodes each significant map syntax element in a context-adaptive manner by using a context that is individually selected according to the positions where the significant transform coefficients are arranged according to the previously extracted and associated significant map syntax elements, and the positions (either the "x" on the right side of FIG. 9 and both sides of FIG. 10, or the marked positions on the left side of FIG. 9) adjacent to the positions where each current significant map syntax element is associated, and is configured to sequentially extract the significant map syntax elements. As shown, the neighborhood of the position associated with each current syntax element consists of positions that are at most in the vertical and / or horizontal directions and that are either directly adjacent positions or positions separated from the positions where each significant map syntax element is associated. Alternatively, only the positions directly adjacent to each current syntax element are considered. As a result, the size of the transform coefficient block is a position of 8×8 or more.
[0071] In a preferred embodiment, the context model used to encode a particular significant_coeff_flag is selected according to the number of significant transform coefficient levels that have already been encoded in the defined neighborhood. Here, the number of available context models can be made smaller than the possible values of the number of significant transform coefficient levels in the defined neighborhood. The encoder and decoder can include a table (or different mapping mechanism) for mapping the number of significant transform coefficient levels in the defined neighborhood to the index of the context model.
[0072] In a further preferred embodiment, the index of the selected context model depends on the number of significant transform coefficient levels in the defined neighborhood and / or one or more additional parameters such as the type of neighborhood used or the scanned position or the quantized value of the scanned position.
[0073] For the coding of the last_significant_coeff_flag, the same context modeling as for the significant_coeff_flag can be used. On the other hand, the probability measurement for the last_significant_coeff_flag mainly depends on the distance to the upper left corner of the transform block at the current scan position. In a preferred embodiment, the context model for the coding of the last_significant_coeff_flag is selected based on the scan diagonal where the current scan position is located (i.e., in the case of the above embodiment of FIG. 8, with x and y being the horizontal and vertical positions of the scan position within the transform block respectively, based on x + y, or based on how many sub-scans are located between the current sub-scan and the upper left DC (such as sub-scan index minus 1)). In a preferred embodiment of the present invention, the same context is used for different values of x + y. The distance measurement, i.e., x + y or the sub-scan index, is mapped onto a set of context models in a predetermined manner (e.g., by quantizing x + y or the sub-scan index), where the number of possible values for the distance measurement is larger than the number of available context models for coding the last_significant_coeff_flag.
[0074] In a preferred embodiment, different context modeling techniques are used for transform blocks of different sizes.
[0075] Next, the coding of the absolute transform coefficient levels will be described.
[0076] In one preferred embodiment, the size of the sub-block is 2×2, and the sub-block Context modeling within the block is disabled. That is, for all transform coefficients within a 2×2 sub-block, one single context model is used. Only blocks larger than 2×2 are affected by the subdivision process. In a further desirable embodiment of the present invention, the size of the sub-block is 4×4, context modeling within the sub-block is performed as in H.264, and only blocks larger than 4×4 are affected by the subdivision process.
[0077] Regarding the scanning order, in a desirable embodiment, the zigzag scan 320 is used for scanning the sub-block 322 of the transform block 256, that is, scanning along the direction of substantially increasing frequency, and the transform coefficients within the sub-block are scanned by the inverse zigzag scan 326 (Figure 11). In a further desirable embodiment of the present invention, both the sub-block 322 and the transform coefficient levels within the sub-block 322 are scanned using the inverse zigzag scan (as shown in Figure 11 where the arrow 320 is reversed). In other desirable embodiments, the same adaptive scan used to encode the significance map is used to process the transform coefficient levels, where, since the adaptation decision is the same, exactly the same scan is used for both the encoding of the significance map and the encoding of the transform coefficient level values. Note that the scan itself usually does not depend on the selected statistics or numbers of the context model set, or on the decision to enable or disable context modeling within the sub-block.
[0078] Next, embodiments regarding context modeling for the coefficient levels will be described.
[0079] In a preferred embodiment, the context modeling for sub-blocks is similar to the context modeling for 4×4 blocks in H.264 described above. The number of context models used for encoding the coeff_abs_greater_one syntax element, and the first bin of the coeff_abs_level_minus_one syntax element, for example, equals 5 by using different sets of context models for two syntax elements. In a more preferred embodiment, the context modeling within a sub-block is disabled and only one predetermined context model is used within each sub-block. For these embodiments, the context model set for sub-block 322 is selected from a predetermined number of context model sets. The selection of the context model for sub-block 322 is based on the statistics of one or more already-encoded sub-blocks. In a preferred embodiment, the statistics used to select a set of context models for a sub-block are taken from one or more already-encoded sub-blocks in the same block 256. How the statistics are used to obtain the selected context model set is described below. In a more preferred embodiment, the statistics are taken from the same sub-blocks in previously-encoded blocks having the same block size, such as blocks 40a and 40a' in FIG. 2B. In other preferred embodiments of the present invention, the statistics are taken from defined neighboring sub-blocks in the same block that depend on the selected scan for the sub-block. Also, it is important to note that the statistics should be independent of the scan order and how the statistics are created to obtain the context model set.
[0080] In a preferred embodiment, the number of context model sets is equal to 4, while in another preferred embodiment, the number of context model sets is equal to 16. In general, the number of context model sets is not fixed and should be adapted according to the selected statistics. In a preferred embodiment, the context model set for sub-block 322 is determined based on the number of absolute transform coefficient levels greater than 2 in one or more already encoded sub-blocks. The index for the context model set is determined by mapping the number of absolute transform coefficient levels greater than 2 in a reference sub-block or a plurality of reference sub-blocks to a set of predetermined context model indices. This mapping can be performed by quantizing the number of absolute transform coefficient levels greater than 2 or by means of a predetermined table. In a more preferred embodiment, the context model set for a sub-block is determined based on the difference between the number of significant transform coefficient levels and the number of absolute transform coefficient levels greater than 2 in one or more already encoded sub-blocks. The index for the context model set is determined by mapping this difference to a set of predetermined context model indices. This mapping can be performed by quantizing the difference between the number of significant transform coefficient levels and the number of absolute transform coefficient levels greater than 2 or by means of a predetermined table.
[0081] In another desirable embodiment, if the same adaptive scan is used to process the absolute conversion coefficient levels and the significance map, the partial statistics of the sub-blocks in the same block are used to determine the set of context models for the current sub-block. Alternatively, if available, the statistics of the previously encoded sub-blocks in the previously encoded transform block may be used. This means that, for example, instead of using the absolute number of absolute conversion coefficient levels greater than 2 in the sub-block to determine the context model, the number of already encoded absolute conversion coefficient levels greater than 2 is multiplied by the ratio of the number of conversion coefficients in the sub-block to the number of already encoded conversion coefficients in the sub-block is used, or instead of using the difference between the number of significant conversion coefficient levels and the number of absolute conversion coefficient levels greater than 2 in the sub-block, the difference between the number of already encoded significant conversion coefficient levels and the number of already encoded absolute conversion coefficient levels greater than 2 is multiplied by the ratio of the number of conversion coefficients in the sub-block to the number of already encoded conversion coefficients in the sub-block is used.
[0082] Regarding context modeling within a sub-block, basically, the reverse of the latest context modeling for H.264 is adopted. This means that when the same adaptive scan is used to process the absolute transform coefficient levels and the significance map, the transform coefficient levels are encoded in the forward scan order basically, instead of the reverse scan order as in H.264. Therefore, the switching of the context model must be adapted accordingly. According to one embodiment, the encoding of the transform coefficient levels starts with the first context model for the coeff_abs_greater_one and coeff_abs_level_minus_one syntax elements, and when two coeff_abs_greater_one syntax elements equal to zero are encoded after the last context model switch, it switches to the next context model in the set. In other words, the context selection depends on the number of already encoded coeff_abs_greater_one syntax elements greater than zero in the scan order. The number of context models for coeff_abs_greater_one and the number of context models for coeff_abs_level_minus_one may be the same as those in H.264.
[0083] Therefore, the above-described embodiments are applicable to the field of digital signal processing, particularly to image and video decoders and encoders. In particular, the above-described embodiments enable the encoding of syntax elements related to transform coefficients in block-based image and video coders, using improved context modeling for syntax elements related to coefficients encoded by an entropy coder that employs probability modeling. In comparison with the prior art, improved coding efficiency is achieved, particularly for relatively large transform blocks.
[0084] Although several aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent corresponding descriptions of methods, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step represent descriptions of corresponding blocks or items or corresponding features of an apparatus.
[0085] The inventive coded signals for representing a conversion block or a significance map respectively can be stored on a digital storage medium or can be transmitted via a transmission medium such as a wireless transmission medium or a wired transmission medium like the Internet.
[0086] According to certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. The implementation can be carried out using a digital storage medium that stores electronically readable control signals thereon, such as a flexible disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM (registered trademark) or flash memory, etc., which cooperate (or are capable of cooperating) with a programmable computer system so that the respective methods are executed. Therefore, these digital storage media are computer-readable.
[0087] Some embodiments according to the present invention comprise a data carrier having electronically readable control signals, which can cooperate with a programmable computer system so that one of the methods described herein is executed.
[0088] Generally, embodiments of the present invention can be implemented as a computer program product having program code, and the program code is operative to execute one of the methods when the computer program product runs on a computer. The program code is stored, for example, on a machine-readable carrier.
[0089] Other embodiments consist of a computer program for performing one of the methods described herein and stored on a machine-readable carrier.
[0090] In other words, an embodiment of the method of the present invention is a computer program having program code for performing one of the methods described herein when the computer program runs on a computer.
[0091] A further embodiment of the method of the present invention is a data carrier (or digital storage medium or computer-readable medium) on which is recorded a computer program for performing one of the methods described herein.
[0092] A further embodiment of the method of the present invention is a data stream or sequence of signals representing a computer program for performing one of the methods described herein. The data stream or sequence of signals may be configured to be transferred via a data communication connection, such as via the Internet or the like.
[0093] A further embodiment consists of processing means, such as a computer or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0094] A further embodiment consists of a computer on which is installed a computer program for performing one of the methods described herein.
[0095] In some embodiments, a programmable logic device (e.g., a field programmable gate array) is used to perform some or all of the functions of the method described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Generally, the method is preferably performed by any hardware device. can. Generally, the method is preferably performed by any hardware device.
[0096] The above embodiments are merely illustrative of the principles of the present invention. It is understood that variations and modifications of the configurations and details described herein will be apparent to those skilled in the art. Therefore, it is intended to be limited only by the claims set forth immediately hereinafter, rather than by the specific details presented by the description and illustration of the embodiments shown herein.
Claims
1. 1. An apparatus for decoding a significance map indicating locations of significant transform coefficients within a transform coefficient block from a data stream, comprising: a decoder (250) for sequentially extracting from said data stream, for associated positions within said transform coefficient blocks (256), at least syntax elements of a first type indicating whether a significant or non-significant transform coefficient is located at each position; an associator (252) for sequentially associating sequentially extracted syntax elements of a first type with positions of the transform coefficient blocks in a scan order dependent on positions of the significant transform coefficients indicated by previously extracted and associated syntax elements of the first type among the positions of the transform coefficient blocks; An apparatus comprising:
2. 2. The apparatus of claim 1, wherein the decoder (250) is further configured to recognize, based on information in the data stream, whether the last significant transform coefficient in the transform coefficient block is located at a position associated with a currently extracted syntax element of a first type indicating that a significant transform coefficient is to be located at this position, independent of multiple positions of the significant transform coefficients indicated by the previously extracted and associated syntax elements of a first type.
3. 3. The apparatus of claim 1 or 2, wherein the decoder (250) is further configured to extract from the bitstream, between a first type syntax element indicating that a significant transform coefficient is located at each of the associated positions and an immediately following first type syntax element, a second type syntax element indicating, for the associated positions at which a significant transform coefficient is located, whether the associated positions are the last significant transform coefficient in the transform coefficient block.
4. 4. The apparatus of claim 1, wherein the decoder (250) is further configured to serially extract values of the significant transform coefficients in the transform coefficient block from the data stream by context adaptive entropy decoding after extraction of all first syntax elements of the transform coefficient block, and the associator (252) is configured to associate the sequentially extracted values with positions of the significant transform coefficients in a predetermined coefficient scanning order among positions of the transform coefficient block, and to associate the transform coefficient blocks with sub-blocks (320) of the transform coefficient block (256) according to the predetermined coefficient scanning order. 22), and further configured to scan positions of the transform coefficients within the sub-blocks (322) in a position sub-scan order (324), and to use a number of contexts of a selected set from a number of contexts of a plurality of sets when the decoder sequentially context-adaptive entropy decodes the values of the significant transform coefficient values, the selection of the selected set being performed for each sub-block depending on values of the transform coefficients within a sub-block of the transform coefficient block already traversed in the sub-block scan order (320) or values of the transform coefficients of co-located sub-blocks in a previously decoded transform coefficient block of equal size.
5. 5. The apparatus of claim 1, wherein the decoder is configured to sequentially extract the first type syntax elements by entropy decoding in a context-adaptive manner by using a context selected individually for each of the first type syntax elements according to a number of positions at which significant transform coefficients are located according to the previously extracted and associated first type syntax elements in a neighborhood of a position with which each of the first type syntax elements is associated.
6. 6. The apparatus of claim 5, wherein the decoder is further configured such that the neighborhood of the location with which each of the first type syntax elements is associated consists of locations that are directly adjacent to, or directly adjacent to or spaced apart from, the location with which each of the first type syntax elements is associated in at most one vertical position and / or one horizontal position, and the size of the transform coefficient clock is greater than or equal to 8x8 positions.
7. 7. The apparatus of claim 5 or 6, wherein the decoder is further configured to map a number of positions at which significant transform coefficients are located according to the previously extracted and associated first type syntax elements to a context index of a predetermined set of possible context indexes, weighted by a number of available positions in the neighborhood of the position at which the respective first type syntax element is associated.
8. 8. The apparatus of claim 1, wherein the associator (252) is further configured to associate the sequentially extracted syntax elements of the first type with positions of the transform coefficient block along a series of sub-paths extending between a first pair of adjacent sides of the transform coefficient block along which a position of the lowest frequency in a horizontal direction and a position of the highest frequency in a vertical direction are respectively located, and a second pair of adjacent sides of the transform coefficient block along which the position of the lowest frequency in the vertical direction and the position of the highest frequency in the horizontal direction are respectively located, the sub-paths having increasing distances from the position of the lowest frequency in both the vertical and horizontal directions, and the associator (252) is configured to determine a direction (300, 302) in which the sequentially extracted syntax elements of the first type are associated with positions of the transform coefficient block based on a position of the significant transform coefficient within the previous sub-scan.
9. 1. An apparatus for decoding a significance map indicating locations of significant transform coefficients within a transform coefficient block from a data stream, comprising: a decoder (250) configured to extract a significance map indicating locations of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block from a data stream, and, in extracting the significance map, sequentially extract from the data stream syntax elements of a first type by context adaptive entropy decoding, the syntax elements of the first type indicating, for an associated position within the transform coefficient block, whether a significant or non-significant transform coefficient is located at each position; an associator (250) configured to sequentially associate the sequentially extracted syntax elements of a first type with positions of the transform coefficient blocks in a predetermined scan order among positions of the transform coefficient blocks; Equipped with The decoder is configured to use, during context-adaptive entropy decoding of the first type syntax elements, a context that is individually selected for each of the first type syntax elements according to a number of positions where significant transform coefficients are located according to the previously extracted and associated first type syntax elements in a neighborhood of a position to which a current first type syntax element is associated. Device.
10. 10. The apparatus of claim 9, wherein the decoder further comprises: determining whether the neighborhood of locations associated with each of the first type syntax elements is at most 100% in a vertical direction; 11. An apparatus configured such that at one position and / or at one horizontal position, each of the first type syntax elements is immediately adjacent to an associated position, or is immediately adjacent or spaced apart, and the size of the transform coefficient block is equal to or greater than 8x8 positions.
11. 11. The apparatus of claim 9 or 10, wherein the decoder (250) is further configured to map a number of positions at which significant transform coefficients are located according to the previously extracted and associated first type syntax elements to a context index of a predetermined set of possible context indexes, weighted by a number of available positions in the neighborhood of the position at which the respective first type syntax element is associated.
12. 1. An apparatus for decoding a transform coefficient block, comprising: a decoder (250) configured to extract a significance map indicating locations of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block from a data stream, and in extracting the values of the significant transform coefficients, sequentially extracting the values by context adaptive entropy decoding; an associator (252) configured to sequentially associate the sequentially extracted values to positions of the significant transform coefficients in a predetermined coefficient scan order between positions of the transform coefficient block, the transform coefficient block being scanned into sub-blocks (322) of the transform coefficient block (256) using a sub-block scan order (320) according to the scan order, and additionally scanning positions of the transform coefficients within the sub-blocks in a position sub-scan order (324); Equipped with The decoder (250) is configured to use a selected set of multiple contexts from a plurality of sets of multiple contexts in sequential context-adaptive entropy decoding of the significant transform coefficient values, the selection of the selected set being performed for each sub-block depending on values of the transform coefficients in sub-blocks of the transform coefficient block already traversed in the sub-block scanning order or values of the transform coefficients of co-located sub-blocks in a previously decoded transform coefficient block of equal size. Device.
13. 13. The apparatus of claim 12, wherein the decoder is configured such that a number of contexts in the plurality of sets of contexts is greater than one, and configured to uniquely assign the contexts of the selected set of number of contexts to positions within the respective sub-blocks when sequentially context-adaptive entropy decoding the significant transform coefficient values within the sub-blocks using a selected set of number of contexts for the respective sub-blocks.
14. 14. An apparatus according to claim 12 or 13, wherein the associator (252) is configured such that the sub-block scanning order proceeds in a zigzag manner from the sub-block containing the location of the lowest frequency in the vertical and horizontal directions to the sub-block containing the location of the highest frequency in both the vertical and horizontal directions, while the position sub-scanning order proceeds in a zigzag manner, in each sub-block, from the location in the respective sub-block associated with the highest frequency in the vertical and horizontal directions to the location of the respective sub-block associated with the lowest frequency in both the vertical and horizontal directions.
15. 12. The method of decoding a transform coefficient block using an apparatus (150) according to any one of claims 1 to 11 for decoding a significance map indicating the location of significant transform coefficients within a transform coefficient block from a data stream. and performing a spectral-domain to spatial-domain transform on the block of transform coefficients.
16. 1. A predictive decoder comprising: a transform-based decoder (150, 152) configured to decode a transform coefficient block using the apparatus of any one of claims 1 to 11 for decoding a significance map indicating the location of significant transform coefficients in a transform coefficient block from a data stream and to perform a spectral-domain to spatial-domain transform on the transform coefficient block to obtain a residual block; a predictor (156) configured to provide predictions for blocks of an array of information samples representing a spatially sampled information signal; a combiner (154) for combining the prediction of said block with said residual block to reconstruct said array of information samples; A predictive decoder comprising:
17. An apparatus for encoding a significance map indicating the locations of significant transform coefficients within a transform coefficient block into a data stream, the apparatus being configured to sequentially encode a first type syntax element into the data stream by entropy encoding, the first type syntax element indicating, for an associated position within the transform coefficient block, at least whether a significant or non-significant transform coefficient is located at each position, the apparatus being further configured to encode the first type syntax elements into the data stream in a scan order according to the positions of the significant transform coefficients indicated by previously encoded first type syntax elements between positions of the transform coefficient block.
18. 11. An apparatus for encoding a significance map indicating locations of significant transform coefficients within a transform coefficient block into a data stream, the apparatus being configured to encode a significance map indicating locations of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block into the data stream, and when encoding the significance map, sequentially encoding first type syntax elements into the data stream by context-adaptive entropy coding, the first type syntax elements indicating, with respect to an associated position within the transform coefficient block, whether a significant or non-significant transform coefficient is located at each position, the apparatus being further configured to sequentially code the first type syntax elements into the data stream in a predetermined scan order between positions of the transform coefficient block, and the apparatus being configured to context-adaptively entropy code each of the first type syntax elements using a context associated with a previously coded first type syntax element in a neighborhood of a position associated with a current first type syntax element.
19. An apparatus for coding a transform coefficient block, the apparatus being configured to code a significance map indicating positions of significant transform coefficients in the transform coefficient block and values of the significant transform coefficients in the transform coefficient block into the data stream, wherein when extracting the values of the significant transform coefficients, the apparatus is configured to code the values into the data stream in a predetermined coefficient scan order between positions of the transform coefficient block, the transform coefficient block being scanned in sub-blocks of the transform coefficient block using a sub-block scan order according to the predetermined coefficient scan order, and further to scan the positions of the transform coefficients in the sub-blocks in a position sub-scan order, the apparatus further being configured to sequentially context-adaptively encode the values of the significant transform coefficient values.
1. An apparatus configured to use a selected set of multiple contexts from a plurality of sets of multiple contexts when entropy encoding, the selection of the selected set being performed for each sub-block depending on values of the transform coefficients in sub-blocks of the transform coefficient block that have already been traversed in the sub-block scanning order or values of the transform coefficients of co-located sub-blocks in a previously coded transform coefficient block of equal size.
20. 1. A method for decoding a significance map indicating locations of significant transform coefficients within a transform coefficient block from a data stream, comprising the steps of: sequentially extracting syntax elements of a first type from the data stream, the syntax elements of the first type indicating, for an associated position within the block of transform coefficients, at least whether a significant or a non-significant transform coefficient is located at each said position; sequentially associating the sequentially extracted syntax elements of a first type to positions of the transform coefficient blocks in a scan order that is dependent on positions of the significant transform coefficients indicated by previously extracted and associated syntax elements of a first type among the positions of the transform coefficient blocks; The method according to claim 1,
21. 1. A method for decoding a significance map indicating locations of significant transform coefficients within a transform coefficient block from a data stream, comprising the steps of: extracting a significance map indicating locations of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block from a data stream, wherein in extracting the significance map, sequentially extracting syntax elements of a first type from the data stream by context adaptive entropy decoding, the syntax elements of the first type indicating, for an associated position within the transform coefficient block, whether a significant or non-significant transform coefficient is located at each position; sequentially associating the sequentially extracted syntax elements of a first type to positions of the transform coefficient blocks in a predetermined scanning order among positions of the transform coefficient blocks; having In the context-adaptive entropy decoding of the first type syntax elements, a context is used that is individually selected for each of the first type syntax elements according to a number of positions where significant transform coefficients are located according to previously extracted and associated first type syntax elements in a neighborhood of a position to which a current first type syntax element is associated. method.
22. 1. A method for decoding a transform coefficient block, comprising the steps of: extracting a significance map indicating locations of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block from a data stream, wherein in extracting the values of the significant transform coefficients, the values are extracted sequentially by context adaptive entropy decoding; sequentially associating the sequentially extracted values to positions of the significant transform coefficients in a predetermined coefficient scan order among positions of the transform coefficient block, the transform coefficient block being scanned into sub-blocks of the transform coefficient block using a sub-block scan order according to the predetermined coefficient scan order, and auxiliary scanning the positions of the transform coefficients within the sub-blocks in a position sub-scan order; having When the significant transform coefficient values are sequentially context adaptive entropy decoded, a selected set of multiple contexts from a plurality of sets of multiple contexts are used, and the selection of the selected set is performed for each sub-block depending on values of the transform coefficients in sub-blocks of the transform coefficient block that have already been traversed in the sub-block scanning order or values of the transform coefficients of co-located sub-blocks in a previously decoded transform coefficient block of equal size. method.
23. 1. A method for encoding into a data stream a significance map indicating locations of significant transform coefficients within a block of transform coefficients, the method comprising the steps of: sequentially encoding by entropy coding syntax elements of a first type into the data stream, the syntax elements of the first type indicating for an associated position within the transform coefficient block at least whether a significant or non-significant transform coefficient is located at each said position, and encoding the syntax elements of the first type into the data stream in a scan order according to the positions of the significant transform coefficients indicated by previously coded syntax elements of the first type among the positions of the transform coefficient blocks. The method according to claim 1,
24. 1. A method for encoding into a data stream a significance map indicating locations of significant transform coefficients within a block of transform coefficients, the method comprising the steps of: coding a significance map indicating positions of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block into the data stream, wherein in coding the significance map, sequentially coding first type syntax elements into the data stream by context-adaptive entropy coding, the first type syntax elements indicating, for an associated position within the transform coefficient block, whether a significant or non-significant transform coefficient is located at each position, the sequential coding of the first type syntax elements into the data stream being performed in a predetermined scan order among positions of the transform coefficient block, and in context-adaptively entropy coding each of the first type syntax elements being selected individually for the first type syntax element according to a number of positions at which significant transform coefficients are located in a neighborhood of a position associated with a current first type syntax element and using a context associated with a previously coded first type syntax element. A method for providing the above.
25. 1. A method for encoding a block of transform coefficients, comprising the steps of: coding a significance map indicating positions of significant transform coefficients within the transform coefficient block and values of the significant transform coefficients within the transform coefficient block into a data stream, wherein when coding the values of the significant transform coefficients, sequentially coding the values by context adaptive entropy coding and coding the values into the data stream is performed in a predefined coefficient scan order between positions of the transform coefficient block, the transform coefficient block being scanned in sub-blocks of the transform coefficient block using a sub-block scan order according to the predefined coefficient scan order, and additionally scanning the positions of the transform coefficients within the sub-blocks in a position sub-scan order and sequentially context adaptive entropy coding the values of the significant transform coefficient values, a selected set of multiple contexts from a plurality of sets of multiple contexts is used, and a selection of the selected set is performed for each sub-block depending on values of the transform coefficients within sub-blocks of the transform coefficient block already traversed in the sub-block scan order or values of the transform coefficients of co-located sub-blocks in a previously coded transform coefficient block of equal size. A method for providing the above.
26. A data stream having a significance map encoded therein indicating the locations of significant transform coefficients within a transform coefficient block, wherein syntax elements of a first type are sequentially coded into the data stream by entropy coding, the syntax elements of the first type indicating, for an associated position within the transform coefficient block, at least whether a significant or non-significant transform coefficient is located at each position, and the syntax elements of the first type are coded into the data stream in a scan order according to the positions of the significant transform coefficients indicated by previously coded syntax elements of the first type between positions of the transform coefficient blocks.
27. 1. A data stream having encoded therein a significance map indicating locations of significant transform coefficients within a transform coefficient block, followed by values of the significant transform coefficients within the transform coefficient block, wherein within the significance map, first type syntax elements are sequentially coded into the data stream by context-adaptive entropy coding, the first type syntax elements indicating, for an associated position within the transform coefficient block, whether a significant or non-significant transform coefficient is located at each respective position, the first type syntax elements being coded into the data stream in a predetermined scanning order between positions of the transform coefficient blocks, and the first type syntax elements being context-adaptively entropy coded into the data stream using a context associated with a previous first type syntax element that was individually selected for the first type syntax element according to a number of positions at which significant transform coefficients are located in a neighborhood of a position associated with a current first type syntax element.
28. 1. A data stream comprising a significance map indicating positions of significant transform coefficients within a transform coefficient block followed by coding of values of the significant transform coefficients within the transform coefficient block, wherein the values of the significant transform coefficients are coded sequentially to the data stream in a predefined coefficient scan order between positions of the transform coefficient blocks by context-adaptive entropy coding, the transform coefficient blocks being scanned in sub-blocks of the transform coefficient block using a sub-block scan order according to the predefined coefficient scan order, and additionally scanning the positions of the transform coefficients within the sub-blocks in a position sub-scan order, and the values of the significant transform coefficient values are coded sequentially to the data stream using a selected set of multiple contexts from a plurality of sets of multiple contexts, the selection of the selected set being performed for each sub-block depending on the values of the transform coefficients in sub-blocks of the transform coefficient block that have already been traversed in the sub-block scan order or the values of the transform coefficients of sub-blocks that are co-located in a previously coded transform coefficient block of equal size.
29. A computer readable digital storage medium storing a computer program having a program code for performing the method according to any one of claims 23 to 25 when the computer program is run on a computer.
Citation Information
Patent Citations
Method and Apparatus for Encoding Transform Coefficients in Image and / or Video Encoders and Decoders and Corresponding Computer Programs and Corresponding Computer Readable Storage Medium
JP2005530375A
Method and arrangement for coding transform coefficients in picture and / or video coders and decoders and a corresponding computer program and a corresponding computer-readable storage medium
US20040114683A1
Method of CABAC Coefficient Magnitude and Sign Decoding Suitable for Use on VLIW Data Processors
US20080266151A1
Early exit techniques for digital video motion estimation
WO2003094528A1
Method and arrangement for encoding transformation coefficients in image and / or video encoders and decoders, corresponding computer program, and corresponding computer-readable storage medium
WO2003094529A2