Entropy coding of motion vector difference
Patent Information
- Application Number
- JP2026085351
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2011-07-15
- Filing Date
- 2026-05-21
- Publication Date
- 2026-09-08
AI Technical Summary
【0008】 さらに、動きベクトル差のエントロピー符号化に関する上述の設定は、それを動きベクトル差の高度な方法と組み合わされ、送信される動きベクトル差の必要な合計を減らすときに特に有用である。たとえば、多重の動きベクトル予測因子は動きベクトル予測因子のオーダーされたリストを得るように提供され、この動きベクトル予測因子のリストへのインデックスは、その予測残余が問題になっている動きベクトル差によって表される実際の動きベクトル予測因子を決定するように用いられる。使用されるリスト·インデックスに関する情報が復号側でデータ·ストリームから導き出されなければならないにもかかわらず、動きベクトルの全体の予測品質は増加し、したがって、動きベクトル差の大きさは更に減少し、全体で、符号化効率は更に増加し、このような改良された動きベクトル予測に向いている動きベクトル差の水平および垂直成分のためのカットオフ値およびコンテキストの共通の使用を減少させる。一方では、データ·ストリームの中で送信される動きベクトル差の数を減らすために結合が使用されることができ、このために、結合情報は、一群のブロックに分類されるブロックの再分割のデコーダ·ブロックに信号を送っているデータ·ストリームの中で伝達される。動きベクトル差はそれから個々のブロックの代わりにこれらの合併されたグループを単位にするデータ·ストリームの中で送信されることができ、それによって、送信されなければならない運動ベクトル差の数を減少させる。このブロックのクラスタリングが隣接した動きベクトル差の間の相関関係を減らすので、上述のビン位置のためのそれぞれのコンテキストの提供の省略が、隣接する動きベクトル差によるコンテキストへのあまりに細かい分類からエントロピー符号化スキームを抑える。むしろ、結合概念はすでに隣接するブロックの動きベクトル差の間の相互関係を利用し、したがって、1つのビン位置-水平および垂直成分のための-のための1つのコンテキストは充分である。 本出願の好ましい実施例は、図面を参照して以下に記載される。
Smart Images

Figure 2026143473000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an entropy encoding concept for encoded video data.
Background Art
[0002] Many video codecs are known in the art. These codecs typically reduce the amount of data required to represent video content, that is, they compress data. It is known in the context of video coding that compression of video data is conveniently achieved by different sequentially applied coding techniques, and motion compensated prediction is used to predict image content. Motion vectors determined by motion compensated prediction and prediction residuals rely on lossless entropy coding. To further reduce the amount of data, the motion vectors themselves are predicted, whereby only the motion vector difference, which merely represents the motion vector prediction residual, needs to be entropy encoded. For example, in H.264, the procedure just outlined is applied to transmit information related to motion vector differences. In particular, motion vector differences are binarized into bin strings corresponding to a combination of truncated unary codes and exponential Golomb codes starting from a specific cutoff value. Several contexts are provided for the first bin, while bins of exponential Golomb codes are easily encoded using an equal-probability bypass mode with a fixed probability of 0.5. The cutoff value is selected to be 9. Accordingly, abundant contexts are provided for encoding motion vector differences.
Summary of the Invention
Problem to be Solved by the Invention
[0003] However, providing a rich context not only increases the complexity of coding but can also negatively impact coding efficiency. If the context is not accessed sufficiently, the probabilistic fit—that is, the fit of probabilistic evaluations associated with each context among the causes of entropy coding—will not be performed effectively. Thus, poorly applied probabilistic evaluations will estimate the actual symbol statistics. Furthermore, if several contexts are provided for a particular bin of binarization, the selection among them requires checking adjacent bin / syntax element values, which would otherwise hinder the execution of the decoding process. On the other hand, if too few contexts are provided, the bins of very diverse actual symbol statistics will be grouped within the scope of a single context, and therefore, the probabilistic evaluations associated with that context will fail to effectively code the bins associated with it.
[0004] There is a continuing need to further improve the encoding efficiency of entropy coding of motion vector differences.
[0005] Therefore, the object of the present invention is to provide this type of coding concept. [Means for solving the problem]
[0006] This objective is achieved by the subject matter of the accompanying separate claim.
[0007] The fundamental discovery of this invention is that the coding efficiency of entropy coding of motion vector differences can be further increased by reducing the cutoff value to 2 such that there are only two bin positions of the shortened unary code until a shortened unary code is used to binarize the motion vector difference, and the order of 1 is used for the exponential Golomb code for motion vector differences from the cutoff value, and furthermore, one context is certainly incomplete. When given to two positions in a unary code, context selection based on bin or syntax element values of adjacent image blocks is unnecessary, avoiding overly fine classification of these bin positions into the context, allowing established fitting to work properly, and reducing the negative effects of overly fine context subdivision when the same context is used for horizontal and vertical components.
[0008] Furthermore, the above-mentioned setting for entropy coding of motion vector differences is particularly useful when combined with advanced methods of motion vector differences to reduce the required sum of the transmitted motion vector differences. For example, multiple motion vector predictors are provided to obtain an ordered list of motion vector predictors, and the index to this list of motion vector predictors is used to determine the actual motion vector predictor represented by the motion vector difference in question, whose prediction residue. Although information about the list index used must be derived from the data stream on the decoding side, the overall prediction quality of the motion vector increases, and therefore the magnitude of the motion vector difference decreases further, resulting in overall, further increased coding efficiency and reduced common use of cutoff values and context for the horizontal and vertical components of the motion vector difference, which is suitable for such improved motion vector prediction. On the one hand, concatenation can be used to reduce the number of motion vector differences transmitted in the data stream, for this purpose the concatenation information is transmitted in the data stream, signaling to the decoder block of the block subdivision, which is classified into a group of blocks. The motion vector differences can then be transmitted in a data stream that uses these merged groups as units instead of individual blocks, thereby reducing the number of motion vector differences that must be transmitted. Since this clustering of blocks reduces the correlation between adjacent motion vector differences, the omission of providing each context for the aforementioned bin positions restrains the entropy coding scheme from being too finely divided into contexts by adjacent motion vector differences. Rather, the coupling concept already utilizes the interrelationships between motion vector differences of adjacent blocks, and therefore one context for one bin position—for horizontal and vertical components—is sufficient. Preferred embodiments of this application are described below with reference to the drawings. [Brief explanation of the drawing]
[0009] [Figure 1] Figure 1 shows a block diagram of an encoder according to an embodiment. [Figure 2a] Figure 2a is an illustration illustrating the image-like subdivision within a block. [Figure 2b] Figure 2b is an illustration illustrating different subdivisions, such as image subdivisions, within a block. [Figure 2c] Figure 2c is an illustration illustrating different subdivisions, such as image subdivisions, within a block. [Figure 3] Figure 3 shows a block diagram of a decoder according to an embodiment. [Figure 4] Figure 4 is a block diagram showing the encoder according to the embodiment in more detail. [Figure 5] Figure 5 is a block diagram showing the decoder according to the embodiment in more detail. [Figure 6] Figure 6 is an illustrative diagram showing the transformation of blocks from the spatial domain to the spectral domain, and the resulting transformed blocks and their re-transformation. [Figure 7] Figure 7 shows a block diagram of the encoder according to the embodiment. [Figure 8] Figure 8 shows a block diagram of a decoder suitable for decoding the bit stream generated by the encoder in Figure 8, according to an embodiment. [Figure 9] Figure 9 is a block diagram showing a data packet having a multiplexed partial bitstream according to an embodiment. [Figure 10] Figure 10 is a block diagram showing a data packet with other divisions using fixed-size segments according to a further embodiment. [Figure 11] Figure 11 shows a decoder that supports mode switching according to the embodiment. [Figure 12] Figure 12 shows a decoder that supports mode switching according to a further embodiment. [Figure 13] Figure 13 shows an encoder that is compatible with the decoder in Figure 11 according to the embodiment. [Figure 14] Figure 14 shows an encoder that is compatible with the decoder in Figure 12 according to the embodiment. [Figure 15]Figure 15 shows the mapping between pStateCtx and fullCtxState / 256**E**. [Figure 16] Figure 16 shows a decoder according to an embodiment of the present invention. [Figure 17] Figure 17 shows an encoder according to an embodiment of the present invention. [Figure 18] Figure 18 is an illustrative diagram showing binarization of motion vector differences according to an embodiment of the present invention. [Figure 19] Figure 19 is an illustrative diagram showing a combination concept according to the present embodiment. [Figure 20] Figure 20 is an illustrative diagram showing a motion vector prediction scheme according to the present embodiment. DETAILED DESCRIPTION OF EMBODIMENTS
[0010] It should be noted that during the description of the drawings, elements appearing in some of these drawings are denoted by the same reference numerals in each of these drawings, and repeated descriptions of these elements are omitted as far as functionality is concerned, so as to avoid unnecessary repetition. Nevertheless, unless the contrary is explicitly indicated, the functions and descriptions provided with respect to one drawing also apply to other drawings.
[0011] In the following, first, embodiments of general video coding concepts are described with reference to Figures 1 to 10. Figures 1 to 6 relate to parts of a video codec that operate at the syntax level. The following Figures 8 to 10 relate to embodiments for parts of coding relating to conversion of syntax elements into a data stream and vice versa. Specific aspects and embodiments of the present invention are described in the form of possible implementations of the general concepts typically outlined with reference to Figures 1 to 10.
[0012] Figure 1 shows an embodiment of an encoder 10 in which aspects of the present application can be implemented.
[0013] The encoder encodes an array of information samples 20 into the data stream. The array of information samples can represent information samples corresponding to, for example, luminance values, brightness values, luma values, saturation values, etc. However, if the sample array 20 is a depth diagram generated over time, such as from a light sensor, the information samples may also be depth values.
[0014] Encoder 10 is a block-based encoder. That is, encoder 10 encodes the sample array 20 into a data stream 30 in units of blocks 40. Encoding in units of blocks 40 does not necessarily mean that encoder 10 encodes these blocks 40 which are independent of each other as a whole. Rather, encoder 10 can use the re-encoded blocks to infer or internally predict the remaining blocks, and can use the block precision to set the encoding parameters, i.e., to set how each sample array region corresponding to each block is encoded.
[0015] Furthermore, the encoder 10 is a transform encoder. That is, the encoder 10 encodes the blocks 40 using transforms in order to transfer the information samples within each block 40 from the spatial domain to the spectral domain. Two-dimensional transforms such as the DCT of the FFT are used. This is possible. Preferably, the block 40 is a quadratic or rectangular shape.
[0016] The subdivision of the sample sequence 20 into blocks 40 shown in Figure 1 is given solely for illustrative purposes. Figure 1 shows the sample sequence 20 as being subdivided into a regular two-dimensional array of non-overlapping quadratic or rectangular blocks 40 that are adjacent to each other. The size of the blocks 40 may be predetermined; that is, the encoder 10 does not transmit information about the block size of the blocks 40 to the decoder in the data stream 30. For example, the decoder can anticipate a given block size.
[0017] However, several modifications are possible. For example, blocks can overlap each other. However, the overlap may be limited to such an extent that each block has a portion that does not overlap any adjacent block, or that each sample of a block overlaps with at most one adjacent block among those arranged to be juxtaposed with the current block along a given direction. The latter means that adjacent blocks to the left or right can overlap the current block so as to completely cover it, but they cannot cover each other, and this applies to adjacent blocks in the vertical and diagonal directions.
[0018] As yet another option, the subdivision of the sample sequence 20 into block 40 is adapted to the contents of the sample sequence 20 by an encoder 10 which has subdivision information about the subdivision that is transmitted to the decoder side via bitstream 30 and used.
[0019] Figures 2a to 2c illustrate different embodiments for the subdivision of the sample array 20 into blocks 40. Figure 2a shows a quadtree-based subdivision of the sample array 20 into blocks 40 of different sizes, with the blocks shown as 40a, 40b, 40c, and 40d as the size increases. According to the subdivision in Figure 2a, the sample array 20 is firstly divided into a uniform two-dimensional arrangement of tree blocks 40d, which have individual subdivision information associated with them, due to the further subdivision of a particular tree block 40d by a quadtree or otherwise. The tree block 40d to the left of the block is subdivision into fewer blocks according to the quadtree structure. The encoder 10 can perform one two-dimensional transformation for each of the blocks shown in solid and dotted lines in Figure 2a. In other words, the encoder 10 can transform the array 20 in units of block subdivisions.
[0020] Instead of a quadtree-based repartition, a more general multitree-based repartition can be used, and the number of child nodes per hierarchy level can differ between different hierarchy levels.
[0021] Figure 2b shows another embodiment for subdivision. According to Figure 2b, the sample array 20 is first divided into macroblocks 40b arranged in a uniform two-dimensional configuration such that they are adjacent to each other but do not overlap, and each macroblock 40b relates to subdivision information relating to whether the macroblock is not subdivision, or, if it is subdivision, to subblocks of the same size in a uniform two-dimensional configuration to achieve different subdivision accuracies for different macroblocks. The result is subdivision of the sample array 20 into blocks 40 of different sizes, as represented by the different sizes shown in 40a, 40b, and 40a. As in Figure 2a, the encoder 10 performs a two-dimensional transformation of each of the blocks shown in Figure 2b, which have solid and dotted lines. Figure 2c will be discussed later.
[0022] Figure 3 shows a decoder 50 capable of decoding the data stream generated by encoder 10 to reproduce a recreated version 60 of sample sequence 20. -Dra 50 reproduces the reproduced version 60 by extracting a transformation factor block for each of the blocks 40 from the data stream 30 and performing an inverse transformation on each of the transformation factor blocks.
[0023] The encoder 10 and decoder 50 can each be configured to perform entropy coding / decoding so that information about conversion coefficients can be inserted into blocks and this information can be extracted from the data stream. Details regarding this will be described later according to different embodiments. It should be noted that the data stream 30 does not necessarily contain information about conversion coefficient blocks for all blocks 40 of the sample sequence 20. Rather, as a subset of blocks, 40 can be coded into the bit stream 30 in other ways. For example, the encoder 10 may refrain from inserting conversion coefficient blocks for specific blocks of block 40 with different coding parameters inserted into the bit stream 30, and instead decide to make the decoder 50 predictable or otherwise allow each block to be filled with a reproduced version 60. For example, the encoder 10 may perform texture analysis to determine the position of blocks in the sample sequence 20 to be filled on the decoder side by the decoder, by demonstrating texture synthesis within the bit stream.
[0024] As will be shown with respect to the following figures, the transformation coefficient blocks do not necessarily represent the spectral domain representation of the original information samples of each block 40 of the sample array 20. Rather, this type of transformation coefficient block can represent the spectral domain representation of the predicted residual of each block 40. Figure 4 shows an embodiment for this type of encoder. The encoder in Figure 4 includes a transformation stage 100, an entropy encoder 102, an inverse transformation stage 104, a predictor 106, a subtractor 108, and an adder 110. The subtractor 108, the transformation stage 100, and the entropy encoder 102 are connected in that order between the input 112 and the output 114 of the encoder in Figure 4. The inverse transformation stage 104, the adder 110, and the predictor 106 are connected in that order between the output of the transformation stage 100 and the inverse input of the subtractor 108, and the output of the predictor 106 is further connected to the input of the adder 110.
[0025] The encoder in Figure 4 is a predictive transform-based block coder. That is, a block of sample sequence 20 receiving input 112 is predicted from previously encoded and reproduced portions of the same sample sequence 20 or from other previously encoded and reproduced sample sequences that precede or follow the current sample sequence 20 at presentation time. The prediction is performed by the predictive means 106. The subtractor 108 subtracts the prediction from the original block of this kind, and the transform stage 100 performs a two-dimensional transform on the prediction residue. The two-dimensional transform itself or the next measurement within the transform stage 100 leads to the quantization of the transform coefficients within the transform coefficient block. The quantized transform coefficient block is encoded in lossless compression by entropy encoding, for example, within the entropy encoder 102, and the resulting data stream is output at output 114. The inverse transform stage 104 reconstructs the quantized residue, and the adder 110 then combines the reproduced residue with the corresponding prediction to obtain a reproduced information sample based on the predictor 106 being able to predict the aforementioned current encoded prediction block. The predictor 106 can use different prediction modes, such as intra-prediction mode and inter-prediction mode, to predict blocks, and the prediction parameters are sent to the entropy encoder 102 for insertion into the data stream. For each inter-predicted prediction block, the respective motion data is inserted into the bit stream via the entropy encoder 114 to allow the decoder to re-predict. The motion data for an image prediction block is compared to a motion vector predictor, which is extracted from the motion vectors of adjacent already encoded prediction blocks in the manner described above. It includes a syntax section containing syntactic elements that indicate the motion vector difference, which encodes the motion vectors differently for the current prediction block.
[0026] In other words, according to the embodiment in Figure 4, the transformation coefficient block represents the spectral representation of the residual sample sequence rather than its actual information sample. That is, according to the embodiment in Figure 4, the sequence of syntactic elements is input to the entropy encoder 102 for entropy encoding into the data stream 114. The sequence of syntactic elements includes, for the transformation block, a motion vector difference syntax for the interprediction block, a syntax for a significance map indicating the location of significant transformation coefficient levels, and a syntax defining the significant transformation coefficient levels themselves.
[0027] Several modifications exist in the embodiment shown in Figure 4, some of which are described in the introductory section of the specification and are incorporated here into the description of Figure 4.
[0028] Figure 5 shows a decoder capable of decoding the data stream generated by the encoder in Figure 4. The decoder in Figure 5 includes an entropy decoder 150, an inverse transform stage 152, an adder 154, and a predictor 156. The entropy decoder 150, the inverse transform stage 152, and the adder 154 are connected in that order sequentially between the input 158 and the output 160 of the decoder in Figure 5. A further output of the entropy decoder 150 is connected to the predictor 156, which is connected between the output of the adder 154 and its further input. The entropy decoder 150 extracts a block of transformation coefficients from the data stream input to the decoder in Figure 5 at input 158, and the inverse transform is applied to the block of transformation coefficients at stage 152 to obtain a residual signal. The residual signal is combined with the prediction from the predictor 156 at adder 154 to obtain a reproduced block of a reproduced version of the sample sequence at output 160. Based on the reproduced version, predictor 156 generates a prediction, thereby restoring the prediction performed by predictor 106 on the encoder side. To obtain the same predictions used on the encoder side, predictor 156 also uses prediction parameters obtained from the data stream at input 158 by entropy decoder 150.
[0029] In the embodiments described above, it should be noted that the residual prediction and transformation do not need to be the same as each other. This is shown in Figure 2C, which shows the subdivision for prediction blocks of prediction accuracy (shown by solid lines) and residual accuracy (shown by dotted lines). As can be seen, the subdivisions can be selected independently by the encoder. For greater accuracy, the data stream syntax may take into account a definition of residual subdivision independent of prediction subdivision. Alternatively, residual subdivision may be an extension of prediction subdivision, such that each residual block is equivalent to or a suitable subset of the prediction blocks. For example, this is shown in Figures 2a and 2b, where prediction accuracy is shown by solid lines and residual accuracy by dotted lines. That is, in Figures 2a-2c, all blocks with associated reference signs are residual blocks on which one two-dimensional transformation is performed, while solid blocks larger than the dotted block 40a are prediction blocks on which, for example, prediction parameter setting is performed individually.
[0030] The above embodiments share the common feature that a block of the sample (residual or original) is converted into a conversion coefficient block on the encoder side, which is then inversely converted into a reproduced block of the sample on the decoder side. This is shown in Figure 6. Figure 6 shows a block of sample 200. In Figure 6, this block 200 is a sample of the second-order 4x4 sample 202 in size. Sample 202 is uniformly distributed along the horizontal x and vertical y directions. Through the two-dimensional conversion T described above, block 200 is converted into a block 204 of the spectral region, i.e., conversion coefficient 206, and the converted block 204 is the same size as block 200. That is, the converted block 200 is in the horizontal and vertical directions In the direction of transformation, block 200 has the same number of transformation coefficients 206 as block 200 has samples. However, since transformation T is a spectral transformation, the positions of the transformation coefficients 206 within transformation block 204 do not correspond to spatial positions, but rather to the spectral components of the contents of block 200. In particular, the horizontal axis of transformation block 204 corresponds to an axis along which the spectral frequency increases monotonically in the horizontal axis, and the vertical axis corresponds to an axis along which the spatial frequency increases monotonically in the vertical axis. The DC component transformation coefficients are located in the corners of block 204—here, the upper left corner as an example—and thus, in the lower right corner, the transformation coefficient 206 corresponding to the highest frequency in both the horizontal and vertical directions is located. Ignoring the spatial direction, the spatial frequencies to which a particular transformation coefficient 206 belongs generally increase from the upper left corner to the lower right corner. Inverse transformation T -1 By doing so, the transformation block 204 is transformed from the spectral domain to the spatial domain, and a copy 208 of block 200 can be obtained again. If no quantization / losses were introduced during the transformation, the reconstruction is perfect.
[0031] As already mentioned above, Figure 6 shows that a larger block size in block 200 increases the spectral resolution of the resulting spectral representation 204. On the one hand, quantization noise tends to spread across all blocks 208, and therefore sudden and very local objects within the range of block 200 tend to lead to deviations in the re-transformed block compared to the original block 200 due to quantization noise. However, the main advantage of using larger blocks is that, on the one hand, the number of significant transformation coefficients, i.e., the ratio between non-zero (quantized) transformation coefficients, i.e., levels, and on the other hand, the number of insignificant transformation coefficients, decreases in larger blocks compared to smaller blocks, thereby enabling superior coding efficiency. In other words, often, significant transformation coefficient levels, i.e., transformation coefficients that are not quantized to zero, are sparsely distributed across the transformed block 204. Thus, according to the embodiment described in more detail below, the location of significant transformation coefficient levels is indicated in the data stream by means of a significant map. Apart from there, the value of a significant transformation coefficient, i.e., the transformation coefficient level in the case of a quantized transformation coefficient, is transmitted in the data stream.
[0032] All the encoders and decoders described above are thus configured to handle specific syntax of syntactic elements. That is, the aforementioned syntactic elements such as the transformation coefficient level, syntactic elements relating to the effectiveness map of transformation blocks, and motion data syntactic elements relating to interprediction blocks are presumed to be arranged sequentially in the data stream in a predetermined manner. This kind of predetermined manner is represented, for example, in the form of pseudocode, as is done in the H.264 standard or other video codecs.
[0033] To put it another way, the above description initially addressed the conversion of video data as media data, in this case, a sequence of syntactic elements according to a predetermined syntactic structure that defines specific syntactic elements, their meaning, and the order between them. The entropy encoder and entropy decoder in Figures 4 and 5 are configured and constructed to operate as outlined below. They are responsible for performing conversions between sequences of syntactic elements and data streams, i.e., symbol or bit streams.
[0034] An entropy encoder according to the embodiment is illustrated in Figure 7. The encoder converts a stream of syntactic elements 301 into a set of two or more partial bit streams 312 in lossless compression.
[0035] In a preferred embodiment of the present invention, each syntactic element 301 is associated with one or more categories, i.e., categories of syntactic element types. For example, a category can define the type of syntactic element. In the context of hybrid video coding, another category could be a macroblock coding mode, a block coding mode, or a reference image index. This may be related to motion vector differences, subdivision flags, coding block flags, quantization parameters, transformation coefficient levels, etc. Different classifications of syntactic elements are possible in other application areas such as audio, spoken language, text, documents, or general data coding.
[0036] In general, each syntactic element can take the value of a finite or countable set of values, and the possible sets of syntactic element values can differ for different syntactic element categories. For example, there are binary syntactic elements as well as integer-value elements.
[0037] To reduce the complexity of the encoding and decoding algorithms, and to allow for general encoding and decoding designs for different syntactic elements and syntactic element categories, syntactic elements 301 are converted into an ordered set of binary decisions, which are then processed by a simple binary encoding algorithm. Thus, the binarizer 302 bijectively maps the value of each syntactic element 301 to a sequence (or string or word) of bins 303. The sequence of bins 303 represents an ordered binary decision. Each bin 303 or binary decision can take one value from a set of two values, for example, one of the values 0 and 1. The binarization scheme may differ for different syntactic element categories. The binarization scheme for a particular syntactic element category may depend on the set of possible syntactic element values and / or other properties of the syntactic element for that particular category.
[0038] Table 1 illustrates three examples of binarization schemes for countable infinite sets. Binarization schemes for countable infinite sets can also be applied to finite sets of syntactic element values. In particular, for large finite sets of syntactic element values, the inefficiencies (resulting from unused sequences in the bins) can be ignored, but the generality of this type of binarization scheme provides benefits in terms of complexity and memory requirements. For small finite sets of syntactic element values, adapting the binarization scheme to the number of possible symbol values is often preferable (in terms of coding efficiency).
[0039] Table 2 illustrates binarization schemes for three embodiments of a finite set of eight values. Binarization schemes for finite sets can be derived from generic binarization schemes for countable infinite sets by modifying some sequences of bins such that a finite set of bin sequences represents a code without redundancy (and potentially reorganizing the bin sequences). As an example, the abbreviated unary binarization scheme in Table 2 was constructed by modifying the bin sequence for syntactic element 7 of the generic unary binarization (see Table 1). The zero-order incomplete and reorganized Exp-Golomb binarizations in Table 2 were created by modifying the bin sequence for syntactic element 7 of the universal Exp-Golomb order 0 binarization (see Table 1) and then reorganizing the bin sequence (the abbreviated bin sequence for symbol 7 was assigned to symbol 1). For finite sets of syntactic elements, it is also possible to use non-systematic / non-universal binarization schemes, as illustrated in the last column of Table 2.
[0040] [Table 1]
[0041] [Table 2]
[0042] Each bin 303 in the sequence of bins created by the binarizer 302 is input to the parameter assigner 304 in sequence order. The parameter assigner assigns one or more sets of parameters to each bin 303 to output bins having sets of parameters 305. The sets of parameters are determined in exactly the same way by the encoder and decoder. The sets of parameters can consist of one or more of the following parameters:
[0043] In particular, the parameter assigner 304 can be configured to assign a context model to the current bin 303. For example, the parameter assigner 304 can select one of the available context indices for the current bin 303. The available set of contexts for the current bin 303 then depends on the type of bin, defined by the type / category of the syntactic element 301, where the current bin 303 is the part of its binarization and the position of the current bin 303 in the subsequent binarization. The selection of a context among the available set of contexts can depend on the previous bin and syntactic element associated with the latter. Each of these contexts has a probabilistic model associated with it, i.e., a measure for evaluating the probability of one of two possible bin values for the current bin. The probabilistic model is measured, in particular, for estimating the probability of bin values that are less likely or more likely for the current bin, and the probabilistic model further This is defined by identifiers that specify estimates for two possible bin values representing less likely or more likely bin values for the current bin 303. Context selection is left aside if there is simply only one context available for the current bin. As outlined in more detail below, the parameter assigner 304 can also perform probabilistic model fitting to adapt probabilistic models associated with various contexts to the actual bin statistics of each bin belonging to each context.
[0044] As will be described in more detail below, the parameter assigner 304 can operate differently depending on whether it is operating in high-efficiency (HE) mode or low-complexity (LC) mode. As outlined below, in both modes, the probabilistic model associates the current bin 303 with one of the bin encoders 310, but the operating mode of the parameter assigner 304 tends to be less complex in LC mode, whereas in high-efficiency mode, the encoding efficiency is increased by the association of individual bins 303 with individual encoders 310 that are more accurately adapted to bin statistics, thereby optimizing entropy in relation to LC mode.
[0045] Each bin having an associated set of parameters 305, which are outputs of the parameter assigner 304, is fed to the bin buffer selector 306. The bin buffer selector 306 modifies the value of the input bin 305 based on the input bin value and the associated parameter 305, and places the output bin 307—which has the potentially modified value—into one of two or more bin buffers 308. The bin buffer 308 to which the output bin 307 is fed is determined based on the value of the input bin 305 and / or the value of the associated parameter 305.
[0046] In a preferred embodiment of the present invention, the bin buffer selector 306 does not modify the bin value, i.e., the output bin 307 always has the same value as the input bin 305. In a further preferred embodiment of the present invention, the bin buffer selector 306 determines the output bin value 307 based on the input bin value 305 and relevant measurements for evaluating the probability of one of two possible bin values for the current bin. In a preferred embodiment of the present invention, if the measurement for the probability of one of two possible bin values for the current bin is less than (or equal to) a certain threshold, the output bin value 307 is set to be equal to the input bin value 305; if the measurement for the probability of one of two possible bin values for the current bin is greater than or equal to (or greater than) a certain threshold, the output bin value 307 is modified (i.e., it is set to the inverse of the input bin value). In a further preferred embodiment of the present invention, if the measurement for the probability of one of the two possible bin values for the current bin is greater than (or greater than or equal to) a certain threshold, the output bin value 307 is set to be equal to the input bin value 305; if the measurement for the probability of one of the two possible bin values for the current bin is less than or equal to (or less than) a certain threshold, the output bin value 307 is modified (i.e., it is set to the inverse of the input bin value). In a preferred embodiment of the present invention, the threshold value corresponds to a value of 0.5 for the estimated probabilities for both possible bin values.
[0047] In a further preferred embodiment of the present invention, the bin buffer selector 306 determines an output bin value 307 based on an input bin value 305 and an associated identifier that identifies an evaluation in which two possible bins are unlikely or more likely to represent a bin value for the current bin. In a preferred embodiment of the present invention, if the identifier indicates that the first of the two possible bin values represents an unlikely (or more likely) bin value for the current bin, the output bin value 307 is set to equal the input bin value 305, and if the identifier indicates that the second of the two possible bin values represents an unlikely (or more likely) bin value for the current bin, the output bin value 307 is modified (i.e., (That is, it is set to the inverse of the input bin value.)
[0048] In a preferred embodiment of the present invention, the bin buffer selector 306 determines the bin buffer 308 to which the output bin 307 is sent based on relevant measurements for evaluating the probability of one of two possible bins for the current bin. In a preferred embodiment of the present invention, the set of possible values for the measurements for evaluating the probability of one of two possible bin values is limited, and the bin buffer selector 306 includes a table that associates exactly one bin buffer 308 with each possible value for evaluating the probability of one of two possible bin values, and different values for the measurements for evaluating the probability of one of two possible bin values can be associated with the same bin buffer 308. In a further preferred embodiment of the present invention, the range of possible values for the measurements for evaluating the probability of one of two possible bin values is divided into many intervals, and the bin buffer selector 306 determines an interval index for the current measurement for evaluating the probability of one of two possible bin values, and the bin buffer selector 306 includes a table that associates exactly one bin buffer 308 with each possible value for the interval index, and different values for the interval index can be associated with the same bin buffer 308. In a preferred embodiment of the present invention, an input bin 305 having an inverse measure for evaluating the probability of one of two possible bin values (where the inverse measure represents probability evaluations P and 1-P) is fed into the same bin buffer 308. In a further preferred embodiment of the present invention, for example, to ensure that the resulting partial bit streams have a similar bit rate, the association of the measures for evaluating the probability of one of two possible bin values for a given bin having a particular bin buffer is adapted over time. Furthermore, as used below, the interval index is also called the pipe index, but together with the improved index and flags indicating more likely bin values, the pipe index indexes the actual probability model (i.e., probability evaluations).
[0049] In a further preferred embodiment of the present invention, the bin buffer selector 306 determines the bin buffer 308 to which the output bin 307 is sent based on relevant measurements for evaluating the probability of an unlikely or more likely bin value for the current bin. In a further preferred embodiment of the present invention, the set of possible values for the measurements for evaluating the probability of an unlikely or more likely bin value is limited, and the bin buffer selector 306 includes a table associated with exactly one bin buffer 308 for each possible value for evaluating the probability of an unlikely or more likely bin value, and different values for the measurements for evaluating the probability of an unlikely or more likely bin value can be associated with the same bin buffer 308. In a further preferred embodiment of the present invention, the range of possible values for the measurements for evaluating the probability of an unlikely or more likely bin value is divided into many intervals, and the bin buffer selector 306 determines an interval index for the current measurement for evaluating the probability of an unlikely or more likely bin value, and different values for the interval index can be associated with the same bin buffer 308. In a further preferred embodiment of the present invention, for example, in order to ensure that the resulting partial bitstream has a similar bit rate, the association of measurements for evaluating the probabilities of unlikely or more likely bin values for the current bin having a particular bin buffer is adapted over time.
[0050] Each of two or more bin buffers 308 is connected to exactly one bin encoder 310, and each bin encoder is connected to only one bin buffer 308. Each bin encoder 310 reads bins from its associated bin buffer 308 and converts the sequence of bins 309 into a codeword 311 representing a sequence of bits. The bin buffer 308 represents a first-in, first-out buffer; bins entered into the bin buffer 308 later (in order of occurrence) are not encoded before bins entered into the bin buffer earlier (in order of occurrence). A specific bin encoder 310 The codeword 311, which is the output of the process, is written to a specific partial bit stream 312. The overall encoding algorithm converts syntactic elements 301 into two or more partial bit streams 312, where the number of partial bit streams is equal to the number of bin buffers and bin encoders. In a preferred embodiment of the present invention, the bin encoder 310 converts bins 309 of varying lengths into codewords 311 of varying bit lengths. One of the above-mentioned and below-mentioned effects of embodiments of the present invention is that bin encoding can be performed in parallel (e.g., for different groups of probabilistic measurements), which reduces processing time for some embodiments.
[0051] Another effect of embodiments of the present invention is that the bin encoding performed by the bin encoder 310 can be specifically designed for different sets of parameters 305. In particular, bin coding and encoding can be optimized (with respect to coding efficiency and / or complexity) for different groups of estimated probabilities. On the one hand, this allows for a reduction in coding / decoding complexity, and on the other hand, it allows for an improvement in coding efficiency. In a preferred embodiment of the present invention, the bin encoder 310 implements different coding algorithms (i.e., mapping of bin sequences over codewords) for different groups of measurements for evaluating the probability of one of two possible bin values 305 for the current bin. In a further preferred embodiment of the present invention, the bin encoder 310 implements different coding algorithms for different groups of measurements for evaluating the probability of unlikely or more likely bin values for the current bin.
[0052] In a preferred embodiment of the present invention, the bin encoder 310—or one or more bin encoders—represents an entropy encoder that directly maps a sequence of input bins 309 to a codeword 310. Such mapping can be performed efficiently and does not require a complex arithmetic coding engine. The inverse mapping of the codeword onto the sequence of bins (as performed in a decoder) should be excellent to ensure complete decoding of the input sequence, but the mapping of the bin sequence 309 to the codeword 310 does not necessarily have to be excellent; that is, a partial sequence of bins can be mapped onto one or more sequences of the codeword. In a preferred embodiment of the present invention, the mapping of the sequence of input bins 309 onto the codeword 310 is bijective. In a further preferred embodiment of the present invention, the bin encoder 310—or one or more bin encoders—represents an entropy encoder that directly maps a variable-length sequence of input bins 309 onto a variable-length codeword 310. In a preferred embodiment of the present invention, the output codeword represents a non-redundant code, such as a Huffman code or a standard Huffman code.
[0053] Two embodiments for a bijective mapping of bin sequences to a non-redundant code are illustrated in Table 3. In a further preferred embodiment of the present invention, the output codeword represents a redundant code suitable for error detection and error recovery. In a further preferred embodiment of the present invention, the output codeword represents an encryption code suitable for encrypting syntactic elements.
[0054] [Table 3]
[0055] [Table 4]
[0056] In a further preferred embodiment of the present invention, the bin encoder 310—or one or more bin encoders—represents an entropy encoder that maps a variable-length sequence of input bins 309 directly onto a fixed-length codeword 310.
[0057] An embodiment of the present invention is illustrated in Figure 8. The decoder essentially performs the inverse operation of an encoder, resulting in the decoding of a (previously encoded) sequence of syntactic elements 327 from a set of two or more partial bit streams 324. The decoder includes two distinct procedural flows: a data request flow that replicates the data flow of the encoder, and a data flow that represents the inverse of the encoder's data flow. In the specific example in Figure 8, dotted arrows represent the data request flow, and solid arrows represent the data flow. The components of the decoder essentially replicate the components of the encoder, but perform the inverse operation.
[0058] Decoding a syntactic element is triggered by a request for a new decoded syntactic element 313, which is sent to the binarizer 314. In a preferred embodiment of the present invention, each request for a new decoded syntactic element 313 is associated with a category from a set of one or more categories. The category associated with the request for a syntactic element is the same category that was associated with the corresponding syntactic element during encoding.
[0059] The binarizer 314 maps the request for the syntactic element 313 to one or more requests for bins that are sent to the parameter assigner 316. As a final response to the requests for bins sent by the binarizer 314 to the parameter assigner 316, the binarizer 314 receives the decoded bin 326 from the bin buffer selector 318. The binarizer 314 compares the received sequence of the decoded bin 326 to the bin sequence of a specific binarization scheme for the requested syntactic element. If the received sequence of the decoded bin 26 matches the binarization of the syntactic element, the binarizer empties its bin buffer and outputs the decoded syntactic element as a final response to the request for a new decoded symbol. If the already received sequence of the decoded bin does not match any of the bin sequences for the binarization scheme of the requested syntactic element, the binarizer sends other requests for bins to the parameter assigner until the sequence of the decoded bin matches one of the bin sequences for the binarization scheme of the requested syntactic element. For each request for a syntactic element, the decoder uses the same binarization scheme used to encode the corresponding syntactic element. The binarization scheme can differ for different syntactic element categories. The binarization scheme for a particular syntactic element category may depend on the set of possible syntactic element values and / or other properties of the syntactic element for that particular category.
[0060] The parameter assigner 316 assigns one or more sets of parameters to each request for a bin and sends a request for a bin with the relevant set of parameters to the bin buffer selector. The set of parameters assigned to a bin requested by the parameter assigner is the same as the one assigned to the corresponding bin during encoding. The set of parameters may consist of one or more of the parameters mentioned in the description of the encoder in Figure 7.
[0061] In a preferred embodiment of the present invention, the parameter assigner 316 associates each bin request with the same parameters as assigneder 304, namely, a measurement for evaluating the probability of an unlikely or more likely bin value for the currently requested bin, and a context and its associated measurement for evaluating the probability of one of two possible bin values for the currently requested bin, such as an identifier that specifies which of the two possible bin values represents the unlikely or more likely bin value for the currently requested bin.
[0062] The parameter assigner 316 can determine one or more of the above-described probabilistic measures (measures for evaluating the probability of one of two possible bin values for the currently requested bin, measures for evaluating the probability of an unlikely or more likely bin value for the currently requested bin, and identifiers that define an evaluation of which of the two possible bin values represents an unlikely or more likely bin value for the currently requested bin) based on one or more already decoded symbols. The determination of probabilistic measures for a particular request for a bin replicates the processing in the encoder for the corresponding bin. The decoded symbols used to determine the probabilistic measures are one or more already decoded symbols of the same symbol category, one or more already decoded symbols of the same symbol category corresponding to a dataset (e.g., a block or group of samples) in adjacent space and / or temporal location (with respect to the dataset associated with the current request for the syntactic element), or temporal location (with respect to the dataset associated with the current request for the syntactic element). It can contain one or more already decoded symbols from different symbol categories corresponding to a dataset at a specific location.
[0063] Each request for a bin having an associated set of parameters 317, which are outputs of the parameter assigner 316, is input to the bin buffer selector 318. Based on the associated set of parameters 317, the bin buffer selector 318 sends the request for bin 319 to one of two or more bin buffers 320 and receives the decoded bin 325 from the selected bin buffer 320. The decoded input bin 325 is potentially modified, and the decoded output bin 326, having the potentially decoded value, is sent to the binarizer 314 as the final response to the request for the bits having an associated set of parameters 317.
[0064] The bin buffer 320 to which the bin request is sent is similarly selected as the output bin of the encoder-side bin buffer selector.
[0065] In a preferred embodiment of the present invention, the bin buffer selector 318 determines the bin buffer 320 to which a request for bin 319 is sent based on relevant measurements for evaluating the probability of one of two possible bin values for the currently requested bin. In a preferred embodiment of the present invention, the set of possible values for the measurements for evaluating the probability of one of two possible bin values is limited, and the bin buffer selector 318 includes a table that associates one bin buffer 320 with each possible value for evaluating the probability of one of two possible bin values, and different values for the measurements for evaluating the probability of one of two possible bin values can be associated with the same bin buffer 320. In a further preferred embodiment of the present invention, the range of possible values for the measurements for evaluating the probability of one of two possible bin values is divided into many intervals, and the bin buffer selector 318 determines an interval index for the current measurement for evaluating the probability of one of two possible bin values, and the bin buffer selector 318 includes a table that precisely associates one bin buffer 320 with each possible value for the interval index, and different values for the interval index can be associated with the same bin buffer 320. In a preferred embodiment of the present invention, a request for bin 317 having an inverse measurement for evaluating the probability of one of two possible bin values (where the inverse measurement represents probabilities P and 1-P) is sent to the same bin buffer. In a further preferred embodiment of the present invention, the association of the measurement for evaluating the probability of one of two possible bin values for a request for a current bin having a particular bin buffer is adapted over time.
[0066] In a further preferred embodiment of the present invention, the bin buffer selector 318 determines which bin buffer 320 to send based on relevant measurements for evaluating the probability of an unlikely or more likely bin value for the currently requested bin. In a preferred embodiment of the present invention, the set of possible values for the measurements for evaluating the probability of an unlikely or more likely bin value is limited, and the bin buffer selector 318 includes a table that precisely associates one bin buffer 320 with each possible value for evaluating the probability of an unlikely or more likely bin value, and different values for the measurements for evaluating the probability of an unlikely or more likely bin value can be associated with the same bin buffer 320. In a further preferred embodiment of the present invention, the range of possible values for a measurement for evaluating the probability of an unlikely or more likely bin value is divided into many intervals, and the bin buffer selector 318 determines the interval index for the current measurement for evaluating the probability of an unlikely or more likely bin value, and the bin buffer selector 318 includes a table that precisely associates one bin buffer with each possible value for the interval index, and different values for the interval index can be associated with the same bin buffer 320. In a further preferred embodiment of the present invention, the range of possible values for the current bin requirement having a particular bin buffer The relevance of the measurement for evaluating the probability of a bin value that is not present or is more likely is adapted over time.
[0067] After receiving the decoded bin 325 from the selected bin buffer 320, the bin buffer selector 318 potentially modifies the input bin 325 and sends the output bin 326—with the potentially modified value—to the binarizer 314. The input / output bin mapping of the bin buffer selector 318 is the inverse of the input / output bin mapping of the bin buffer selector on the encoder side.
[0068] In a preferred embodiment of the present invention, the bin buffer selector 318 does not modify the bin value, i.e., the output bin 326 always has the same value as the input bin 325. In a further preferred embodiment of the present invention, the bin buffer selector 318 determines the output bin value 326 based on the input bin value 325 and a measurement for evaluating the probability of one of two possible bin values for the currently requested bin associated with the request for bin 317. In a preferred embodiment of the present invention, if the measurement for the probability of one of two possible bin values for the current bin request is less than (or equal to) a certain threshold, the output bin value 326 is set to be equal to the input bin value 325; if the measurement for the probability of one of two possible bin values for the current bin request is greater than or equal to (or greater than) a certain threshold, the output bin value 326 is modified (i.e., it is set to the inverse of the input bin value). In a further preferred embodiment of the present invention, if the measurement for the probability of one of the two possible bin values for the current bin request is greater than (or greater than or equal to) a certain threshold, the output bin value 326 is set to be equal to the input bin value 325; if the measurement for the probability of one of the two possible bin values for the current bin request is less than or equal to (or less than) a certain threshold, the output bin value 326 is modified (i.e., it is set to the inverse of the input bin value). In a preferred embodiment of the present invention, the threshold value corresponds to a value of 0.5 for the evaluation probability of both possible bin values.
[0069] In a further preferred embodiment of the present invention, the bin buffer selector 318 determines the output bin value 326 based on the input bin value 325 and an identifier that defines an evaluation of which of two possible bin values represents a less likely or more likely bin value for the current bin request associated with the request for bin 317. In a preferred embodiment of the present invention, if the identifier indicates that the first of the two possible bin values represents a less likely (or more likely) bin value for the current bin request, the output bin value 326 is set equal to the input bin value 325; if the identifier indicates that the second of the two possible bin values represents a less likely (or more likely) bin value for the current bin request, the output bin value 326 is modified (i.e., it is set to the reverse of the input bin value).
[0070] As described above, the bin buffer selector sends a request for bin 319 to one of two or more bin buffers 320. Bin buffer 20 represents a first-in, first-out buffer given by the sequence of bins 321 decoded from the connected bin decoder 322. In response to the request for bin 319 sent from the bin buffer selector 318 to bin buffer 320, bin buffer 320 moves the bin whose contents will be placed in bin buffer 320 first and sends it to the bin buffer selector 318. Bins sent to bin buffer 320 early are moved early and sent to the bin buffer selector 318.
[0071] Each of the two or more bin buffers 320 is connected to exactly one bin decoder 322, and each GON decoder is connected to only one GON buffer. Each bin decoder 322 reads a codeword 323 representing a sequence of bits from another partially bit stream 324. The bin decoder converts the codeword 323 into a sequence of bins 321 that are sent to the connected bin buffer 320. The overall decoding algorithm converts two or more partial bit streams 324 into many decoded syntactic elements, where the number of partial bit streams is equal to the number of bin buffers and bin decoders, and the decoding of syntactic elements is triggered by the request of a new syntactic element. In a preferred embodiment of the present invention, the bin decoder 322 converts a codeword 323 of a variable number of bits into a sequence of a variable number of bins 321. One effect of the embodiment of the present invention is that the decoding of bins from two or more partial bit streams can be done in parallel (e.g., for different groups of probabilistic measurements), which reduces processing time for some implementations.
[0072] Another effect of embodiments of the present invention is that the bin decoding performed by the bin decoder 322 can be specifically designed for different sets of parameters 317. In particular, bin coding and decoding can be optimized (with respect to coding efficiency and / or complexity) for different groups of probabilities being evaluated. On the one hand, this can reduce the coding / decoding complexity over a top-level entropy coding algorithm with comparable coding efficiency. On the other hand, it can improve coding efficiency over a top-level entropy coding algorithm with similar coding / decoding complexity. In a preferred embodiment of the present invention, the bin decoder 322 performs different decoding algorithms (i.e., mapping of bin sequences onto codewords) for different groups of measurements for evaluating the probability of one of two possible bin values 317 for the current bin request. In a further preferred embodiment of the present invention, the bin decoder 322 performs different decoding algorithms for different groups of measurements for evaluating the probability of unlikely or more likely bin values for the current requested bin.
[0073] The bin decoder 322 performs the reverse mapping of the corresponding bin encoder on the encoder side.
[0074] In a preferred embodiment of the present invention, bin encoder 322—or one or more bin encoders—represents an entropy decoder that directly maps a codeword 323 onto a sequence of bins 321. Such mapping can be performed efficiently and does not require a complex arithmetic coding engine. The mapping of a codeword onto a sequence of bins must be unique. In a preferred embodiment of the present invention, the mapping of a codeword 323 onto a sequence of bins 321 is bijective. In a further preferred embodiment of the present invention, bin encoder 310—or one or more bin encoders—represents an entropy decoder that directly maps a variable-length codeword 323 onto a variable-length sequence of bins 321. In a preferred embodiment of the present invention, the input codeword represents a non-redundant code, such as a general Huffman code or a standard Huffman code. Two examples of bijective mappings of non-redundant codes onto bin sequences are illustrated in Table 3.
[0075] In a further preferred embodiment of the present invention, the bin decoder 322—or one or more bin decoders—represents an entropy decoder that directly maps fixed-length codewords 323 onto a variable-length sequence 321 of bins.
[0076] Thus, Figures 7 and 8 show embodiments of an encoder for encoding a sequence of symbols 3 and a decoder for reproducing it. The encoder includes an assignor 304 configured to assign several parameters 305 to each symbol in the sequence of symbols. The assignment is syntactic to the representation to which the current symbol belongs—such as a binarized representation. Based on the information contained in previous symbols of a sequence of symbols like category 1, it is predicted that, according to the syntactic structure of syntactic element 1, which predictions can now be inferred sequentially from the history of previous syntactic elements 1 and symbols 3. Furthermore, the encoder includes a plurality of entropy encoders 10, each of which translates a symbol 3 sent to its respective entropy encoder 10 into its respective bit stream 312, and a selector 306 configured to send each symbol 3 to one of the plurality of entropy encoders 10, the selection depending on the number of parameters 305 assigned to each symbol 3. The assignor 304 is thought to be aggregated into selector 206 to obtain each selector 502.
[0077] The decoder for reconstructing the sequence of symbols includes multiple entropy decoders 322, each configured to translate its respective bit stream 323 into a symbol 321; an assignor 316 (see 326 and 327 in Figure 8) configured to assign multiple parameters 317 to each symbol 315 of the sequence of symbols to be reconstructed based on information contained in previously reconstructed symbols of the sequence of symbols; and a selector 318 configured to look up each symbol of the sequence of symbols to be reconstructed from one of the multiple entropy decoders 322 (selection based on the number of parameters defined for each symbol), the selection depending on the multiple parameters defined for each symbol. The assignor 316 is configured such that the number of parameters assigned to each symbol is included, or is a measure for evaluating the establishment of a distribution among the possible symbol values that each symbol assumes. The assignor 316 and the selector 318 can also be thought of as being incorporated into a single block, selector 402. The sequence of symbols to be reproduced may be a binary alphabet, and the assignor 316 includes a measure for evaluating the probability of an unlikely or more likely bin value for two possible bin values of the binary alphabet, and the identifier determines the evaluation of which of the two possible bin values is unlikely or more likely. The assignor 316 is configured to reproduce based on information contained in previously reproduced symbols of the sequence of symbols to be reproduced in each context, each having its own associated probability distribution evaluation, and to internally assign each symbol of the sequence of symbols 315 that adapts the probability distribution evaluation for each context to actual Sibol statistics based on symbols reproduced before each context was assigned. The context may consider the spatial relationships or neighborhoods of the locations to which syntactic elements belong, for example in video or image coding, or even in tables in the case of financial applications.Then, the measurement for evaluating the probability distribution for each symbol can be determined, for example, based on the probability distribution evaluation used as an index to each table, which is related to the context assigned to each symbol by quantization, and the probability distribution evaluation is related to the context assigned to each symbol (in the following embodiment, indexed by the pipe index together with the improved index) to one of a plurality of probability distribution evaluation representations (clipping away from the improved index) in order to obtain a measurement for evaluating the probability distribution (pipe index that indexes a partial bit stream 312). The selector can define a bijective association between a plurality of entropy encoders and a plurality of probability distribution representations. Selector 18 can be configured to change the quantization mapping over time from a range of probability distribution evaluations to a plurality of probability distribution evaluation representations in a predetermined deterministic manner, corresponding to symbols reproduced before the sequence of symbols. That is, selector 318 can change the quantization step size, i.e., the interval of the probability distributions mapped to the individual probability indices that are omni-projectively associated with the individual entropy decoders. Multiple entropy decoders 322 can be configured to adapt their methods for converting symbols into bit streams that respond to changes in quantization mapping. For example, each entropy decoder 322 may be optimized for certain probability distribution evaluations within the quantization interval of each probability distribution evaluation. That is, it is possible to have an optimal compression ratio, and the sequence mapping of the codeword / symbol can be changed to adapt the position of a particular probability distribution evaluation so that it is optimal within the range of the quantization interval of each probability distribution evaluation with respect to the latter change. The selector can be configured to change the quantization mapping so that the rate at which a symbol is retrieved from multiple entropy decoders is less dispersed. With respect to the binarizer 314, if the syntactic element is already binary, it is set aside. Furthermore, depending on the type of decoder 322, the presence of buffer 320 is not required. Furthermore, buffer 320 can be accumulated within the scope of the decoder.
[0078] End of finite syntax element array
[0079] In a preferred embodiment of the present invention, encoding and decoding are performed for a finite set of syntactic elements. Often, a specific amount of data, such as a still image, a frame, or a region of a video sequence, a slice of an image, a slice of a frame, or a region of a video sequence, or a set of consecutive audio samples, etc., is encoded. For a finite set of syntactic elements, generally, the partial bitstream created on the encoder side must be terminated, i.e., it must be ensured that all syntactic elements can be transmitted or decoded from the stored partial bitstream. After the last bin is input to the corresponding bin buffer 308, the bin encoder 310 must ensure that the complete codeword is written to the partial bitstream 312. If the bin encoder 310 represents an entropy encoder that performs a direct mapping of bin sequences onto a codeword, the bin sequence stored in the bin buffer after the last bin is written to the bin buffer may not represent the bin sequence associated with the codeword (i.e., it may represent a prefix of two or more bin sequences associated with the codeword). In such cases, one of the codewords associated with a bin sequence containing the bin sequence of the bin buffer as a prefix must be written to the partial bit stream (the bin buffer must be flushed). Until a codeword is written, this can be done by inputting bins having specific or arbitrary values into the bin buffer. In a preferred embodiment of the present invention, the bin encoder selects one of the codewords having the minimum length (in addition to the characteristic that the associated bin sequence must contain the bin sequence of the bin buffer as a prefix). On the decoder side, the bin decoder 322 may decode more bins than are required for the last codeword of the partial bit stream; these bins are not requested by the bin buffer selector 318, are discarded, and ignored. Decoding of a finite set of symbols is controlled by requests for decoded syntactic elements; decoding terminates when no further syntactic elements are requested due to the amount of data.
[0080] Partial bitstream transmission and multiplexing
[0081] The partial bitstreams 312 created by the encoder can be transmitted separately, or they can be multiplexed into a single bitstream, or the codewords of the partial bitstreams can be arranged alternately in a single bitstream.
[0082] In embodiments of the present invention, each partial bitstream for a data amount is written to a single data packet. The data amount can be any set of syntactic elements such as still photographs, fields or frames of video sequences, slices of still photographs, slices of fields or frames of video sequences, or frames of audio samples.
[0083] In other preferred embodiments of the present invention, two or more partial bitstreams or all partial bitstreams of a data amount are multiplexed into a single data packet. The structure of a data packet containing multiplexed partial bitstreams is illustrated in Figure 9.
[0084] A data packet 400 includes a header and one partition for the data of each partial bit stream (due to the considered amount of data). The header 400 of the data packet includes an indication for dividing the data packet (the rest of it) into segments of bit stream data 402. In addition to the indication for division, the header may include additional information. In a preferred embodiment of the present invention, the indication for division of the data packet is the starting position of a data segment, in units of bits or bytes or a variety of bits or a variety of bytes. In a preferred embodiment of the present invention, the starting position of a data segment is encoded as the absolute value of the title of the data packet, in relation to the beginning of the data packet, or in relation to the end of the header, or in relation to the beginning of a previous data packet. In a further preferred embodiment of the present invention, the starting position of a data segment is encoded differently, i.e., only the difference between the actual beginning of a data segment and the predicted beginning of a data segment is encoded. The prediction can be derived based on already known or transmitted information, such as the overall size of the data packet, the size of the header, the number of data segments in the data packet, or the starting position of a previous data segment. In a preferred embodiment of the present invention, the starting position of the first data packet is not encoded and is estimated based on the size of the data packet header. On the decoder side, the transmitted partition representation is used to extract the beginning of the data segment. The data segment is then used as a partial bitstream, and the data contained in the data segment is fed into the corresponding bin decoder in segment order.
[0085] There are several options for multiplexing a partial bitstream into a data packet. One option that can reduce the amount of necessary side information, especially when the partial bitstream is very small, is illustrated in Figure 10. The payload of the data packet, i.e., the data packet 410 without its header 411, is divided into segments 412 in a predetermined manner. As an example, the data packet payload can be divided into segments of the same size. Each segment then relates to a partial bitstream or a first portion of a partial bitstream 413. If the partial bitstream is larger than the data segment it relates to, the remainder 414 is placed in the unused space at the end of the other data segment. This can be done in such a way that the remainder of the bitstream is entered in reverse order (starting from the end of the data segment), which reduces side information. When a partial bitstream is related to a data segment and one or more remainders are appended to the data segment, one or more remainders must be signaled within the bitstream, for example, in the data packet header.
[0086] Interleaving of variable-length codewords
[0087] Regarding some applications, the above-mentioned multiplexing of a partial bitstream (for the amount of syntactic elements) of a single data packet may have the following disadvantages: On the one hand, for small data packets, the number of bits required for side information to signal the division can be important in relation to the actual data of the partial bitstream, which ultimately reduces encoding efficiency. On the other hand, multiplexing is not suitable for applications that require low latency (e.g., for video conferencing applications). Regarding the multiplexing described, the position of the start of the partition is before it. Because it is not known, the encoder cannot begin transmitting data packets before a partial bitstream is fully formed. Furthermore, generally, the decoder must wait until it receives the beginning of the last data segment before it can begin decoding the data packets. For video conferencing system applications, this results in further overall delays in some video image systems (especially for encoders / decoders that require close time intervals between two images for encoding / decoding), which critically impacts this type of application. To overcome the disadvantages for specific applications, an encoder in a preferred embodiment of the present invention can be configured such that codewords generated by two or more bin encoders are alternately arranged in a single bitstream. The bitstream with the alternately arranged codewords can be transmitted directly to the decoder (when small buffer delays are ignored, see below). On the decoder side, if two or more bin decoders read the codewords directly from the bitstream in decoding order, decoding can begin with the first received bit. Furthermore, side information is not required to send signals for partial bitstream multiplexing (or alternating arrangement). A further way to reduce the complexity of the decoder can be achieved when the bin decoders 322 do not read variable-length codewords from the global bit buffer, but instead, they always read a fixed-length sequence of bits from the global bit buffer, add these fixed-length sequences of bits to a local bit buffer, and each bin decoder 322 is connected to a separate local bit buffer. The variable-length codeword is then read from the local bit buffer. Thus, the parsing of variable-length codewords can be done in parallel, and only the access to the fixed-length sequences of bits must be done in a synchronous manner, but such access to fixed-length sequences of bits is usually very fast, and as a result, the complexity of the overall decoding can be reduced for some architecture.The fixed number of bins sent to a particular local bit buffer may differ for different local bit buffers, and it may also change over time depending on certain parameters as events of the bin decoder, bin buffer, or bit buffer. However, the number of bits read by a particular access does not depend on the actual bits read during that particular access, which is a key difference from reading variable-length codewords. Reading a fixed-length sequence of bits is triggered by a specific event in the bin buffer, bin decoder, or local bit buffer. As an example, it is possible to request the reading of a new fixed-length sequence of bits when the number of bits present in a connected bit buffer decreases below a predetermined threshold, and different thresholds can be used for different bit buffers. The encoder must ensure that the fixed-length sequences of bins are input to the bit stream in the same order, and they are read from the bit stream on the decoder side. This alternating arrangement of fixed-length sequences can also be coupled with low-latency control similar to that described above. Preferred embodiments for the alternating arrangement of fixed-length sequences of bits are described below. For further details regarding the latter alternating arrangement scheme, see WO2011 / 128268A1.
[0088] Following the description of embodiments where previously encoded data is used for video data compression, further embodiments for carrying out embodiments of the present invention demonstrate particularly effective implementations with respect to a good trade-off between compression ratio and reference tables and computational overhead. In particular, the following embodiments entropy encode bit streams individually and enable the use of computationally uncomplicated variable-length codes to effectively cover the probabilistic evaluation portion. In the embodiments described below, symbols are binary, and the VLC codes presented below are, for example, R that extends between [0;0.5]. LPS by This effectively covers probability evaluations expressed in this way.
[0089] In particular, the embodiments outlined below each show individual entropy encoders 310 and decoders 322, as shown in Figures 7-17. They are suitable for encoding bins, i.e., binary symbols, as this occurs in image or video compression applications. Thus, these embodiments can be applied to image or video encoding, where such binary symbols are separated into one or more bins 307 to be encoded and bit streams 324 to be decoded, and each such bin stream can be considered an implementation of the Bernoulli process. The embodiments described below use one or more different so-called variable-to-variable codes (v2v codes) to encode the bin streams. A v2v code can be thought of as two prefix codes having the same number of codewords: a first prefix code and a second prefix code. Each codeword of the first prefix code is associated with one codeword of the second prefix code. According to the embodiments outlined below, at least some of the encoders 310 and decoders 322 operate as follows: Whenever a codeword of the first prefix is read from buffer 308 to encode a specific sequence in bin 307, the corresponding codeword of the second prefix is written to bitstream 312. A similar procedure is used to decode such bitstreams, but replaced with the first and second prefixes. That is, whenever a codeword of the second prefix is read from each bitstream 324 to decode bitstream 324, the corresponding codeword of the first prefix is written to buffer 320.
[0090] Conveniently, the codes described below do not require a reference table. The codes are executable in the form of a finite state machine. The v2v-codes that appear here can be generated by simple structural rules that do not require storing a large table for the codewords. Instead, simple algorithms can be used to perform encoding or decoding. Three structural rules are described below, two of which can be parameterized. They cover different or disparate portions of the aforementioned probability intervals and are therefore particularly advantageous when used in parallel (each for different encoder / decoder 11 and 22) or together, such as two of them for all three codes. For the structural rules described below, it is possible to design a set of v2v-codes such that one of the codes works well with respect to excessive code length for a Bernoulli process with any probability p.
[0091] As described above, the encoding and decoding of streams 312 and 324 can also be performed for each stream individually or in an alternating arrangement. However, this is not specific to the types of v2v codes shown, and therefore, the encoding and decoding of specific codewords are described only for each of the following three structural rules. However, all of the above embodiments relating to the alternating solution, which is emphasized, can be coupled to the codes or encoders and decoders 310 and 322 described, respectively.
[0092] Structural Rule 1: "Unary bin pipe" code or encoder / decoder 310 and 322
[0093] The unary bin-pipe code (PIPE = probability interval partitioning entropy) is the so-called "bin-pipe" code, that is, the individual bits · A special version of the code suitable for encoding either Stream 12 or 24, which transmits binary symbolic statistics data, each belonging to a specific probability subinterval of the aforementioned probability range [0;0.5]. The structure of the binpipe code is first described. The binpipe code can also consist of any prefix code having at least three codewords. To form a v2v code, prefixes are used as the first and second codes. A code is used, but it is swapped with two codewords of a second prefix code. This means that, except for the two codewords, the bin is written without being converted into a bit stream. For this technique, only one prefix code is required to be stored with information that has been swapped with two codewords, reducing memory consumption. Note that swapping codewords of different lengths makes sense, because otherwise the bit stream would have the same length as the bin stream (a negligible effect that can occur at the end of the bin stream).
[0094] Due to this structural rule, a notable characteristic of binpipe codes is that when the first and second prefix codes are swapped (while the codeword mapping is preserved), the resulting v2v-code is identical to the original v2v-code. Therefore, the coding and decoding algorithms are identical for binpipe codes.
[0095] A unary binpipe code is composed of special prefix codes. These special prefix codes are created as follows: Firstly, a prefix code consisting of n codewords is generated starting with "01", "001", "0001", ..., until a codeword is created, where n is a parameter of the unary binpipe code. A sequence of 1s is removed from the longest codeword. This corresponds to a shortened unary code (but without the codeword "0"). Then, n-1 unary codewords are generated starting with "10", "110", "1110", ..., until an n-1 codeword is created. A sequence of 0s is removed from the longest of these codewords. The combined set of these two prefix codes is used as input to generate a unary binpipe code. The two codewords to be exchanged are one consisting only of 0s and one consisting only of 1s.
[0096] Example for n=4: Nr 1st 2nd 1 0000 111 2 0001 0001 3 001 001 4 01 01 5 10 10 6 110 110 7 111 0000
[0097] Structural Rule 2: "Unary to rice" code and Unary to rice encoder / decoder 10 and 22:
[0098] The Unary to Rice code uses abbreviated unary codes as the first code. That is, unary codewords start with "1", "01", "001", ... and then 2 nCodewords are generated until a +1 codeword is produced, and a sequence of 1s is removed from the longest codeword. n is a parameter of the unary to rice code. The second prefix code is constructed from the codewords of the first prefix code as follows: The codeword "1" is assigned to the first codeword that consists only of 0s. All other codewords are constructed from a concatenation of codewords "0" that have an n-bit binary representation of the number of 0s corresponding to the first prefix code.
[0099] Examples for n=3: Nr 1st 2nd 1 1 0000 2 01 0001 3 001 0010 4 0001 0011 5 00001 0100 6 000001 0101 7 0000001 0110 8 00000001 0111 9 000000001 1 Note that this is equivalent to mapping an infinite unary code to a rice code with rice parameters.
[0100] Structural Rule 3: Three-bin code The three-bin code is given as follows: Nr 1st 2nd 1 000 0 2 001 100 3 010 101 4 100 110 5 110 11100 6 101 11101 7 011 11110 8 111 11111
[0101] The first code (symbol sequence) is of fixed length (always 3 bins), and the codeword is selected by an ascending field number of 1s.
[0102] An efficient embodiment of the three-bin code is described below. Encoders and decoders for the three-bin code can be made without a storage table as follows.
[0103] In the encoder (one of the 10), three bins are read from the bin stream (i.e., 7). If these three bins contain exactly one 1, the codeword "1" is written to the bit stream, followed by two bins consisting of a binary representation of one position (starting from the right with 00). If the three bins contain exactly one 0, the codeword "111" is written to the bit stream, followed by two bins consisting of a binary representation of one position (starting from the right with 00). The remaining codewords "000" and "111" are mapped to "0" and "11111", respectively.
[0104] In one of the decoders (22), one bin or bit is read from its respective bit stream 24. If it is equal to "0", the codeword "000" is decoded into bin stream 21. If it is equal to "1", two more bins are read from bit stream 24. If these two bits are not equal to "11", they are interpreted as a binary representation of a number, and the two 0s and one 1 are decoded into the bit stream so that the position of the 1 is determined by the number. If the two bits are equal to "11", two more bits are read and interpreted as a binary representation of a number. If this number is less than 3, two 1s and one 0 are decoded so that the number determines the position of the 0. If it is equal to 3, "111" is decoded into the bin stream.
[0105] The efficient implementation of unary binpipe codes is described below. Encoders and decoders for unary binpipe codes can be efficiently constructed using counters. Due to the structure of binpipe codes, encoding and decoding of binpipe codes is easy to perform.
[0106] In one of the encoders (10), if the first bin of the codeword is equal to "0", the bins are processed until a "1" occurs or until n zeros are read (including the first "0" of the codeword). If a "1" occurs, the read bins are written to the immutable bit stream. Otherwise (i.e., n zeros are read), n-1 ones are written to the bit stream. If the first bin of the codeword is equal to "1", the bins are processed until a "0" occurs or until n-1 ones are read (including the first "1" of the codeword). If a "0" occurs, the read bins are written to the immutable bit stream. Otherwise (i.e., n-1 ones are read), n zeros are written to the bit stream.
[0107] In the decoder (one of the 322s), since this is the same bin-pipe code as described above, the same algorithm is used as with the encoder.
[0108] An efficient implementation from unary to rice code is described below. Encoders and decoders for the code can be efficiently created using counters, as will be described below.
[0109] In the encoder (one of 310), until 1 occurs, or 2 n The number of 0s Until loaded, bins are read from the bin stream (i.e., 7). Numbers of 0 are counted. If the counted number is 2 n If equal to, the codeword "1" is written to the bitstream. It is written. Otherwise, "0" is written, followed by the binary representation of the counted number written in n bits.
[0110] In the decoder (one of the 322), 1 bit is read. If it is equal to "1", then 2 n The number of zeros is decoded into a binstring. If it is equal to "0", then n More bits are read and interpreted as a binary representation of the number. This number of zeros is decoded into a binstream, followed by "1".
[0111] TIFF2026143473000006.tif113169
[0112] TIFF2026143473000007.tif108169
[0113] TIFF2026143473000008.tif67169
[0114] TIFF2026143473000009.tif51169
[0115] TIFF2026143473000010.tif92169
[0116] Furthermore, one of the encoder's entropies, when converting the symbols sent to a predetermined entropy encoder into their respective bit streams, determines whether (1) if the triplet consists of a, the predetermined entropy encoder is configured to write a codeword (c) to each bit stream, or (2) if the triplet consists of a single b, each bit stream The system is configured to check whether a given entropy encoder is configured to write a codeword with (d) as a prefix and a 2-bit representation of the position of b in the triplet as a suffix, (3) if the triplet is indeed composed of one a, then for each bit stream, the given entropy encoder is configured to write a codeword with (d) as a prefix and a sequence of first 2-bit words that are not elements of the first set and a 2-bit representation of the position a in the triplet as a suffix, or (4) if the triplet is composed of b, then the given entropy encoder is configured to write a codeword with (d) as a prefix and a sequence of second 2-bit words that are not elements of the first set and a sequence of first 2-bit words that are not elements of the second set as a suffix for each bit stream.
[0117] TIFF2026143473000011.tif108169
[0118] Each of the first subset of entropy encoders, when converting each bit stream into a symbol, is configured such that (1) if the first bit is equal to 0{0,1}, each entropy encoder is configured such that (1.1) if b ≠ a and b0{0,1} occurs in the next n-1 bits following the first bit, each entropy decoder reproduces the sequence of symbols up to bit b such that the next bit of each bit stream is equal to the first bit, or (1.2) if b does not occur in the next n-1 bits following the first bit, each entropy decoder reproduces the sequence of symbols equal to (b,···,b)n-1. (2) Each bit stream is configured to examine the next bit to determine whether it is configured to reproduce the sequence of symbols up to symbol a, where the next bit of each bit stream is equal to the following first bit, or (2) each entropy decoder is configured to examine the next bit of each bit stream n This will reproduce a sequence of symbols equal to [the specified value]. To determine whether it is configured to look at the next bit of each bit stream, in order to determine whether it is configured to look at each bit... It is configured to examine the first bit of the stream.
[0119] TIFF2026143473000012.tif72169
[0120] TIFF2026143473000013.tif57169
[0121] TIFF2026143473000014.tif92169
[0122] TIFF2026143473000015.tif124169
[0123] Now, after describing the general concept of video encoding schemes, embodiments of the present invention will be described in relation to those embodiments. In other words, the embodiments outlined below can be carried out using the above scheme, and vice versa, and the above encoding schemes can be carried out using and effectively employed by the embodiments outlined below.
[0124] In the embodiments described with respect to Figures 7-9, the entropy encoders and decoders in Figures 1-6 were implemented according to the PIPE concept. One particular embodiment used arithmetic encoders / complexers 310 and 322 of a single stochastic state. As will be discussed later, according to another embodiment, components 306-310 and corresponding components 318-322 can be replaced with a general entropy coding engine. For example, as will be discussed further later, if we consider an arithmetic encoding engine, it simply manages one general state R and L and encodes all symbols into one common bit stream, thereby abandoning the advantageous aspects of the current PIPE concept with respect to parallel processing, but avoiding the need for alternating partial bit streams. In this case, the number of stochastic states whose context probabilities are estimated by updates (e.g., lookups into a table) may be higher than the number of stochastic states on which the probability interval subdivisions are performed. That is, similar to quantizing the probability interval width values before indexing the table Rtab, the stochastic state index can be quantized. The foregoing description for possible embodiments for one encoder / decoder 310 and 322 can thus be extended to embodiments of realizations of entropy encoder / decoder 318-322 / 306-310, such as context-adaptive binary coding / decoding edging.
[0125] More precisely, according to the embodiment, the entropy encoder attached to the output of the parameter assigner (which here acts as a context assigner) operates as follows: It is possible.
[0126] TIFF2026143473000016.tif134170
[0127] Similarly, an entropy decoder attached to the output of a parameter assigner (which here acts as a context assigner) can operate as follows:
[0128] TIFF2026143473000017.tif167169
[0129] As described above, assignor 4 assigns pState_current[bin] to each bin. Association is based on context selection; that is, assignor 4 can select a context using the context index ctxIdx which has the respective pState_current associated with it. Probabilistic updates are performed at each time, and the established pState_current[bin] is applied to the current bin. Updates to the probabilistic state pState_current[bin] are performed according to the value of the encoded bit.
[0130] TIFF2026143473000018.tif26162
[0131] If one or more contexts are provided, the matching is done contextually, i.e., pState_current[ctxIdx] is used for encoding, and Then, it is updated using the current bin values (which are either encoded or decoded).
[0132] As will be outlined in more detail below, according to the embodiments described herein, the encoder and decoder can be optionally configured to operate in different modes, namely, low complexity (LC) and high efficiency (HE) modes. This is primarily illustrated with respect to PIPE coding below (and refers to LC and HE PIPE modes), but the detailed explanation of complexity extensions is readily transferable to other embodiments of entropy encoding / decoding engines, such as the embodiment using a single common context adaptive arithmetic encoder / decoder.
[0133] According to the embodiments outlined below, both entropy coding modes can be shared. • Same syntax and behavior (for syntax element sequences 301 and 327, respectively) • The same binarization scheme for all syntactic elements (as currently specified for CABAC) (i.e., the binarizer can operate regardless of the mode in which it is activated). • Use of the same PIPE code (i.e., the bin encoder / decoder can operate regardless of the mode in which it is activated) • Use of 8-bit probability model initialization (instead of 16-bit initialization as currently specified for CABAC)
[0134] Generally speaking, LC-PIPE differs from HE-PIPE in terms of processing complexity (for example, the complexity of selecting PIPE route 312 for each bin).
[0135] For example, the LC mode can operate under the following constraints: For each bin (binIdx), there is indeed one probabilistic model (i.e., one ctxIdx). That is, context selection / fitting cannot be provided in the LC PIPE. Certain syntactic elements, such as those used for residual coding, are coded with context, as will be further outlined below. Furthermore, all probabilistic models may be non-adaptable, i.e., all models can be initialized at the beginning of each slice with appropriate model probabilities (depending on the choice of slice type and slice QP) and kept fixed throughout the processing of the slice. For example, for both context modeling and coding, only eight different model probabilities corresponding to eight different PIPE codes 310 / 322 may be supported. Certain syntactic elements for residual coding, i.e., significance_coeff_flag and coeff_abs_level_greaterX (X=1,2), whose operation will be outlined below, for example. At least four groups of syntactic elements can be assigned to a probabilistic model so that they are coded / decoded with the same probability. Compared to CAVLC, LC-PIPE mode achieves substantially the same RD performance and throughput.
[0136] HE-PIPE can be constructed to be conceptually similar to H.264's CABAC, with the following differences: Binary arithmetic coding (BAC) is replaced by PIPE coding (as in the LC-PIPE case). Each probabilistic model, i.e., each ctxIdx, can be represented by pipeIdx and refineIdx, where pipeIdx, having values ranging from 0 to 7, represents the model probability of eight different PIPE codes. This change affects only the internal representation of the state, and not the movement of the state machine (i.e., the probabilistic evaluation) itself. As outlined in more detail below, the initialization of the probabilistic model can use an 8-bit initialization value, as described above. Syntactic elements coeff_abs_level_greaterX(X=1,2), coeff_abs The backward scanning of _level_minus3 and coeff_sign_flag (whose operation will become clear from the following considerations) can be performed along the same scanning path as the forward scan (for example, used in importance map coding). Context derivation for coding coeff_abs_level_greaterX (X=1,2) can also be simplified. Compared to CABAC, the proposed HE-PIPE achieves substantially the same RD performance with better throughput.
[0137] It is easy to see that the mode just mentioned can be immediately achieved, for example, by rendering the aforementioned context-adaptive binary coding / decoding engine to operate in a different mode.
[0138] Thus, according to an embodiment of the first aspect of the present invention, a decoder for decoding a data stream can be fabricated as shown in Figure 11. The decoder is for decoding a data stream 401, such as an alternating bit stream 340, in which media data, such as video data, is encoded. The decoder includes a mode switch 400 configured to operate either a low-complexity mode or a high-efficiency mode depending on the data stream 401. For this purpose, the data stream 401 includes syntactic elements, such as binary syntactic elements, which have a binary value of 1 when the low-complexity mode is in operation and a binary value of 0 when the high-efficiency mode is in operation. Obviously, the association between binary values and encoding modes can be switched, and non-binary syntactic elements having two or more possible values can be used as well. Since the actual choice between the two modes is not yet apparent prior to the acceptance of each syntactic element, this syntactic element can be included within some of the main headers of data stream 401, either encoded, for example, with a fixed-probability evaluation or a probabilistic model, or written directly to data stream 401 using bypass mode.
[0139] Furthermore, the decoder in Figure 11 includes multiple entropy decoders 322, each configured to convert the codeword of the data stream 401 into a partial sequence of symbols 321. As mentioned above, the deinterleaver 404 is connected on the one hand between the inputs of the entropy decoders 322 and on the other hand to the input of the decoder in Figure 11 to which the data stream 401 is applied. Furthermore, as already mentioned above, each of the entropy decoders 322 is associated with its own probability interval, and in the case of entropy decoders 322 that deal with MPS and LPS rather than absolute symbol values, the probability intervals of the various entropy decoders range from 0 to 1 - or 0 to 0.5 together to cover all probability intervals. Details regarding this issue have been mentioned above. Later, it will be assumed that the number of decoders 322 with PIPE indices assigned to each decoder is 8, but any other number is possible. Furthermore, one of these encoders, which has pipe_id 0 as an example below, is optimized for bins with a similar probability of having a certain statistical value, i.e., their bin values are assumed to be equally likely to be 1 and 0. Thus, the decoder can simply pass the bins. Each encoder 310 operates in the same way. Any bin operation by the most likely bin value, valMPS, can be set aside by selectors 402 and 502, respectively. In other words, the entropy of each partial stream is already optimal.
[0140] Furthermore, the decoder in Figure 11 includes a selector 402 configured to look up each symbol of the symbol sequence 326 from one of a selection of entropy decoders 322. As described above, the selector 402 may be separated into a parameter assigner 316 and a selector 318. The desymbolizer 314 is configured to desymbolize the symbol sequence 326 in order to obtain a sequence of syntactic elements 327. The regenerator 404 is configured The media data 405 is configured to reproduce the sequence of statement elements 327. The selector 402 is configured to make a selection depending on whether it is operating in low complexity mode or high efficiency mode, as indicated by the arrow 406.
[0141] As already mentioned above, the regenerator 404 acts on the constant syntax and behavior of syntactic elements, i.e., it may be part of a block-based video decoder that is a precursor fixed in relation to mode selection by the mode switch 400. That is, the structure of the regenerator 404 is not plagued by mode switchability. To be more precise, the regenerator 404 does not increase implementation overhead due to the mode switchability indicated by the mode switch 400, and at least functional and predictive data remains with respect to the residual data, which is independent of the mode selected by the switch 400. However, the same is true with respect to the entropy decoder 322. All these decoders 322 are reused in both modes, and therefore there is no additional implementation overhead despite the decoders in Figure 11 being compatible with both modes, low complexity and high efficiency modes.
[0142] As an additional aspect, it should be noted that the decoder in Figure 11 is not only capable of acting on a self-sufficient data stream of one mode or another. Rather, the decoder in Figure 11 is configured similarly to the data stream 401 so that it is possible to switch between both modes between one element of media data, such as between video or some audio elements, using a feedback channel from the decoder to the encoder to control the encoding complexity on the decoding side in response to external or environmental conditions, such as the state of the battery, and consequently to perform fixed-loop control of model selection.
[0143] Thus, in the case of the selected LC mode or the selected HE mode, the decoder in Figure 11 operates similarly in both cases. The reproducer 404 requests a current syntactic element of a given syntactic element type by performing a reproducer using the syntactic elements and processing or adhering to some syntactic structure rules. The desymbolizer 314 requests multiple bins to produce a valid binarization for the syntactic element requested by the reproducer 404. Obviously, in the case of a binary alphabet, the binarization performed by the desymbolizer 314 is reduced to simply passing each bin / symbol 326 to the reproducer 404 as the currently requested binary syntactic element.
[0144] However, each selector 402 acts according to the mode selected by the mode switch 400. The modes of operation of selector 402 tend to be more complex in the high-efficiency mode and less complex in the low-complexity mode. Furthermore, the following explanation shows that the modes of operation of selector 402 in the less complex mode also tend to decrease the rate at which selector 402 changes its selection within the entropy decoder 322 when searching for a continuous symbol from the entropy decoder 322. In other words, in the low-complexity mode, there is an increased probability that a continuous symbol will be immediately found in the same entropy decoder among multiple entropy decoders 322. This, in turn, allows for faster retrieval of symbols from the entropy decoder 322. In high-efficiency mode, the mode of operation of selector 402 then tends to lead to the selection of an entropy decoder 322 that more closely matches the actual symbolic statistics of the symbol currently being searched by selector 402, with the probability interval associated with each selected entropy decoder 322 being more closely aligned, thereby resulting in a better compression ratio on the encoding side when generating each data stream following high-efficiency mode.
[0145] For example, the different responses of selector 402 in both modes can be understood as follows: For example, selector 402, for a given symbol, high efficiency mode The selector 402 can be configured to perform a selection among multiple entropy decoders 322 depending on the symbols searched before the symbol sequence 326 when the D is activated, and independently of the previously searched symbols in the symbol sequence when the low complexity mode is activated. The dependence of the symbol sequence 326 on previously searched symbols can be due to contextual adaptation and / or probabilistic adaptation. Both adaptations can be switched off during the low complexity mode of the selector 402.
[0146] In a further embodiment, the data stream 401 can be constructed into continuous parts such as slices, frames, groups of images, or sequences of frames, where each symbol in the sequence of symbols is associated with one of several symbol types. In this case, the selector 402 is configured to change its selection for symbols of a given symbol type in the current part, depending on previously retrieved symbols in the sequence of symbols of that given symbol type when the high-efficiency mode is activated, and to keep the selection constant within the current part when the low-complexity mode is activated. That is, the selector 402 can change the selection in the entropy decoder 322 to a given symbol type, but these changes are limited to occurring between transitions between continuous parts. This measurement reduces coding complexity over most of the time, and the evaluation of actual symbol statistics is limited to rare time instances.
[0147] Furthermore, each symbol in the symbol sequence 326 is associated with one of several symbol types, and the selector 402 is configured to select one of several contexts for a given symbol of a given symbol type, depending on the previously searched symbols in the symbol sequence 326, and to perform a selection between the entropy decoder 322 depending on the selected context and the associated probabilistic model when the high-efficiency mode is activated, and when the low-complexity mode is activated, to select one of several contexts depending on the previously searched symbols in the symbol sequence 326, and to perform a selection between the entropy decoder 322 depending on the selected context and the associated probabilistic model, moving away from the selected context constant.
[0148] Alternatively, instead of completely suppressing stochastic matching, selector 402 can simply reduce the update rate of stochastic matching for LC mode in relation to HE mode.
[0149] Furthermore, possible LC-pipe-specific aspects, i.e., aspects of the LC mode, can be described in other words as follows: In particular, a non-adaptive probabilistic model can be used in LC mode. The non-adaptive probabilistic model can be hardcoded, i.e., a constant overall probability or either of these probabilities can be kept fixed throughout the entire processing of slices, and thus can be set according to the slice type and QP, i.e., the quality parameter signaled in the data stream 401 for each slice. By assuming that consecutive bins assigned to the same context follow the fixed probabilistic model, it is possible to decode some of those bins in one step, since they are encoded using the same pipe code, i.e., using the same entropy decoder, and the probabilistic update after each decoded bin is omitted. Omitting the probabilistic update saves operation during the encoding and decoding process, leading to reduced complexity and significant simplification of hardware design.
[0150] The non-adaptive constraints are all or some selected probabilistic models, such that a certain number of bins are encoded / decoded using this model and then probabilistic updates are allowed. This can be easily done. An appropriate update interval allows for probabilistic fitting, and the ability to quickly decode several bins is gained.
[0151] A more detailed description of possible general and complexity-measurable embodiments of LC-pipe and HE-pipe is given below. In particular, embodiments used for LC-pipe mode and HE-pipe mode in the same manner, or in a complexity-measurable manner, are described below. The complexity-measurable method shows that LC-case is derived from HE-case by removing certain parts or replacing them with something less complex. However, before dealing with that, it should be noted that the embodiment in Figure 11 is readily transferable to the context-adaptive binary coding / decoding embodiments described above: the selector 402 and entropy decoder 322 collate to become a context-adaptive binary decoder that directly receives the data stream 401 and selects the context for the bins now to be derived from the data stream. This is true in particular for context-adaptive and / or probabilistic adaptive. During low-complexity modes, both functions / adaptives can be switched off or designed to be more relaxed.
[0152] For example, when implementing the embodiment in Figure 11, the pipe entropy coding stage, including the entropy decoder 322, can use eight organized variable-to-variable codes, i.e., the entropy decoder 322 can be of the v2v type described above. The PIPE coding concept using organized v2v codes is simplified by limiting the number of v2v codes. In the case of a context-adaptive binary decoder, it can manage the same stochastic state for different contexts and can use it—or its decoded version. The mapping of PIPEids or stochastic indices for lookups of CABAC or stochastic model states—i.e., states used for stochastic updates—to Rtab is as illustrated in Table 5.
[0153] [Table 5]
[0154] This modified coding scheme can be used as the basis for a complexity-measurable video coding approach. When performing stochastic mode fitting, the selector 402 or the context-adaptive binary decoder selects the PIPE decoder 322, respectively, i.e., the pipe index used based on the stochastic state index. The probability index is then retrieved into Rtab and updated according to the currently decoded symbol using the mapping shown in Table 5—for example, via context—in relation to the currently decoded symbol—here, ranging from 0 to 62 as an example—based on the probability state index, and respectively, using a specific table that shows the transition values shown in the next probability state index to visit for MPS and LPS.
[0155] However, any entropy coding setup can be used, and the techniques in this document can also be employed with small fits.
[0156] The above explanation in Figure 11 is more typically related to syntactic elements and syntactic element types. The following describes variable complexity coding at the conversion coefficient level.
[0157] For example, the reproducer 404 is configured to reproduce the conversion coefficient levels 202 based on a sequence of syntactic elements independent of the operating high-efficiency or low-complexity mode, such that a portion of the sequence 327 of syntactic elements defines an importance map in which importance map syntactic elements indicate the location of non-zero conversion coefficient levels within the conversion block 200, in a non-alternating manner, followed by level syntactic elements defining non-zero conversion coefficient levels. In particular, the following elements may be included: terminal position syntactic elements (last_significant_pos_x, last_significant_pos_y) indicating the position of the last non-zero conversion coefficient level within the conversion block; a first syntactic element (coeff_significant_flag) that defines a significance map and indicates whether the conversion coefficient level at each position is non-zero for each position along a one-dimensional path (274) that leads from the DC position to the position of the last non-zero conversion coefficient level within the conversion block (200); a second syntactic element (coeff_abs_greater1) that indicates whether the conversion coefficient level at each position is greater than the first binary syntactic element for each position along the one-dimensional path (274) where a non-zero conversion coefficient level is located; and a third syntactic element (coeff_abs_greater2, coeff_abs_minus3) that, according to the first binary syntactic element, reveals the numerical value that the respective conversion coefficient level at each position exceeds for each position along the one-dimensional path where a greater conversion coefficient level is located.
[0158] The order between end-position syntactic elements, the first, second, and third syntactic elements are common to both high-efficiency and low-complexity modes, and the selector 402 can be configured to perform a selection in the entropy decoder 322 for symbols from which the desymbolizer 314 obtains the end-position syntactic elements, the first syntactic element, the second syntactic element, and / or the third syntactic element, depending on whether the low-complexity or high-efficiency mode is operating.
[0159] In particular, in low complexity mode, so that selection is constant through consecutive subparts of a subsequence, the selector 402 is configured to select one of several contexts for each symbol of a given symbol type in the subsequence of symbols from which the desymbolizer 314 obtains the first and second syntactic elements, depending on previously searched symbols of the given symbol type in the subsequence of symbols, and to perform the selection in a partially constant manner depending on the probabilistic model associated with the selected context when high efficiency mode is activated. As described above, subparts can be measured in terms of the number of positions to which each subpart extends when measured along the one-dimensional path 274, or in terms of the number of syntactic elements of each type already encoded by the current context. That is, for example, the binary syntactic elements coeff_significant_flag, coeff_abs_greater1 and coeff_abs_greater2 are selected in HE mode This is an encoding context that adapts to selecting decoder 322 based on a probabilistic model of the text. Probabilistic fitting is used similarly. In LC mode, there are also different contexts used for each of the binary syntactic elements coeff_significant_flag, coeff_abs_greater1, and coeff_abs_greater2. However, for each of these syntactic elements, the context is kept unchanged for the first part along path 274, relating to simply changing the context in the transition to the next, immediately following part along path 274. For example, each part is defined by the length of 4, 8, or 16 positions in block 200, independently of whether there is a syntactic element for each position. For example, coeff_abs_greater1 and coeff_abs_greater2 exist only for significant positions, i.e., positions where coeff_significant_flag=1. Or, independently of whether each of the resulting parts is an extension beyond a further number of block positions, it is defined by the length of 4, 8, or 16 syntactic elements. For example, coeff_abs_greater1 and coeff_abs_greater2 exist only for their important positions, and thus each of the four syntactic element parts can extend its position by more than four blocks for positions along path 274 where such syntactic elements are not sent, such as coeff_abs_greater1 and coeff_abs_greater2, because the level of each of these positions is zero.
[0160] Selector 402 is configured to select one of a plurality of contexts for each symbol of a given symbol type in a subsequence of symbols from which the desymbolizer obtains a first syntactic element and a second syntactic element, depending on a plurality of previously searched symbols of the given symbol type within the subsequence of symbols, which have a given symbol value and belong to the same subpart, or are a plurality of previously searched symbols of the given symbol type within a sequence of symbols that belong to the same subpart. The first variation is applicable to coeff_abs_greater1, and the second variation is applicable to coeff_abs_greater2 according to the particular embodiment described above.
[0161] Furthermore, the third syntactic element to be examined includes, according to the first syntactic element, an integer syntactic element, namely coeff_abs_minus3, for each position in the one-dimensional path where a greater transformation coefficient level is located, and the amount by which the respective transformation coefficient level at each position exceeds it is an integer syntactic element. The desymbolizer 314 is configured to use a mapping function controlled by a control parameter that maps the domain of the symbol sequence word to the range of the integer syntactic element, and is configured to set the control parameter for each integer syntactic element according to the integer syntactic element of the previous third syntactic element when the high-efficiency mode is activated, and to perform the setting in a piecewise constant manner so that the setting is constant across successive sub-parts of the subsequence when the low-complexity mode is activated. The selector 402 is configured to select one of a predetermined entropy decoders (322) for the symbols of the symbol sequence word that are mapped to integer syntactic elements associated with the same probability distribution in both the high-efficiency mode and the low-complexity mode. That is, even the desymbolizer operates according to the mode selected by switch 400, as shown by the dotted line 407. Instead of setting control parameters to a fixed, segmental value, the desymbolizer 314 keeps the control parameters constant, for example, during the current slice, or while they remain constant overall over time.
[0162] Next, context modeling that allows for complexity measurement is described.
[0163] For example, for motion vector difference syntax elements, the derivation of context model index Evaluating the same syntactic element above and to the left is a common method, often used in HE cases. However, this evaluation requires more buffer memory and cannot directly encode the syntactic element. Furthermore, to achieve higher encoding performance, more readily available neighbors can be evaluated.
[0164] In a preferred embodiment, all context modeling stages evaluating syntactic elements in adjacent square or rectangular blocks or prediction units are determined to a single context model. This is equivalent to disabling the adaptation of the context model selection stage. For that preferred embodiment, context model selection based on the bin index of the bin string after binarization is not modified compared to the current design for CABAC. In another preferred embodiment, in addition to a fixed context model for syntactic elements adopting the evaluation of neighbors, context models are fixed for different bin indices. Note that the description does not include motion vector differences and binarization and context model selection for syntactic elements related to the coding of transformation coefficient levels.
[0165] In a preferred embodiment, only the leftmost adjacent unit is allowed to be evaluated. This leads to a reduced buffer in the processing chain, as the last block or coded unit line no longer needs to be stored. In another preferred embodiment, only adjacent units residing in the same coding device are evaluated.
[0166] In a preferred embodiment, all available neighbors are evaluated. For example, in addition to the top and left neighbors, the top-left, top-right, and bottom-left neighbors are evaluated if effective.
[0167] In other words, the selector 402 in Figure 11 can be configured to use previously retrieved symbols of a sequence of symbols relating to a higher number of different adjacent blocks of media data when the high-efficiency mode is activated, to select one of several contexts for a given symbol relating to a given block of media data and to perform a selection in the entropy decoder 322 according to the probabilistic model associated with the selected context. That is, adjacent blocks can be close in the temporal and / or spatial domain. Spatially adjacent blocks can be seen, for example, in Figures 1-3. Then, as just described, the selector 402 may respond to mode selection by the mode switch 400 to perform contact fitting based on previously retrieved symbols or syntactic elements relating to a higher number of adjacent blocks in the case of the HE mode, compared to the LC mode, which thereby reduces memory overhead.
[0168] Next, a complex coding method with reduced motion vector difference, according to an embodiment, is described.
[0169] In the H.264 / AVC video codec standard, motion vectors associated with macroblocks are transmitted by sending a signal of the difference (motion vector difference - mvd) between the motion vector of the current macroblock and the intermediate motion vector predictor. When CABAC is used as the entropy encoder, mvd is encoded as follows: The integer value of mvd is divided into an absolute part and a signed part. The absolute part is binarized using a combination of a shortened unary cubic Exp-Golomb, called a prefix and a suffix. The bins associated with the Exp-Golomb are encoded in bypass mode, i.e., with a fixed probability of 0.5 in CABAC, while the bins associated with the shortened unary binarization are encoded using a context model. The unary binarization works as follows: If the absolute integer value of mvd is n, the resulting bin string consists of n "1"s and one following "0". For example, if n=4, the bin string is "11110". In the case of an omitted unary, a restriction exists, and if the value exceeds this restriction, the bin string It consists of n+1 instances of "1". In the case of mvd, the restriction is equal to 9. That is, it is equal to 9. For absolute mvd values of 9 or greater, the result is nine "1"s, and the bin string consists of a prefix and suffix with Exp-Golomb binarization. Context modeling for omitted unary parts is performed as follows: For the first bin of the bin string, the absolute mvd values from the upper and left adjacent macroblocks are required, if available (if unavailable, the value is assumed to be 0). If the sum of a particular component (horizontally or vertically) is greater than 2, the second context model is selected; if the absolute sum is greater than 32, the third context model is selected; otherwise (the absolute sum is less than 3), the first context model is selected. Furthermore, the context model differs for each component. For the second bin of the bin string, the fourth context model is used, and the fifth context model is used for the remaining bins of the unary part. If the absolute mvd is equal to or greater than 9, for example, all bins of the omitted unary part are equal to "1", and the difference between the absolute mvd value and 9 is encoded in bypass mode with cubic Exp-Golomb binarization. In the final stage, the MVD code is encoded in bypass mode.
[0170] The latest coding techniques when using CABAC as the entropy encoder are high This is defined in the current Test Model (HM) of the Efficiency Video Coding (HEVC) project. In HEVC, the block size is variable, and the shape specified by the motion vector is called a Prediction Unit (PU). The PU sizes of the top and left neighbors can have different shapes and sizes than the current PU. Therefore, whenever relevant, the definitions of top and left neighbors refer to the left neighbors of the top and upper left corners of the current PU. For the coding itself, only the induction method for the first bin can vary according to the embodiment. Instead of evaluating the absolute sum of MVs from neighbors, each neighbor can be evaluated separately. If the absolute MV of a neighbor is available and greater than 16, the context model index increases to the same number as the context model for the first bin, while the coding of the remaining absolute MVD levels and codes becomes exactly the same as H.264 / AVC.
[0171] In the above outlined techniques for encoding mvd, bins 9 or less must be encoded using the context model, while the remaining values of the mvd can be encoded in low-complexity bypass mode along with the coding information. This current embodiment describes a technique for reducing the number of bins encoded using the context model, resulting in an increased number of bypasses and a reduction in the number of context models required for encoding the mvd. For this purpose, the cutoff value is reduced from 9 to 1 or 2. That is, only the first bin specifying whether the absolute mvd is greater than zero is encoded using the context model, or the first and second bins specifying whether the absolute mvd is greater than zero and they are encoded using the context model, while the remaining values are encoded in bypass mode and / or using VLC code. All bins obtained by binarizing using VLC code without using unary or abbreviated unary coding are encoded using low-complexity bypass mode. In the case of PIPE, direct insertion into and from the bitstream is possible. Furthermore, if any, a different definition can be used next to the first bin, either above or to the left, to elicit a better contextual model selection.
[0172] In a preferred embodiment, the Exp-Golomb code is used to binarize the remaining portion of the absolute MVD component. For this purpose, the order of the Exp-Golomb codes is variable. The order of the Exp-Golomb codes is derived as follows: After the context model for the first bin, and therefore the index of that context model, is derived and encoded, the index is used to find the Exp-Golomb binarized portion. It is used as an instruction. In this preferred embodiment, the context model for the first bin is in the range of 1-3, resulting in indices 0-2 which are used as instructions in the Exp-Golomb code. This preferred embodiment can be used for the HE case.
[0173] In a variation of the technique outlined above, which uses two sets of five contexts for absolute MVD coding, 14 context models (7 for each component) can similarly be used to code nine unary code binarized bins. For example, the first and second bins of the unary part can be coded with the four different contexts described above, while the fifth context can be used for the third bin, the sixth context can be used for the fourth bin, and the fifth through ninth bins are coded using the seventh context. Thus, in this case, a total of 14 contexts are required, and only the remaining values can be coded in low-complexity bypass mode. Techniques to reduce the number of bins coded in context models, which result in an increased number of bypasses and a reduction in the number of context models required for MVD coding, involve reducing the cutoff value, for example, from 9 to 1 or 2. This indicates that only the first bin, which specifies whether its absolute MVD is greater than zero, is encoded using the context model, or the first and second bins, which specify whether their absolute MVD is greater than zero, are encoded using their respective context models, while the remaining values are encoded by VLC code. All bins obtained from binarization using VLC code are encoded using the low-complexity bypass mode. In the case of PIPE, direct insertion into and from the bit stream is possible. Furthermore, the given embodiment uses the other definitions of the upper and left neighbors to derive a better context model selection for the first bin. In addition to this, the context modeling is modified to some extent so that the number of context models required for the first, or the first and second bins, is reduced, leading to further memory reduction. Also, evaluation of neighbors as described above can be disabled, resulting in a reduction in the line buffer / memory required for accumulating neighbor mvd values. Finally, the encoding order of the components can be split in a way that allows the encoding of the prefix bins to be followed by the encoding of the bypass bins for both components (i.e., bins encoded with the context model).
[0174] In a preferred embodiment, the Exp-Golomb code is used to binarize the remaining portion of the absolute mvd component. For this purpose, the order of the Exp-Golomb code is variable. The order of the Exp-Golomb code can be derived as follows: After the context model for the first bin, and therefore the index of that context model, is derived, the index is used as the instruction to request the Exp-Golomb binarization. In this preferred embodiment, the context model for the first bin is in the range of 1 to 3, resulting in indices 0 to 2 being used as the instruction for the Exp-Golomb code. This preferred embodiment can be used for the HE case, where the number of context models is reduced to 6. Also, in order to reduce the number of context models and thereby save memory, the horizontal and vertical components can share the same context models as in a further preferred embodiment. In that case, only three context models are required. Furthermore, only the leftmost neighbor can be considered for evaluation of further preferred embodiments of the invention. In this preferred embodiment, the thresholds do not need to be changed (for example, one threshold 16 yields an Exp-Golomb parameter of 0 or 1, and one threshold 32 yields an Exp-Golomb parameter of 0 or 2). This preferred embodiment saves the line buffer required for MVD storage. In another preferred embodiment, the thresholds are modified to be equal to 2 and 16. For that preferred embodiment, a total of three context models are required for MVD encoding, and the possible Exp-Golomb parameters are in the range of 0 to 2. In yet another preferred embodiment... The thresholds are equal to 16 and 32. Furthermore, the described embodiments are suitable for HE cases.
[0175] In a further preferred embodiment of the present invention, the cutoff value is reduced from 9 to 2. In this preferred embodiment, the first and second bins are encoded using context models. Context model selection for the first bin can be carried out in the highest-level or modified manner as described in the above preferred embodiments. For the second bin, another context model is selected, such as the highest-level technique. In another preferred embodiment, the context model for the second bin is selected by evaluating the left neighbor mvd. In this case, the context model index is the same with respect to the first bin, while the available context models are different from those for the first bin. In total, six context models are required (note the components that share context models). Also, the Exp-Golomb parameter may depend on the selected context model index for the first bin. In other preferred embodiments of the present invention, the Exp-Golomb parameter depends on the context model index for the second bin. The embodiments described in the present invention can be used for the HE case.
[0176] In a further preferred embodiment of the present invention, the context models for both bins are fixed and are not derived by evaluating the left or upper neighbor. For this preferred embodiment, the total number of context models is equal to 2. In a further preferred embodiment of the present invention, the first and second bins share the same context model. As a result, only one context model is required for the MVD encoding. In both preferred embodiments of the present invention, the Exp-Golomb parameter is fixed and equal to 1. The preferred embodiments described in the present invention are suitable for both HE and LC configurations.
[0177] In another preferred example, the order of the Exp-Golomb parts is derived from the context model index of the first bin, respectively. In this case, the absolute sum of the normal context models in H.264 / AVC is used to derive the instruction to find the Exp-Golomb parts. This preferred embodiment can be used for the HE case.
[0178] In another preferred embodiment, the order of the Exp-Golomb codes is fixed and set to 0. In yet another preferred embodiment, the order of the Exp-Golomb codes is fixed and set to 1. In a preferred embodiment, the order of the Exp-Golomb codes is fixed to 2. In a further embodiment, the order of the Exp-Golomb codes is fixed to 3. In a further embodiment, the order of the Exp-Golomb codes is fixed according to the shape and size of the current PU. The preferred embodiments shown can be used for the LC case. Note that the fixed order of the Exp-Golomb parts is considered in terms of the reduced number of bins encoded in the context model.
[0179] In a preferred embodiment, the neighbor is determined as follows: For the above PU, all PUs covering the current PU are considered, and the PU with the largest MV is used. This is also done for the left neighbor. All PUs covering the current PU are evaluated, and the PU with the largest MV is used. In another preferred example, the average absolute motion vector values from all PUs covering the top and left boundaries of the current PU are used to derive the first bin.
[0180] For the preferred embodiment shown, the encoding order can be changed as follows: mvd must be specified for the horizontal and vertical directions, one after the other (or vice versa). No. Thus, the two bin strings must be encoded. Mode switching for the entropy coding engine (i.e., switching between bypass and normal modes) allows for encoding the bins encoded in the context model for both components in the first stage, followed by the bins encoded in bypass mode in the second stage. Note that this is simply a rearrangement.
[0181] It should be noted that bins obtained from the binarization of a unary or abbreviated unary can also be represented by the binarization to an equivalent fixed length of one flag per bin index that specifies whether the value is greater than the current bin index. For example, the cutoff value for the binarization of an abbreviated unary mvd is set to 2, resulting in the codewords 0, 10, 11 for values 0, 1, 2. In the corresponding fixed-length binarization with one flag per bin index, one flag for bin index 0 (i.e., the first bin) specifies whether the absolute mvd value is greater than 0, and one flag for the second bin with bin index 1 specifies whether the absolute mvd value is greater than 1. This results in the same codewords 0, 10, 11 when the second flag is encoded only when the first flag is equal to 1.
[0182] Next, a measurable representation of the complexity of the internal state of a probabilistic model is described according to the examples.
[0183] In an HE-PIPE setup, the internal state of a probabilistic model is updated after encoding the bins that possess it. The updated state is retrieved by state transitions using table lookups that utilize the old state and the encoded bin values. In the case of CABAC, the probabilistic model can take 63 different states corresponding to the model probability in the interval (0.0, 0.5). Each of these states is used to realize two model probabilities. In addition to the probabilities assigned to the states, a minus 1 probability is also used, and a flag called valMps stores information on whether the probability or the minus 1 probability is used. This results in a total of 126 states. To use such a probabilistic model with the PIPE encoding concept, each of the 126 states needs to be mapped to one of the available PIPE encoders. In the current implementation of the PIPE encoder, this is done using a reference table. An example of such mapping is shown in Table 5.
[0184] The following examples illustrate how the internal states of a probabilistic model can be shown to avoid using a reference table to translate the internal states into PIPE indices. Simply some simple bit-masking operations are required to obtain the PIPE indices from the internal state variables of the probabilistic model. This novel complexity-measurable representation of the internal states of a probabilistic model is designed in two levels. For applications where low complexity operation is essential, only the first level is used. It lists only the pipe index and flag valMps used to encode or decode the relevant bins. For the PIPE entropy coding scheme described, the first level can be used to distinguish eight different model probabilities. Thus, the first level requires 3 bits for pipeIdx and one additional bit for the valMps flag. In the second level, each of the coarse probability ranges of the first level is refined into several smaller intervals to support probability presentation with higher resolution. This more detailed presentation allows for more accurate operation of the probabilistic estimator. In general, it is suitable for coding applications aiming toward high RD-performance. For example, an example of how this complexity criterion can be expressed for the internal state of a probabilistic model that uses PIPE is shown below.
[0185] [Table 6]
[0186] The first and second levels are stored in a single 8-bit memory. Four bits are required to store the first level—the index that determines the PIPE index with the MPS value on the most important bit—and another four bits are used to store the second level. To satisfy the response of the CABAC probability estimator, each PIPE index has a certain number of allowed refined indices depending on how many CABAC states are mapped to the PIPE index. For example, for the mapping in Table 5, the number of CABAC states per PIPE index is shown in Table 7.
[0187] [Table 7]
[0188] During the bin encoding or decoding process, the PIPE index and valMps can be directly accessed by using simple bit mask or bit shift operations. Low-complexity encoding processes require only 4 bits at the first level, while high-efficiency encoding processes can utilize an additional 4 bits at a second level to perform probabilistic model updates of the CABAC stochastic estimator. To implement this update, a state transition reference table can be designed that has the same transitions as the original table but uses a two-level representation of state complexity measurable. The original state transition table consists of 2 × 63 components. For each input state, it contains two output states. When using the complexity measurable representation, the size of the state transition table does not exceed 2 × 128 components, which is an acceptable increase in the table size. This increase depends on how many bits are used to represent the improved index and how many are used to accurately emulate the CABAC stochastic estimator's response, requiring 4 bits. However, different stochastic estimators can be used to work with a reduced set of CABAC states, such that only 8 states are allowed per pipe index. Therefore, memory consumption can be matched to a given level of complexity in the encoding process by adapting the number of bits used to represent the improved index. Compared to the internal state of model probabilities with a CABAC having 64 probability state indices, the use of table lookups to map model probabilities to specific PIPE codes is avoided, and no further transformation is required.
[0189] Next, a complexity-measurable context model, updated according to the examples, is described.
[0190] To update the context model, its stochastic state index can be updated based on one or more previously encoded bins. In an HE-PIPE setup, this update occurs after encoding or decoding each bin. Conversely, in an LC-PIPE setup, this update never occurs.
[0191] However, it is possible to update context models in a way that allows for complexity measurement. That is, the decision of whether or not to update a context model is based on various factors. For example, an encoding setup cannot be updated only for a specific context model, such as the context model for the syntactic element coeff_significant_flag, but can always be updated for all other context models.
[0192] In other words, the selector 402 is configured to perform a sorting within the entropy decoder 322 for each symbol of the given number of symbol types, according to the respective probabilistic model associated with each given symbol, such that the number of given symbol types is lower in the low complexity mode compared to the high efficiency mode.
[0193] Furthermore, criteria for controlling whether or not to update the context model could be, for example, the size of the bitstream packet, the number of bins decoded so far, or the update could only be performed after encoding a specific fixed or variable number of bins for the context model.
[0194] For this scheme of deciding whether or not to update the context model, a complexity-measurable context model can be implemented. It allows for increasing or decreasing the portion of the bitstream bin in which the context model update occurs. The more context model updates there are, the better the encoding efficiency and the higher the computational complexity. Thus, a complexity-measurable context model update can be provided by the described scheme.
[0195] In a preferred embodiment, the context model is updated for all syntactic element bins except for the syntactic elements coeff_significant_flag, coeff_abs_greater1, and coeff_abs_greater2.
[0196] In another preferred embodiment, the context model update is performed only for the bins coeff_significant_flag, coeff_abs_greater1, and coeff_abs_greater2.
[0197] In another preferred embodiment, context model updates are performed for all context models when the encoding or decoding of the slice begins. After a certain predetermined number of transformation blocks being processed, context model updates are unavailable for all context models until the end of the slice is reached.
[0198] For example, the selector 402 is configured to perform selection in the entropy decoder 322 according to a probabilistic model associated with a given symbol type, with or without updating the associated probabilistic model, such that the length of the learning phase of the sequence of symbols in which selection for symbols of a given symbol type is performed with updates is shorter in the low complexity mode than in the high efficiency mode.
[0199] A further preferred embodiment is identical to the previously described preferred embodiment, but it uses a somewhat measurable representation of the complexity of the internal state of the context model, such that one table stores the "first part" (valMps and pipeIdx) of all context models and the "second part" (refineIdx) of all context models. In this embodiment, the context model being updated is all context models As described in the previous preferred embodiment, the table for storing the "second part" is no longer needed and can be discarded.
[0200] Next, the context model being updated for the bin sequence according to the example is described.
[0201] In the LC-PIPE configuration, syntactic element bins of types coeff_significant_flag, coeff_abs_greater1, and coeff_abs_greater2 are classified into subsets. For each subset, a single context model is used to encode its bins. In this case, the context model is updated after a fixed number of bins have been encoded in this sequence. This is illustrated below in the multi-bin update example. However, this update differs from an update using only the last encoded bin and the internal state of the context model. For example, one context model update step for each encoded bin. This will be executed.
[0202] The following example illustrates the encoding of a typical subset consisting of eight bins. The letter "b" signifies bin decoding, and the letter "u" signifies context model update. In the case of LC-PIPE, only bin decoding is performed without updating the context model.
[0203] bbbbbbbb
[0204] In the case of HE-PIPE, the context model is updated after each bin is decoded.
[0205] bububububububu
[0206] To reduce complexity somewhat, context model updates can be performed after the bin sequence (in this example, updates to these four bins are performed after each of the four bins).
[0207] bbbbuuuubbbbuuuu
[0208] In other words, the selector 402 is configured to perform selections in the entropy decoder 322 in accordance with a probabilistic model associated with a given symbol type, with or without updating the associated probabilistic model, such that the frequency at which selections for symbols of a given symbol type are performed with updates is lower in the low complexity mode compared to the high efficiency mode.
[0209] In this case, after decrypting the four bins, four update steps follow, based on the four decrypted bins. Note that these four update steps can be performed in a single step using a special lookup table. For each possible combination of the four bins and each possible internal state of the context model, this lookup table stores the new state resulting after the four conventional update steps.
[0210] In certain modes, multibin updates are used for the syntax element coeff_significant_flag. For all other syntax element bins, The text model update is not used. The number of bins to be encoded before the multi-bin update step is set to n. When the number of bins in a set is not divisible by n, 1 to n-1 bins remain at the end of the subset after the last multi-bin update. For each of these bins, the conventional single-bin update is performed after all of these bins have been encoded. The number n can be any positive number greater than 1. Another mode may be the same as the previous mode, except that the multi-bin update is performed for any combination of coeff_significant_flag, coeff_abs_greater1 and coeff_abs_greater2 (instead of just coeff_significant_flag). Thus, this mode is more complex than the others. All other syntactic elements (where the multi-bin update is not used) are divided into two separate subsets, where the single-bin update is used for one subset and the context model update is not used for the other subset. Any possible other subset is valid (including an empty subset).
[0211] In another embodiment, the multibin update may be based only on the last m-bin of powder that is encoded immediately before the multibin update, where m is any natural number less than n. Thus, the decoding can be done as follows:
[0212] bbbbuubbbbuubbbbuubbbb Here, n=4 and m=2.
[0213] In other words, the selector 402 is configured to perform a selection in the entropy decoder 322 according to the probabilistic model associated with a given symbol type, along with updating the nth symbol of all associated probabilistic models based on the nearest symbol m of the given symbol type, such that the n / m ratio is higher in the low complexity mode compared to the high efficiency mode for symbols of a given symbol type.
[0214] In another preferred embodiment, for the syntactic element coeff_significant_flag, the context modeling scheme using local templates as described above for HE-PIPE configuration is used to assign the context model to the syntactic element bin. However, for these bins, no context model updates are used.
[0215] Furthermore, the selector 402 is configured to select one of many contexts in a sequence of symbols such that the number of contexts and / or previously searched symbols is lower in the low complexity mode than in the high efficiency mode for a given symbol type of symbol, and is configured to perform a selection in the entropy decoder 322 according to the probabilistic model associated with the selected context.
[0216] Probability model initialization using 8-bit initialization values
[0217] This section illustrates the initialization process of the measurable internal state of a probabilistic model that uses so-called 8-bit initialization values instead of two 8-bit values, as in the case of the highest-level video coding standard H.265 / AVC. It consists of two parts that are equivalent to the pair of initialization values used for the CABAC probabilistic model of H.264 / AVC. The two parts show two parameters of a linear equation that computes the initial state of the probabilistic model, representing a specific probability (e.g., in the form of a PIPE index) from QP.
[0218] The first part shows the slope, which takes advantage of the dependence of the internal state with respect to the quantization parameter (QP) used during encoding or decoding. The second part determines the PIPE index with a predetermined QP, similar to valMps.
[0219] Two different modes can be used to initialize the probabilistic model using predetermined initialization values. The first mode represents QP-independent initialization. It simply uses the PIPE index and valMps defined in the second part of the initialization value for all QPs. This is equivalent to the case where the slope is equal to 0. The second mode represents QP-dependent initialization, and furthermore, it uses the slope of the first part of the initialization value to determine an improved index by changing the PIPE index. The two parts of an 8-bit initialization value are exemplified below.
[0220] [Table 8]
[0221] It consists of two 4-bit parts. The first part contains an index indicating one of 16 different predetermined slopes stored in an array. A predetermined slope consists of one of seven negative slopes (slope indices 0-6), a slope equal to zero (slope index 7), and eight positive slopes (slope indices 8-15). The slopes are shown in Table 9.
[0222] [Table 9]
[0223] All values are estimated with 256 factors to avoid the use of floating-point arithmetic. The second part is the PIPE index which embodies the probability of valMps=1 rising between probability intervals p=0 and p=1. In other words, PIPE encoder n operates at a higher model probability than PIPE encoder n-1. For any probability model, one PIPE probability index is available, which is p for probability intervals QP=26. valMPs=1 Check the PIPE encoder, which includes the probability.
[0224] [Table 10]
[0225] The initialization of the internal state of the probabilistic model requires calculating the initialization of QP and the 8-bit initialization value by computing a simple linear equation of the form y = m * (QP - QPref) + 256 * b. Note that m determines the slope taken from Table 9 using the slope index (the first part of the 8-bit initialization value), and b represents the PIPE encoder with QPref = 26 (the second part of the 8-bit initialization value: "PIPE probability index"). If y is greater than 2047, then valMPS is 1 and pipeIdx is equal to (y - 2048) >> 8. Otherwise, valMPS is 0 and pipeIdx is equal to (2047 - y) >> 8. If valMPS is equal to 1, then the improved index is equal to (((y - 2048) & 255) * numStates) >> 8. Otherwise, the improved index is equal to (((2047―y)&255)*numStates)>>8. In either case, numStates is equal to the number of CABAC states in pipeIdx, as illustrated in Table 7.
[0226] The scheme described above is used not only in combination with the PIPE encoder, but also in relation to the CABAC scheme described above. Without PIPE, the number of CABAC states, i.e., the (pState_current[bin]) probabilistic states, in which probabilistic update state transitions occur for each PIPE Idx (i.e., each most important bit of pState_current[bin]), is in fact just a set of parameters that implement a piecewise linear interpolation of CABAC states depending on the QP. Furthermore, this piecewise linear interpolation may be effectively invalid if the parameter numStates uses the same value for all PIPE Idx. For example, setting numStates for 8 for all cases yields a total of 16*8 aspects and an improved index calculation that simplifies the index to either valMPS equals 1, ((y-2048)&255)>>5 or valMPS equals 0, ((2047-y)&255)>>5. In this case, mapping the representation using valMPS, PIPE idx, and improved idx to the representation used by the original H.264 / AVC CABAC is very straightforward. The CABAC state is given as (PIPE Idx << 3) + refinement Idx. This aspect is described further below with respect to Figure 16.
[0227] Unless the slope of the 8-bit initialization value is equal to zero, or QP is equal to 26, it is necessary to calculate the internal state by using a linear equation with QP for the encoding or decoding process. If the slope is equal to zero or QP for the current encoding method is equal to 26, a second part of the 8-bit initialization value can be used directly to initialize the internal state of the probabilistic model. Otherwise, the fractional part of the resulting internal state can be further used to determine an improved index for a high-efficiency encoding application by linear interpolation between the limits of a particular PIPE encoder and the limits of the current PIPE encoder. In this preferred embodiment, linear interpolation is simply a modification available to the current PIPE encoder. This is done by multiplying the fractional part, which represents the total number of good indexes, by the result and mapping it to the nearest integer improved index.
[0228] The process of initializing the internal state of a probabilistic model can vary with respect to the number of PIPE probabilistic index states. In particular, the duplication of using two different PIPE indices that distinguish between 1 and 0 for PIPE encoder E1, i.e., MPS, which uses a mode with equal probability, can be avoided as follows. The process can also be triggered during the beginning of parsing of slice data, and the input to this process can be an 8-bit initialization value, as illustrated in Table 11, which is sent within the range of the bit stream for any context model to be initialized.
[0229] [Table 11]
[0230] The first four bits are used to define the slope index, and the search is performed by masking bits b4-b7. For every slope index, the slope (m) is specified and shown in Table 12.
[0231] [Table 12]
[0232] Bits b0-b3, the last four bits of the 8-bit initialization value, determine probIdx, where a given QP.probIdx0 indicates the highest probability for a symbol with the value 0, and probIdx14 indicates the highest probability for a symbol with the value 1. Table 13 shows the corresponding pipe encoder and its valMps for each probIdx.
[0233] [Table 13]
[0234] For both values, the internal state can be calculated using a linear equation of the form y=m*x+256*b, where m represents the slope, x represents the QP of the current slice, and b is derived from probIdx as shown in the description below. All values in this process are scaled by a factor of 256 to avoid the use of floating-point operations. The output (y) of this process represents the internal state of the probability model at the current QP and is stored in 8-bit memory. As shown in G, the internal state consists of valMPs, pipeIdx and refineIdx.
[0235] [Table 14]
[0236] The assignment of refineIdx and pipeIdx is similar to the internal state of the CABAC probability model (pStateCtx), and is shown in H.
[0237] [Table 15]
[0238] In a preferred embodiment, probIdx is defined by QP26. Based on the 8-bit initialization values, the internal states of the probability model (valMps, pipeIdx and refineIdx) are processed as described in the following pseudo-code.
[0239] TIFF2026143473000030.tif123146
[0240] As shown in the pseudocode, refineIdx is calculated by linearly inserting it between intervals of pipeIdx and quantizing the result to the corresponding refineIdx. The offset identifies the total number of refineIdx for each pipeIdx. The interval [7, 8) of fullCtxState / 256 is split in half. Interval [7, 7.5) maps to pipeIdx=0 and valMps=0, and interval [7.5, 8) maps to pipeIdx=0 and valMps=1. Figure 16 illustrates how the internal state is extracted and shows the mapping of fullCtxState / 256 to pStateCtx.
[0241] Note that the slope exhibits a dependency on probIdx and QP. When the 8-bit initialization value slopeIdx is equal to 7, the resulting internal state from the probabilistic model is common to all slice QPs—and therefore the initialization process of the internal state is independent of the slice's current QP.
[0242] In other words, selector 402 initializes a pipe index used to decode the next portion of a data stream, such as the entire stream or the next slice, using a syntactic element that indicates the quantization step size QP used to quantize the data in this portion, such as the transformation coefficient level contained within a table common to both modes LC and HC, using this syntactic element as an index. A table like Table 10 contains a pipe index for each symbol type, its respective reference QPref, or other data for each symbol type. Depending on the actual QP of the current portion, the selector can calculate the pipe index value, such as multiplying a by (QP-QPref), using the actual QP and the respective table item a indexed by the QP itself. The only difference between LC and HE modes: the selector is different in LC compared to HE mode. In some cases, the result is simply calculated with low precision. The selector can, for example, simply use the integer part of the calculation result. In HE mode, high-precision residuals, such as fractional parts, are used to select one of the available improved indices for each pipe index, as indicated by the low-precision or integer part. Improved indices are used in HE mode (and potentially less frequently in LC mode) to perform probabilistic fitting, for example, by using the table walk described above. When leaving an available index for the current pipe index at a higher boundary, the higher pipe index minimizes the improved index and is then selected. When leaving an available index for the current pipe index at a lower boundary, the next lower pipe index maximizes the available improved index for the new pipe index and is then selected. The pipe index, together with the improved index, defines the probabilistic state, but for selection within a partial stream, the selector simply uses the pipe index. Improved indices are only useful for tracking probabilities more closely or for finer precision.
[0243] However, the above explanation also showed that complexity scalability can be achieved using the decoder shown in Figure 12, separate from the PIPE coding concept in Figures 7-10. The decoder in Figure 12 is for decoding a data stream 601 on which media data has been encoded and includes a mode switch 600 configured to activate a low-complexity mode or a high-efficiency mode depending on the data stream 601, and a desymbolizer configured to symbolically represent a sequence of symbols 603 obtained from the data stream 601—either directly or, for example, by entropy decoding—to obtain integer-valued syntax elements 604, which use a mapping function controllable by a control parameter to map a region of symbol sequence words to a common region of integer-valued syntax elements. A reproducer 605 is configured to reproduce the media data 606 based on the integer-valued syntax elements. The desymbolizer 602 is configured to perform desymbolization such that, when the high-efficiency mode is activated, the control parameters change according to the data stream at a first rate, and when the low-complexity mode is activated, the control parameters remain constant regardless of the data stream or changes according to the data stream, except at a second rate lower than the first rate, as indicated by arrow 607. For example, the control parameters can change according to previously desymbolized symbols.
[0244] Some of the embodiments described above utilized the configuration shown in Figure 12. The syntactic elements coeff_abs_minus3 and MVD in sequence 327 are binarized in the desymbolizer 314 depending on the selected mode, for example, as shown in 407, and the regenerator 605 uses these syntactic elements for regeneration. Clearly, both configurations in Figures 11 and 19 are readily integrable, but the configuration in Figure 12 can also be combined with other encoding environments.
[0245] For example, see the motion vector difference coding shown above. The desymbolizer 602 is configured to use a combination of a shortened unary code that performs mapping within a first interval of integer-value syntactic elements smaller than the cutoff value, and a prefix in the form of the shortened unary code for the cutoff value and a suffix in the form of a VLC codeword within a second interval of integer-value syntactic elements including or exceeding the cutoff value. The decoder includes an entropy decoder 608 configured to extract multiple first bins of the shortened unary code from the data stream 601 using unary entropy decoding with a changing probabilistic evaluation and multiple second bins of the VLC codeword using a constant equiprobability bypass mode. In HE mode, as indicated by arrow 609, entropy coding is more complex than in LC coding. That is, context adaptation and / or probabilistic fitting are applied in HE mode and suppressed in LC mode, or the complexity is expanded or reduced as described above with respect to the various embodiments.
[0246] An encoder conforming to the decoder of Figure 11, which encodes media data into a data stream, is shown in Figure 13. It includes an inserter 500 configured to signal the activation of either a low-complexity mode or a high-efficiency mode in the data stream 501; a constructor 504 configured to pre-encode media data 505 into a sequence of syntactic elements 506; a symbolizer 507 configured to symbolically represent the sequence of syntactic elements 506 into a sequence of symbols 508; a plurality of entropy encoders 310, each configured to translate a partial sequence of symbols into a codeword in the data stream; and a selector 502 configured to send each symbol of the sequence of symbols 508 to a selected one of the plurality of entropy encoders 310, the selector 502 configured to perform a selection depending on which of the low-complexity mode and high-efficiency mode is operating, as indicated by the arrow 511. An interleaver 510 may optionally be provided to sandwich the codeword of the encoder 310.
[0247] An encoder adapted to the decoder of Figure 12 for encoding media data into a data stream, shown in Figure 14, includes an inserter 700 configured to signal operation in low complexity mode or high efficiency mode within the data stream 701, a constructor 704 configured to pre-encode media data 705 into a sequence of integer-value syntactic elements 706, and a constructor 707 configured to represent integer-value syntactic elements with symbols using a mapping function controllable by a control parameter to map the region of integer-value syntactic elements to a common region of symbol sequence words, wherein when high efficiency mode is activated, the control parameter changes according to the data stream at a first rate, and when low complexity mode is activated, the symbolizer 707 performs symbolization at a second rate lower than the first rate, although the control parameter remains constant regardless of the data stream or the data stream-dependent change, as indicated by arrow 708. The result of symbolization is encoded into the data stream 701.
[0248] Furthermore, it should be noted that the embodiment in Figure 14 is readily transferable to the implementation of the context-adaptive binary coding / decoding described above: the selector 509 and entropy encoder 310 are condensed into a context-adaptive binary encoder that outputs directly to the data stream 401, selecting to extract the context for the current bin from the data stream. This is especially true for context-adaptive and / or probabilistic adaptation. During low complexity modes, both functions / adaptives can be switched off or designed to be more relaxed.
[0249] As briefly shown above, the ability to switch modes, as described in some of the above embodiments, is separated according to another embodiment. To make this clear, Figure 16 shows an example summarizing the above description, where the only difference between the embodiments of Figure 16 and the above embodiments is the removal of the ability to switch modes. Furthermore, the following description highlights the advantages that arise from initializing the probabilistic evaluation of the context with less precise parameters of slope and offset compared to H.264, for example.
[0250] Figure 16 shows a decoder for decoding video 405 from data stream 401, where the horizontal and vertical components of the motion vector difference are encoded using binarization of the horizontal and vertical components, in particular, within a first interval of combinations of horizontal and vertical component regions below the cutoff value and prefixes in the form of abbreviated unary codes, the binarization is equal to the abbreviated unary codes of the horizontal and vertical components, respectively. The cutoff value and suffixes in the form of exponential Golomb codes of the horizontal and vertical components are in a second interval of horizontal and vertical component regions that are included in or above the cutoff value, where the cutoff value is 2 and the exponential Golomb code is in the order of 1. The decoder includes an entropy decoder 409 configured to extract short unary codes from a data stream using context-fitted binarization entropy decoding, which ensures there is one context for each bin position of a shortened unary code common to exponential Golomb codes, using a constant equal-probability bypass mode for the horizontal and vertical components of the motion vector difference and to obtain the binarization of the motion vector difference, i.e., using several parallel operating entropy decoders 322 together with a selector / assignor A desymbolizer 314 that de-binarizes the binarization of motion vector difference syntactic elements to obtain integer values of the horizontal and vertical components of the motion vector difference, respectively, and a reconstructor 404 reconstructs the video based on the integer values of the horizontal and vertical components of the motion vector difference.
[0251] To illustrate this in more detail, an embodiment is briefly shown in Figure 18. 800 represents, as a representative example, a vector representing a motion vector difference, i.e., the predicted residual between the predicted motion vector and the actual / reproduced motion vector. Horizontal and vertical components 802x and 802y are exemplified. They may be transmitted in units of pixel position, i.e., pixel pitch or a more precise position than one pixel unit (e.g., half or a quarter of the pixel pitch). The horizontal and vertical components 802x,y are integers to be evaluated. Their range extends from zero to infinity. The marker values are treated separately and are not considered here. In other words, the explanation outlined here focuses on the magnitude of the motion vector difference 802x,y. The region is exemplified by 804. To the right of the region axis 804, Figure 19 illustrates binarization, where each possible value is mapped (binarized) in relation to the possible values of the components 802x,y, each positioned perpendicular to it. As can be seen, below the cutoff value of 2, only the abbreviated unary code 806 occurs, but as a suffix, there is also the exponential Golomb code of order 808 from possible values equal to or greater than the cutoff value of 2, as a reminder of integer values above the cutoff value minus 1. For all bins, only two contexts are provided: one is the first bin position for the binarization of the horizontal and vertical components 802x,y, and the other is the second bin position for the abbreviated unary code 806 of both the horizontal and vertical components 802x,y. For the bin positions of the exponential Golomb code 808, the equal-probability bypass mode is used by the entropy decoder 409. That is, both bin values are assumed to occur equally likely. Probability estimates for these bins are determined. In comparison, the probability estimates for the bins of the abbreviated unary code 806, associated with the two just mentioned contexts, are applied sequentially during decoding.
[0252] Before going into further detail, regarding how the entropy decoder 409 can perform the work just mentioned in accordance with the above description, the description focuses on a possible implementation of a reproducer 404 using motion vector differences and their integer values, as obtained by the desymbolizer 314 by debinarization of bins of codes 106 and 108 using debinarization shown in Figure 18. In particular, the reproducer 404 can retrieve from the data stream 401 information about the subdivision of the currently reproduced image into blocks, at least some of which depend on motion compensation predictions, as described above. Figure 19 shows the subdivision of the just mentioned image 120, which is typically reproduced in 820 and in 822, where motion compensation predictions are used to predict the image content therein. The subdivision and the size of the block 122 may vary, as described in Figures 2A-2C. To avoid transmitting the motion vector difference 800 for each of these blocks 122, the regenerator 404 can utilize the coupling concept, which, in addition to the fact that the data stream is fixed as a subdivision, further transmits coupling information in addition to or without subdivision information. The coupling information signals to the regenerator 404 which blocks 822 form a group. This measurement makes it possible for the regenerator 404 to apply a specific motion vector difference 800 to the entire coupling group of blocks 822. Naturally, on the encoding side, Transmission of combination information depends on the trade-off among re-split transmission overhead (if any), combination information transmission overhead and motion vector difference transmission overhead, which decreases as the size of the combination group increases. On the other hand, an increase in the number of blocks per combination group reduces the fit of the motion vector differences for these combination groups to the actual needs of individual blocks of each combination group, thereby resulting in less accurate motion compensated prediction for the motion vector differences of these blocks, which requires higher transmission overhead for transmission of prediction residuals, for example, in the form of transmission coefficient levels. Therefore, the trade-off is found on the encoding side by an appropriate method. However, in any case, the combination concept results in motion vector differences for combination groups that exhibit low internal correlation. For example, refer to Figure 19 which shows this by shading the membership for a given combination group. Clearly, the actual motion of the image content of these blocks was similar to that which the encoding side decided to combine the respective blocks. However, the correlation with the motion of the image content of other combination groups is low. Therefore, the limitation of using only one context per bin of the shortened unary code 806 does not negatively affect the entropy coding efficiency of such a combination concept that is already adapted to the spatial entropy coding efficiency between motion of sufficiently adjacent image content. The context can only be selected based on the fact that the bin is part of the binarization of the motion vector difference component 802x,y and the bin position which is 1 or 2 due to the cut-off value being 2. Therefore, other already decoded bins / syntax elements / mvd components 802x,y do not affect context selection.
[0253] Similarly, the regenerator 404 is configured, firstly, to reduce the information content transmitted via the motion vector difference (beyond the spatial and / or temporal prediction of the motion vectors) by using a multi-hypothetical prediction concept generated for each block or combined group, which is explicitly or implicitly transmitted in the data stream information on the index of the predictor actually used to predict the motion vector difference. See, for example, the unshaded block 122 in Figure 20. The regenerator 404 can provide different predictors for the motion vector of this block by spatially predicting the motion vector from, for example, from the left, from the top, a combination of both, etc., and by temporally predicting the motion vector from the motion vector of the portion of the previously decoded video image and further combinations of the aforementioned predictors that are located at the same position. These predictors are sorted by the regenerator 404 in a predictable manner that is predictable on the encoding side. Some information is transmitted for this purpose in the data stream and used by the regenerator. That is, some hints are included in the data stream so that, with respect to them, the predictors from this ordered list of predictors are actually used as predictors for the motion vector of this block. This index can clearly be transmitted within the data stream for this block. However, it is possible for the index to be predicted first and simply convey its prediction. Other possibilities exist as well. In any case, the prediction scheme just mentioned allows for a very accurate prediction of the motion vector of the current block, and thus the information content requirement imposed on the motion vector difference is reduced. Thus, the constraints of context-adaptive entropy coding on two bins of the shortened unary code and a reduction of the cutoff value to 2, as described in Figure 18, along with the selection of the order of the exponential Golomb code, which is 1, along with the selection of two bins of the shortened unary code, show a frequency histogram in which higher values of the motion vector difference components 802x,y are accessed less frequently, for high prediction efficiency, and the constraints of context-adaptive entropy coding on a reduction of the cutoff value to 2, as described in Figure 18, do not negatively affect coding efficiency.Because predictions tend to work equally well in both directions with high prediction accuracy, even omitting any distinctive features between the horizontal and vertical components can lead to effective predictions.
[0254] In the foregoing description, as far as the functions of the desymbolizer 314, the regenerator 404, and the entropy decoder 409 are concerned, all the details of which are included in Figures 1-15 are shown, for example, in Figure 16. It is important to note that these components can also be transferred. Nevertheless, for the sake of completeness, these details are outlined again below.
[0255] For a better understanding of the prediction scheme outlined, please refer to Figure 20. As just described, the constructor 404 can obtain different predictors for the current block 822 or combined group of the current block, and these predictors are shown by the solid vector 824. The predictors are obtained by spatial and / or temporal prediction, and furthermore, arithmetic mean operations, etc., can be used so that the individual predictors are obtained to some extent by the reproducer 404 so that they correlate with one another. Apart from the way the vector 826 is obtained, the reproducer 404 sequences or sorts these predictors 126 into an ordered list. This is illustrated in numbers 1-4 of Figure 21. When the sorting process is independently decidable, it is preferable that the encoder and decoder operate synchronously. Then, the index just mentioned can be obtained explicitly or implicitly from the data stream by the reproducer 404 for the current block or combined group. For example, if a second predictor "2" is selected, the regenerator 404 adds a motion vector difference 800 to this selected predictor 126, thereby obtaining the last reproduced motion vector 128, which is used by motion-compensated prediction to predict the contents of the current block / joint group. In the case of a joint group, it is possible that the regenerator 404 includes further motion vector differences provided for the blocks in order to further refine the motion vector 128 with respect to the individual blocks of the joint group.
[0256] Thus, continuing with the description of the implementation of the components shown in Figure 16, it may also be that the entropy decoder 409 is configured to extract the abbreviated unary code 806 from the data stream 401 using binary arithmetic decoding or binary PIPE coding. Both concepts are described above. Furthermore, the entropy decoder 409 can be configured to use the abbreviated unary code 806 or different contexts for two bin positions in the same context for both bins. The entropy decoder 409 can be configured to perform a stochastic state update. The entropy decoder 409 can do this by transitioning from the current stochastic state associated with the context selected for the bin being extracted to a new stochastic state corresponding to the bin being extracted from the abbreviated unary code 806 for the bin being extracted. Refer to the tables Next_State_LPS and Next_State_MPS above, which are table looks performed by the entropy decoder in addition to the other steps 0-5 above. In the above description, the current stochastic state is referred to by pState_current. It is determined for each context of interest. The entropy decoder 409 is configured to quantize the bin currently drawn from the abbreviated unary code 806 by indicating the current probability interval to obtain the current probability interval width value, i.e., the probability interval index q_index, and performing interval repartition by quantizing R, which is currently drawn from the binThe entropy decoder 409 is configured to select between two partial intervals based on the current probability interval, i.e., the offset state value from within V, update the probability interval width value R and the offset state value, estimate the value of the bin currently drawn using the selected partial interval, and obtain the updated probability interval width value R and the sequence of bits read from the data stream 401. The renormalization of the offset value R is performed. For example, the entropy decoder 409 is configured to perform binary arithmetic decoding of the bins from the exponential Golomb code by having the current probability interval width value to obtain a subdivision of the current probability interval width value into two partial intervals. Halving corresponds to fixing the probability evaluation to 0.5 and equalizing it. This can be done by a simple bit shift. For each motion vector difference, the entropy decoder is configured to extract the shortened unary codes of the horizontal and vertical components of each motion vector difference from the data stream 401, before the exponential Golomb code of the horizontal and vertical components of each motion vector difference. By this measurement, the entropy decoder 409 can take advantage of the fact that a higher number of bins together form a continuation of bins with a fixed probability evaluation, i.e., fixed to 0.5. This can speed up the entropy decoding procedure. On the other hand, the entropy decoder 409 prefers to maintain the order in the motion vector differences by first extracting the horizontal and vertical components of one motion vector difference, and then extracting the horizontal and vertical components of the next motion vector difference. This measurement allows the desymbolizer 314 to immediately proceed with de-binarization of the motion vector difference without waiting for further motion vector difference scans, thus reducing the memory requirements imposed on the decoding components, i.e., the decoder in Figure 16. This is made possible by context selection: only exactly one context is available for the bin position of code 806.
[0257] As described above, the reproducer 404 can spatially and / or temporally predict the horizontal and vertical components of the motion vector by obtaining the predictor 126 for the horizontal and vertical components of the motion vector and reproducing the horizontal and vertical components of the motion vector difference by improving the predictor 826, for example, by simply adding the motion vector difference to each predictor.
[0258] Furthermore, the reproducer 404 can be configured to reproduce the horizontal and vertical components of a motion vector by predicting the horizontal and vertical components of the motion vector in different ways to obtain an ordered list of predictors for the horizontal and vertical components of the motion vector, obtaining list indices from the data stream, and improving the predictors into list predictors whose list indices indicate the horizontal and vertical components of the motion vector.
[0259] Furthermore, as described above, the regenerator 404 is configured to regenerate the video using motion compensation prediction by applying the horizontal and vertical components 802x,y of the motion vector with spatial accuracy defined by the re-dividing of the video image into blocks, and the regenerator 404 uses the join syntax elements in the data stream 401 to group the blocks into join groups and apply integer values of the horizontal and vertical components 802x,y of the motion vector difference obtained by the binarizer 314 to the units of the join group.
[0260] The regenerator 404 can extract the re-partition of the video image into blocks from some data stream 401, excluding the join syntax elements. The regenerator 404 adapts the horizontal and vertical components of a given motion vector for all blocks of the associated join group, or improves it with the horizontal and vertical components of the motion vector difference associated with the blocks of the join group.
[0261] For completeness only, Figure 17 shows an encoder adapted to the decoder of Figure 16. The encoder of Figure 17 includes a constructor 504, a symbolizer 507, and an entropy decoder 513. The encoder is configured such that the constructor 504 encodes video 505 by motion compensation prediction using motion vectors, predictively encodes motion vectors by predicting motion vectors, and sets integer values 506 of the horizontal and vertical components of the motion vector difference indicating the prediction error of the predicted motion vectors; and the horizontal component of the motion vector difference To obtain the binarization of the horizontal and vertical components 508, integer values are binarized, the binarization being equal to a combination of a prefix for the cutoff value and a suffix in the form of the Exp-Golomb code for the horizontal and vertical components in a first interval of regions of horizontal and vertical components smaller than the cutoff value, and a symbolizer 507 in a second interval of regions of horizontal and vertical components including or greater than the cutoff value, where the cutoff value is 2 and the Exp-Golomb code is 1; and an entropy encoder 513 configured to encode the shortened unary codes into a data stream using context-adapted binary entropy coding, ensuring that there is one context for each bin position of the shortened unary code, common to the horizontal and vertical components of the motion vector difference and the Exp-Golomb code using a constant equiprobability bypass mode. Further details of possible implementations can be transferred directly from the description of the decoder in Figure 16 to the encoder in Figure 17.
[0262] Although several embodiments are described in the context of the apparatus, it is clear that these embodiments can also be applied to the description of the corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, embodiments described in the context of a method step represent a description of the corresponding block or component or feature of the corresponding apparatus. Some or all of the method steps can be performed by (using) a hardware device, such as a microprocessor, a programmable computer, or electronic circuitry. In some embodiments, one or more of the most important method steps can be performed by this type of apparatus.
[0263] The encoded signal of the invention can be stored in a digital storage medium or transmitted over a transmission medium such as a wireless or wired transmission medium, such as the Internet.
[0264] Depending on the specific implementation requirements, embodiments of the present invention may be implemented in hardware or in software. Implementation may be carried out using a digital storage medium having electronically readable control signals stored thereon, such as a flexible disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, and may (or may) cooperate with a programmable computer system so that each method is performed. Therefore, the digital storage medium may also be computer-readable.
[0265] Some embodiments of the present invention include a data carrier having an electronically readable control signal, which can cooperate with a programmable computer system so that one of the methods described in this specification is performed.
[0266] Typically, embodiments of the present invention can be implemented as a computer program product having program code, where the program code is implemented to perform one of the methods when the computer program product runs on a computer. The program code can be stored, for example, in a machine-readable carrier.
[0267] Other embodiments include computer programs for performing one of the methods described in this specification, which are stored in a machine-readable carrier.
[0268] In other words, an embodiment of the method of the invention is, therefore, a computer program that, when run on a computer, has program code for performing one of the methods described in this specification.
[0269] A further embodiment of the method of the invention is a data carrier (or digital storage medium or computer-readable medium) having a computer program recorded thereon for performing one of the methods described in this specification. The data carrier, digital storage medium, or recorded material is typically tangible and / or non-transferable.
[0270] A further embodiment of the method of the invention is a data stream or sequence of signals representing a computer program for performing one of the methods described in this specification. The data stream or sequence of signals may be configured to be transmitted over a data communication connection, for example, over the Internet.
[0271] Further embodiments include processing means configured or adapted to perform one of the methods described herein, such as a computer or a programmable logic device.
[0272] Further embodiments include a computer on which a computer program for performing one of the methods described in this specification is installed.
[0273] Further embodiments of the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) a computer program to perform one of the methods described in this specification to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may be, for example, a file server for transmitting the computer program to the receiver.
[0274] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field-programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. Typically, the methods are preferably performed by any hardware device.
[0275] The embodiments described above are for illustrative purposes only, in order to illustrate the principles of the present invention. Modifications and changes in arrangement and details described herein will be obvious to those skilled in the art. Therefore, it is not intended to be limited solely by the scope of the imminent patent claims, or solely by the description of the embodiments and the specific details shown herein as descriptive.
Claims
1. A decoder for decoding video from a data stream encoded using binarization of the horizontal and vertical components of a motion vector difference, wherein the binarization is equal to a combination of a prefix in the form of a shortened unary code for the horizontal and vertical components in a first interval of horizontal and vertical component regions lower than a cutoff value, and a suffix in the form of an Exp-Golumb code for the horizontal and vertical components in a second interval of horizontal and vertical component regions equal to or higher than the cutoff value, the cutoff value is 2, and the Exp-Golumb code has order 1, An entropy decoder configured to extract a shortened unary code from a data stream using context-adaptive binary entropy decoding, which has exactly one context for each bin position of a shortened unary code, common to Exp-Golomb codes using a constant equiprobability bypass mode to obtain the horizontal and vertical components of the motion vector difference and the binarization of the motion vector difference; A desymbolizer configured to debinarize the binarization of motion vector difference syntax elements in order to obtain integer values for the horizontal and vertical components of the motion vector difference; A decoder including a reproducible that is configured to reconstruct video based on integer values of the horizontal and vertical components of the motion vector difference.
2. The decoder according to claim 1, wherein the entropy decoder (409) is configured to extract a shortened unary code (806) from a data stream (401) using binary arithmetic decoding or binary PIPE decoding.
3. The decoder according to claim 1 or 2, wherein the entropy decoder (409) is configured to use different contexts for two bin positions of the abbreviated unary code 806.
4. The decoder according to any one of claims 1 to 3, wherein the entropy decoder (409) is configured to perform a stochastic state update for the bin currently drawn from the abbreviated unary code (806) by transitioning from the current stochastic state associated with the context selected for the bin currently drawn to a new stochastic state corresponding to the bin currently drawn.
5. The decoder according to any one of claims 1 to 4, wherein the entropy decoder (409) is configured to perform binary decoding of the bin currently drawn from a shortened unary code (806) by obtaining a probability interval index by deciphering a table item among a plurality of table items using a probability interval index and a probability state index depending on the current probability state associated with the context selected for the bin currently drawn, in order to obtain a subdivision of the current probability interval into two partial intervals, and by quantizing the current probability interval width value indicating the current probability interval in order to perform the interval subdivision.
6. The decoder according to claim 5, configured to use an 8-bit representation for the current probability interval width value in the quantization of the current probability interval width value, and to use the most important bit of the 8-bit representation for grab-out 2 or 3.
7. The entropy decoder (409) is configured to select from within the current probability interval one of two partial intervals based on an offset value, update the probability interval width value and offset state value, and use the selected partial interval to continue reading bits from the data stream (401) with the updated probability interval width value and offset state The decoder according to claim 5 or 6, which performs value renormalization to estimate the value of the bin currently being drawn.
8. The decoder according to any one of claims 5 to 7, wherein the entropy decoder (409) is configured to perform binary decoding of bins from the Exp-Golomb code by halving the current probability interval width value to obtain a subdivision of the current probability interval into two partial intervals in a constant equal probability bypass mode.
9. The decoder according to any one of claims 1 to 8, wherein the entropy decoder (409) is configured to extract, for each motion vector difference, the codes of the shortened unary horizontal and vertical components of each motion vector difference from the data stream, prior to the Exp-Golumb codes of the horizontal and vertical components of each motion vector difference.
10. The decoder according to any one of claims 1 to 9, wherein the reproducer is configured to spatially and / or temporally predict the horizontal and vertical components of a motion vector in order to reproduce the horizontal and vertical components of a motion vector by obtaining predictors for the horizontal and vertical components of a motion vector and improving the predictor 826 with the horizontal and vertical components of the motion vector difference.
11. The decoder according to any one of claims 1 to 10, wherein the reproducer is configured to predict the horizontal and vertical components of a motion vector in different ways so as to obtain an ordered list of predictors for the horizontal and vertical components of the motion vector, obtain a list index from a data stream, and reproduce the horizontal and vertical components of the motion vector by improving the list of predictors indicated by the list index using the horizontal and vertical components of the motion vector difference.
12. The decoder according to claim 10 or 11, wherein the regenerator is configured to reproduce the video using motion compensation prediction by using the horizontal and vertical components of the motion vector.
13. The decoder according to claim 12, wherein the reproducer is configured to reproduce the video using motion compensation prediction by applying the horizontal and vertical components of the motion vector with spatial precision defined by the subdivision of the video images of the blocks, and the reproducer uses join syntax elements present in the data stream in the units of the join group to group the blocks into join groups and apply the horizontal and vertical components of the motion vector difference obtained by the non-binarizer.
14. The decoder according to claim 13, wherein the regenerator is configured to extract the re-fraction of block video images from portions of the data stream that exclude concatenation syntax elements.
15. The decoder according to claim 13 or 14, wherein the reproducible is configured to employ the horizontal and vertical components of a predetermined motion vector for all blocks of the associated coupling group, or to improve it by the horizontal and vertical components of the motion vector difference associated with the blocks of the coupling group.
16. An encoder for encoding video into a data stream, A constructor configured to predictively encode video using motion vectors and motion compensation prediction, predictively encode motion vectors by predicting motion vectors, and set integer values for the horizontal and vertical components of the motion vector difference representing the prediction error of the predicted motion vector; The system is configured to binarize integer values to obtain binarization of the horizontal and vertical components of the motion vector difference, where the binarization is equivalent to a combination of prefixes in the form of abbreviated unary codes for the horizontal and vertical components in a first interval of the regions of the horizontal and vertical components lower than the cutoff value, and suffixes in the form of Exp-Golumb codes for the horizontal and vertical components in a second interval of the regions of the horizontal and vertical components equal to or higher than the cutoff value, where the cutoff value is 2 and the Exp-Golumb codes are symbolizers with order 1; and An encoder comprising an entropy encoder that encodes a shortened unary code into a data stream using context-adaptive binary entropy coding, where the horizontal and vertical components of the motion vector difference and the Exp-Golom code using a constant equiprobability bypass mode are common, and there is exactly one context for each bin position of the shortened unary code.
17. The encoder according to claim 16, wherein the entropy encoder is configured to encode a shortened unary code into a data stream using binary arithmetic coding or binary PIPE coding.
18. The encoder according to claim 16 or 17, wherein the entropy encoder is configured to use different contexts for two bin positions of a shortened unary code.
19. The encoder according to any one of claims 16 to 18, wherein the etropy encoder is configured to perform a stochastic state update in accordance with the bin currently drawn out by transferring from the current stochastic state associated with the context selected for the bin currently to be encoded, for the bin currently to be encoded from a shortened unary code.
20. The encoder according to any one of claims 16 to 19, wherein the entropy encoder is configured to obtain a probability interval index by deciphering a table item among a plurality of table items using a probability interval index and a probability state index depending on the current probability state associated with the context selected for the bin to be drawn, in order to obtain a subdivision of the current probability interval into two partial intervals, and to perform a binary computation encoding of the bin to be currently encoded from a shortened unary code by quantizing the current probability interval width value indicating the current probability interval.
21. The encoder according to claim 20, configured to use an 8-bit representation for the current probability interval width value in the quantization of the current probability interval width value, and to use the most important bit of the 8-bit representation for grab-out 2 or 3.
22. The encoder according to claim 20 or 21, wherein the entropy encoder is configured to select from two partial intervals based on the integer value of the bin currently being encoded, to update the probability interval width and probability interval offset using the selected partial interval, and to perform a renormalization of the probability interval width and probability interval offset, which includes continuing to write bits to the data stream.
23. The encoder according to any one of claims 20 to 21, wherein the entropy encoder is configured to binary encode bins from an Exp-Golom code by halving the current probability interval width value to obtain a subdivision of the current probability interval into two partial intervals.
24. The entropy encoder adds each motion vector difference to the data stream before the Exp-Golumb code for the horizontal and vertical components of each motion vector difference. The encoder according to any one of claims 16 to 23, configured to encode the codes of the shortened unary horizontal and vertical components of a vector difference.
25. The encoder according to any one of claims 16 to 24, wherein the constructor is configured to obtain predictors for the horizontal and vertical components of a motion vector, thereby improving the predictor with respect to the horizontal and vertical components of the motion vector, and to spatially and / or temporally predict the horizontal and vertical components of the motion vector difference.
26. The encoder according to any one of claims 16 to 25, wherein the constructor is configured to predict the horizontal and vertical components of a motion vector in different ways to obtain an ordered list of predictors for the horizontal and vertical components of the motion vector, to determine the list indices and to insert information revealing them into the data stream to improve the list predictors that the list indices indicate with respect to the horizontal and vertical components of the motion vector, and to determine the horizontal and vertical components of the motion vector difference.
27. The encoder according to any one of claims 16 to 26, wherein the constructor is configured to encode the video using motion compensation prediction by applying the horizontal and vertical components of the motion vector with spatial precision defined by the subdivision of the video images in the block, and the constructor determines and inserts into the data stream a join syntax element in a unit of the join group so as to group the blocks into join groups and apply the horizontal and vertical components of the motion vector difference which depend on binarization by a non-binarizer.
28. The encoder according to claim 27, wherein the constructor is configured to encode the subdivision of a block of video images into portions of a data stream that exclude associative syntax elements.
29. The encoder according to claim 27 or 28, wherein the constructor is configured to take the horizontal and vertical components of a predetermined motion vector for all blocks of the associated coupling group, or to improve it by the horizontal and vertical components of the motion vector difference associated with the blocks of the coupling group.
30. A method for decoding video from a data stream encoded using binarization of the horizontal and vertical components of a motion vector difference, wherein the binarization is equal to a combination of a prefix in the form of a shortened unary code for the horizontal and vertical components in a first interval of regions of the horizontal and vertical components lower than a cutoff value, and a suffix in the form of an Exp-Golumb code for the horizontal and vertical components in a second interval of regions of the horizontal and vertical components equal to or higher than the cutoff value, wherein the cutoff value is 2 and the Exp-Golumb code has order 1. For the horizontal and vertical components of the motion vector difference, a step of extracting the abbreviated unary code from the data stream using context-adaptive binary entropy decoding, which has exactly one context for each bin position of the abbreviated unary code, common to Exp-Golomb codes using a constant equiprobability bypass mode to obtain the horizontal and vertical components of the motion vector difference and the binarization of the motion vector difference; A step of de-binarizing the binarization of the motion vector difference syntax elements in order to obtain integer values of the horizontal and vertical components of the motion vector difference; A method comprising the step of reconstructing a video based on integer values of the horizontal and vertical components of a motion vector difference.
31. An encoder for encoding video into a data stream, The steps include: predictively encoding the video using motion vectors by motion compensation prediction; predictively encoding the motion vectors by predicting the motion vectors; and setting integer values for the horizontal and vertical components of the motion vector difference that represent the prediction error of the predicted motion vectors; A step of binarizing an integer value to obtain binarization of the horizontal and vertical components of a motion vector difference, wherein the binarization is equal to a combination of a prefix in the form of a shortened unary code for the horizontal and vertical components in a first interval of the regions of the horizontal and vertical components lower than the cutoff value, and a suffix in the form of an Exp-Golumb code for the horizontal and vertical components in a second interval of the regions of the horizontal and vertical components equal to or higher than the cutoff value, wherein the cutoff value is 2 and the Exp-Golumb code has order 1; and An encoder comprising the step of encoding a shortened unary code into a data stream using context-adaptive binary entropy coding, where the horizontal and vertical components of the motion vector difference and the Exp-Golom code using a constant equiprobability bypass mode have in common for each bin position of the shortened unary code.
32. A computer program having program code for execution when the method of either claim 30 or claim 31 is performed on a computer.