Context-based entropy encoding of sample values of spectral envelope
By integrating spectral time prediction with context-based entropy coding of residuals, the encoding of spectral envelope sample values is optimized, addressing inefficiencies in existing speech coders and enhancing encoding performance through adaptive context selection.
Patent Information
- Application Number
- JP2025064854
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2013-10-18
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-30
- Estimated Expiration
- 2034-07-15
AI Technical Summary
Existing speech coders face inefficiencies in encoding the sample values of the spectral envelope, particularly due to the complexity of context selection and quantization required for the spectral time-domain coefficients, leading to high overhead and suboptimal entropy coding performance.
Combining spectral time prediction with context-based entropy coding of residuals, where the context for the current sample value is determined by the deviation between adjacent encoded/decoded sample values, effectively reducing spectral time correlation and improving entropy coding efficiency.
This approach results in a compact prediction residual distribution, reducing overhead and enhancing entropy coding efficiency by adapting contexts based on spectral time correlation, thus improving encoding performance.
Smart Images

Figure 2025111519000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to context-based entropy coding of sample values of a spectral envelope and its use in speech coding / compression.
Background Art
[0002] For example, as described in [1] and [2], many modern state-of-the-art irreversible speech coders are based on the MDCT transform and use both decorrelation and redundancy reduction to minimize the bitrate required for a given perceptual quality. Decorrelation generally exploits the perceptual limitations of the human auditory system to reduce the display accuracy or to reduce frequency information that is not perceptually relevant. Redundancy reduction is generally applied using a statistical model associated with entropy coding to exploit statistical structures or correlations to achieve a minimum compact representation of the remaining data.
[0003] In particular, parametric coding concepts are used to efficiently code speech content. Using parametric coding, parts of a speech signal, for example parts of its spectrogram, are described using parameters rather than using actual time-domain speech samples or the like. For example, a part of the spectrogram of a speech signal can be synthesized on the decoder side with a data stream consisting of parameters such as, for example, a spectral envelope and further parameters that control the synthesis to adapt the synthesized spectrogram part to the transmitted spectral envelope. This kind of novel technology has as its core codec the spectral band replication (SBR) that encodes and transmits the low-frequency components of the speech signal, but the transmitted spectral envelope is used on the decoding side to spectrally shape / form a spectral replication of the reproduction of the low-frequency band components of the speech signal to synthesize the high-frequency band components of the speech signal.
[0004] As outlined above, the spectral envelope within the framework of the encoding technique is transmitted in the data stream with some appropriate spectral time resolution. In a manner similar to the transmission of the sample values of the spectral envelope, the scaling factors for scaling spectral line coefficients or frequency domain coefficients, such as MDCT coefficients, are transmitted in a somewhat coarser spectral time resolution, coarser than the original spectral line resolution, and with some appropriate spectral time resolution for the examples in the context of the spectrum.
[0005] The fixed Huffman coding table can be used to convey information about the samples that describe the spectral envelope or the scaling factors or the frequency domain coefficients. An improved method is, for example, to use context coding as described in [2] and [3], where the context used to select the probability distribution for encoding the values spans both time and frequency. Individual spectral lines, such as MDCT coefficient values, are the actual projections of the composite spectral lines, and even when the magnitude of the composite spectral line is constant over time, it can appear somewhat random in fact. However, the phase changes from one frame to the next. This requires a very complex scheme of context selection, quantization, and mapping for good results, as described in [3].
[0006] In image coding, the context used is usually two-dimensional across the x and y axes of the image, as described in [4] for example. In image coding, the values exist in a linear region or a power region, for example by using gamma correction. In addition, a single fixed linear prediction can be used in each context as a plane approximation and a basic edge detection mechanism, and the prediction error can be encoded. Parametric Golomb or Golomb-Rice coding can be used to encode the prediction error. Run-length coding is additionally used to compensate for the difficulty of directly encoding ultra-low entropy signals with less than 1 bit per sample, for example using a bit-based encoder.
[0007] However, despite improvements related to the scaling factor and / or the encoding of the spectral envelope, there is still a need for an improved concept for encoding the sample values of the spectral envelope. Accordingly, it is an object of the present invention to provide a concept for the encoded spectral values of the spectral envelope.
[0008] This object is achieved by the subject matter of the independent claims in question.
[0009] The embodiments described herein are based on the discovery that an improved concept for the encoded sample values of the spectral envelope can be obtained by combining, on the one hand, spectral time prediction and, on the other hand, context-based entropy coding of the residuals, and in particular determining a context for the current sample value that depends on the amount of deviation between pairs of already encoded / decoded sample values of the spectral envelope in the spectral time vicinity of the current sample value. The combination with context-based entropy coding of the prediction residuals, which involves selecting a context depending on the spectral time prediction on the one hand and the deviation amount on the other hand, is in harmony with the nature of the spectral envelope. The smoothness of the spectral envelope results in a compact prediction residual distribution such that the spectral time correlation is almost completely removed after prediction and can be ignored in context selection for the entropy coding of the prediction result. This then reduces the overhead for managing the context. The use of the amount of deviation between already encoded / decoded sample values in the spectral time vicinity of the current sample value, however, still enables the provision of context adaptability that improves the entropy coding efficiency in a manner that justifies the additional overhead caused thereby.
[0010] According to the embodiments described below, linear prediction is combined with the use of difference values as the amount of deviation, thereby keeping the overhead for encoding low.
[0011] According to the embodiments, the positions of the already encoded / decoded sample values used to determine the difference value that was last used to select / determine the context are such that they are adjacent to each other, spectrally or temporally, in a manner that they are aligned in a row with the current sample value, i.e., they exist along a single line parallel to the time or spectral axis, and are selected such that when determining / selecting the context, the sign of the difference value is further considered. By this measurement, a kind of "trend" in the prediction residual only significantly increases the overhead for simply managing the context, and can be considered when determining / selecting the context for the current sample value.
[0012] Preferred embodiments of the present application are described below with respect to the drawings:
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
[0014] As a motivation for one kind of the embodiments outlined in the present specification below, which can typically be applied to the encoding of spectral envelopes, some ideas leading to the advantageous embodiments outlined below are presented by way of example using intelligent gap filling (IGF). IGF is a novel method that significantly improves the quality of encoded signals even at very low bitrates. For details, the reader is referred to the following description. In any case, IGF addresses the fact that important parts of the spectrum in the high-frequency region are typically quantized to zero due to insufficient bit allocation. To preserve as much as possible, in the IGF information in the low-frequency region, the fine structure in the higher-frequency region is used as a source for adaptively replacing the target region in the high-frequency region where most is quantized to zero. An important requirement for achieving good perceptual quality is the matching of the decoded energy envelope of the spectral coefficients having that of the original signal. To achieve this, the average spectral energy is calculated based on the spectral coefficients from one or more continuous AAC scaling coefficient bands. Calculating the average energy using the boundaries defined by the scaling coefficient bands is motivated by the careful adjustment already present at those boundaries up to a part of the important band, which is characteristic of human hearing. The average energy is converted to a dB scale representation using a formula similar to the one for AAC scaling coefficients and then uniformly quantized. In IGF, different quantization precisions can be arbitrarily used depending on the total bitrate required. Since the average energy constitutes an important part of the information generated by IGF, its efficient representation is highly important for the overall performance of IGF.
[0015] Therefore, in the IGF, the scaling factor energy describes the spectral envelope. The scaling factor energies (SFE) indicate that the spectral values describe the spectral envelope. When decoding the same, the special properties of the SFE can be utilized. In particular, in contrast to [2] and [3], the SFE represents the average value of the MDCT spectral lines, and thus, those values are understood to be much "smoother" and linearly correlated to the average magnitude of the corresponding composite spectral lines. Taking advantage of this situation, the following examples use a combination of context-based entropy coding of prediction residuals that uses, on the one hand, spectral envelope sample value prediction and, on the other hand, a context that depends on the amount of deviation of pairs of adjacent already encoded / decoded sample values of the spectral envelope. The use of this combination is particularly suitable for this type of data to be encoded, i.e., the spectral envelope.
[0016] To facilitate the understanding of the examples outlined further below, FIG. 1 shows the spectral envelope 10 and its components from sample values 12 that take samples of the spectral envelope 10 of an audio signal with a specific spectral time resolution. In FIG. 1, the sample values 12 are exemplarily arranged along the time axis 14 and the spectral axis 16. Each sample value 12 describes or defines the height of the spectral envelope 10 within a corresponding spatio-temporal tile that covers, for example, a specific rectangle in the spatio-temporal region of the spectrogram of the audio signal. The sample values are thus integrated values obtained by integrating the spectrogram over its associated spectral time tile. The sample values 12 can measure the height or strength of the spectral envelope 10 with respect to energy or some other physical quantity and can be defined in the non-logarithmic or linear domain or in the logarithmic domain, which can further provide an additional effect for its characteristic of additionally smoothing the sample values along axes 14 and 16, respectively.
[0017] For the following description, it should be noted that only the sample values 12 being spectrally and temporally regularly arranged, i.e., the corresponding spatio-temporal tile corresponding to the sample value 12 covering the frequency band 18 periodically from the spectrogram of the audio signal, this kind of regularity is assumed for the convenience of explanation not to be obligatory. Rather, an irregular sampling of the spectral envelope 10 by the sample values 12 can also be used. And each sample value 12 represents the average of the height of the spectral envelope 10 within its corresponding spatio-temporal tile. Furthermore, the neighborhood definition outlined below can nevertheless be transferred to this kind of another embodiment of the irregular sampling of the spectral envelope 10. A short statement regarding this kind of possibility is provided below.
[0018] Previously, however, it should be noted that the above-described spectral envelope can be subject to encoding and decoding for transmission from the encoder to the decoder for various reasons. For example, the spectral envelope can be used for scalability purposes, for example, to extend the core encoding of the low-frequency band of the audio signal, i.e., to extend the low-frequency band towards a higher frequency, i.e., the high-frequency band regarding the spectral envelope. In that case, for example, the context-based entropy decoder / encoder described later can be part of, for example, an SBR decoder / encoder. Alternatively, the same can be part of an audio encoder / decoder using IGF as already described above. In IGF, the high-frequency part of the audio signal spectrogram can be additionally described using the spectral values describing the spectral envelope of the high-frequency part of the spectrogram in order to fill the zero-quantized region of the spectrogram within the range of the high-frequency part using the spectral envelope. Details regarding this point are further described below.
[0019] FIG. 2 shows a context-based entropy encoder for encoding sample values 12 of a spectral envelope 10 of an audio signal according to an embodiment of the present application.
[0020] The context-based entropy encoder of FIG. 2 is typically shown using reference numeral 20 and includes a predictor 22, a context determiner 24, an entropy encoder 26, and a residual determiner 28. The context determiner 24 and the predictor 22 have inputs that access the sample values 12 of the spectral envelope (FIG. 1) as described above. The entropy encoder 26 has a control input connected to the output of the context determiner 24 and a data input connected to the output of the residual determiner 28. The residual determiner 28 has two inputs, one of which is connected to the output of the predictor 22 and the other of which provides the residual determiner 28 access to the sample values 12 of the spectral envelope 10. In particular, the residual determiner 28 receives the sample value x to be currently encoded at its input, while the context determiner 24 and the predictor 22 receive at their inputs the sample values 12 that have already been encoded and are present within the spectral time neighborhood of the current sample value x.
[0021] TIFF2025111519000002.tif116169
[0022] As already outlined above, although the sample values 12 are assumed to be regularly arranged along the time and spectral axes 14 and 16, this regularity is not obligatory, and the definition of the neighborhood and the identification of adjacent sample values can be extended to such irregular cases. For example, an adjacent sample value "a" can be defined as adjacent to the upper left corner of the spectral time tile of the current sample along the time axis that is temporally preceding in the upper left corner. A similar definition can be used to define other adjacencies, for example, the adjacency b for e.
[0023] As will be outlined in more detail below, predictor 22 may use different subsets of all sample values within the spectral time neighborhood, i.e., a subset of {a, b, c, d, e}, depending on the spectral time position of the current sample value x. Which subset is actually used may depend, for example, on the availability of adjacent sample values within the spectral time neighborhood defined by the set {a, b, c, d, e}. The adjacent sample values a, d, and c may not be used, for example, for the current sample value x that directly follows the point in time at which the decoder is enabled to start decoding such that dependence on previous portions of the spectral envelope 10 is prohibited / forbidden, i.e., a random access point. Alternatively, the adjacent sample values b, c, and e may not be used for the current sample value x representing the low frequency end of interval 18 such that the position of each adjacent sample value lies within the outer interval 18. In any case, predictor 22 may spectrally predict the current sample value x by linearly combining already encoded sample values within the spectral time neighborhood.
[0024] TIFF2025111519000003.tif65169
[0025] TIFF2025111519000004.tif85169
[0026] As an intermediate note, it must be stated that the definition of the spectral time neighborhood can be adapted to the encoding / decoding order in which the context-based entropy encoder 20 sequentially encodes the sample value 12. As shown in FIG. 1, for example, the context-based entropy encoder can be configured to sequentially encode the sample value 12 using a decoding order 30 that traverses the sample value 12 at each time, proceeding from the lowest frequency to the highest frequency. Hereinafter, "time" is shown as "frame", however, time can alternatively be referred to as a time slot, a time unit, etc. In any case, when using this kind of spectral traversal before temporal feed-forward, the definition of the spectral time neighborhood that extends to the previous time and towards the lower frequencies provides the greatest possibility that the corresponding sample values have already been encoded / decoded and can be utilized. In this case, if they exist, the values within the neighborhood are always already encoded / decoded, however, this can be different for other neighborhood and decoding order pairs. Of course, the decoder uses the same decoding order 30.
[0027] The sample value 12 can represent the spectral envelope 10 in the logarithmic domain as already shown above. In particular, the spectral value 12 can already be quantized to integer values using a logarithmic quantization function. Therefore, for quantization, the deviation amount determined by the context determiner 24 can essentially already be an integer. This is the case, for example, when using a difference as the deviation amount. Regardless of the inherent integer nature of the deviation amount measured by the context determiner 24, the context determiner 24 can subordinate the deviation amount to quantization and determine the context using the quantized amount. In particular, as outlined below, the quantization function used by the context determiner 24 can be constant for values of the deviation amount outside a predetermined interval, for example, a predetermined interval including zero.
[0028] 3 exemplarily shows such a quantization function 32 that maps unquantized deviation magnitudes to quantized deviation magnitudes, in this example the just-mentioned predetermined interval 34 extending from -2.5 to 2.5, with unquantized deviation magnitude values larger than the interval always being mapped to a quantized deviation magnitude value of 3 and unquantized deviation magnitude values smaller than the interval 34 always being mapped to a quantized deviation magnitude value of -3. Thus, only seven contexts should be distinguished and supported in a context-based entropy coder. In the example outlined below, as just illustrated, the length of the interval 34 is 5 and the cardinality of the set of possible values of the spectral envelope sample values is 2. n (e.g., =128), i.e., greater than 16 times the length of the interval. As will be explained later, for the escape coding used, the range of possible values of the spectral envelope sample values is [0;2 n [where n is 2 n+1 is an integer selected to be smaller than the base of the codable values of the prediction residual values, which is 311, according to a particular embodiment described below.
[0029] TIFF2025111519000005.tif113170
[0030] For completeness, FIG. 2 already shows that a quantizer 36 may be connected before the input of the residual determiner 28 from which the current sample value x arrives to obtain the current sample value x, e.g., as outlined above, with a logarithmic quantization function applied to the unquantized sample value x.
[0031] FIG. 4 shows a context-based entropy decoder according to an embodiment, which is compatible with the context-based entropy encoder of FIG.
[0032] TIFF2025111519000006.tif116170
[0033] The entropy decoder 46 performs an inverse transformation of the entropy encoding executed by the entropy encoder 26. That is, the entropy decoder also manages a number of contexts and, for the current sample value x, uses the context selected by the context determiner 44. Each context has an associated corresponding probability distribution that assigns to each possible value of a specific probability r the same as that selected by the context determiner 24 for the entropy encoder 26.
[0034] When using arithmetic coding, the entropy decoder 46, for example, reverses the interval subdivision sequence of the entropy encoder 26. The internal state of the entropy decoder 46 is defined, for example, by the probability interval width of the current interval, and the offset value indicates a sub-interval from the same above where the actual value of r of the current sample value x corresponds within the current probability interval. The entropy decoder 46 uses the arriving arithmetic coding bit stream output by the entropy encoder 26 to update the probability interval and the offset value, for example, by a renormalization process, and examines the offset value to obtain the actual value of r by confirming the sub-interval to which the same applies.
[0035] As previously described, it may be beneficial to limit the entropy encoding of the residual value to some small sub-intervals of the possible values of the prediction residual r. FIG. 5 shows a modified example of the context-based entropy encoder of FIG. 2 to achieve this. In addition to the elements shown in FIG. 2, the context entropy encoder of FIG. 5 consists of controls connected between the residual determiner 28 and the entropy encoder 26, that is, the control 60, as well as an escape coding handler 62 controlled via the control 60.
[0036] TIFF2025111519000007.tif130170
[0037] In the case of an initial prediction residual r that exists within interval 68, control 60 causes entropy encoder 26 to directly entropy-encode this initial prediction residual r. No special measures are to be taken. However, if r exists outside interval 68, as provided by residual determiner 28, an escape encoding procedure is initialized by control 60. In particular, the directly adjacent values that are directly adjacent to interval boundaries 70 and 72 of interval 68 may, according to one embodiment, belong to the symbol alphabet of entropy encoder 26 and function as the escape code itself. That is, as indicated by parentheses 74, the symbol alphabet of entropy encoder 26 includes all values of interval 68 and its directly adjacent values below and above that interval 68, and control 60, in the case of a residual value r greater than upper limit 72 of interval 68, simply decreases the value to be entropy-encoded until it reaches the largest alphabet value 76 that is directly adjacent to upper limit 72 of interval 68, and if the initial prediction residual r is less than the lower limit of interval 68, sends the smallest alphabet value 78 that is directly adjacent to lower limit 70 of interval 68 to entropy encoder 26.
[0038] TIFF2025111519000008.tif136170
[0039] Obviously, escape encoding is not more complex than encoding of normal prediction residuals that exist within interval 68. Context adaptation, for example, is not used. Rather, encoding of the values encoded in the case of escape can simply be performed by directly describing the binary representation, the binary representation for values such as |r| plus x. However, interval 68 is preferably chosen such that the escape procedure occurs statistically rarely and simply represents statistical "outliers" of the sample values x.
[0040] FIG. 7 shows a modified example of the context-based entropy decoder of FIG. 4 and corresponds to, or is adapted to, the entropy encoder of FIG. 5. Similar to the entropy encoder of FIG. 5, the context-based entropy decoder of FIG. 7 is different from that shown in FIG. 4 in that the control 71 is connected between the entropy decoder 46 on one hand and the combiner 48 on the other hand. The entropy decoder of FIG. 7 further includes an escape code handler 73. Similar to FIG. 5, the control 71 performs a check 74 as to whether the entropy decoded value r output by the entropy decoder 46 exists within the interval 68 or corresponds to some escape code. If the latter situation applies, the escape code handler 73 is triggered by the control 71 to extract from the data stream that also carries the entropy encoded data stream entropy decoded by the entropy decoder 46, and the aforementioned code may represent, for example, a self-sufficient mode independent of the escape code indicated by the entropy decoded value r or a real prediction residual r in a mode dependent on the real escape code as already explained in connection with FIG. 6. A binary representation of sufficient bit length is inserted by the escape code handler 62. For example, when the escape code handler 73 reads the binary representation of the value from the data stream, it adds the same to the absolute value of the escape code, i.e., the absolute value of the upper or lower limit, respectively, and uses the sign of each boundary, i.e., the plus sign for the upper limit and the minus sign for the lower limit, as the sign of the read value. Conditional coding may be used. That is, if the entropy decoded value r output by the entropy decoder 46 is located outside the interval 68, the escape code handler 73 may first read, for example, the p-bit absolute value from the data stream, and the same is 2 pIt can be verified whether it is -1. Otherwise, the entropy decoding value r is updated by adding the p-bit absolute value to the entropy decoding value r when the escape code is 72 at the upper limit, and by subtracting the p-bit absolute value from the entropy decoding value r when the escape code is 70 at the lower limit. However, when the p-bit absolute value is 2 p -1, the other q-bit absolute values are read from the bit stream, and when the escape code is 72 at the upper limit, the entropy decoding value r is updated by adding q-bit absolute value + 2 p -1 to the entropy decoding value r, and when the escape code is 70 at the lower limit, the entropy decoding value r is updated by subtracting p-bit absolute value + 2 p -1 from it.
[0041] However, FIG. 7 also shows other variations. According to this variation, in the case of the escape code, the escape code procedure realized by the escape code handlers 62 and 72 directly encodes the complete sample value x so that the estimated value is more than necessary. For example, 2 n bit representation may be sufficient in that case and can indicate the value of x.
[0042] Note that as a precaution only, other ways of implementing escape coding are similarly possible by these other embodiments by not entropy decoding anything for the spectral values. And the prediction residue is beyond or outside the interval 68. For example, for each syntax element, a flag can be transmitted indicating whether the same above is encoded using entropy coding or escape coding is used. In that case, for each sample value, the flag indicates the selected method of coding.
[0043] The following are specific examples for implementing the above embodiments. In particular, the clear examples presented below illustrate a method for handling the above-mentioned difficulty in obtaining certain previously encoded / decoded sample values in the vicinity of the spectral time. Further, the specific examples are shown for setting ranges of possible values 66, intervals 68, quantization function 32, range 34, and others. It will be described later that specific embodiments can be used in connection with IGF. However, note that the explanations presented below can be easily transferred to other cases where the time grid in which the sample values of the spectral envelope are arranged is defined by other time units rather than a frame such as a group of QMF slots, and the spectral resolution is similarly defined by the sub-grouping of sub-bands into spectral time tiles.
[0044] Let the frame number over time be indicated by t (time), and the position of each sample value of the spectral envelope of the entire scale factor (or group of scale factors) be indicated by f (frequency). The sample values are hereinafter referred to as SFE values. We want to encode the value of x using the information already available from the frames already decoded at positions (t - 1), (t - 2), …, and from the current frame at position (t) at frequencies (f - 1), (f - 2), …. The situation is represented again in FIG. 8.
[0045] For an independent frame, we set t = 0. An independent frame is a frame suitable as a random access point for the decoding entity. It thus represents the time when random access to decoding is possible on the decoding side. As far as the spectral axis 16 is concerned, the first SFE 12 associated with the lowest frequency has f = 0. In FIG. 8, the neighbors in time and frequency used for calculating the context (available for both the encoder and the decoder) are like in the cases of a, b, c, d, and e in FIG. 1.
[0046] TIFF2025111519000009.tif19170
[0047] TIFF2025111519000010.tif102130
[0048] TIFF2025111519000011.tif97169
[0049] TIFF2025111519000012.tif70170
[0050] TIFF2025111519000013.tif59170
[0051] Regarding the following figures, various possibilities are described regarding how the above-described context-based entropy encoder / decoder can be incorporated into respective audio decoders / encoders. Figure 9 shows, for example, a parametric decoder 80 into which a context-based entropy decoder 40 according to any of the above-described embodiments can be advantageously incorporated. The parametric decoder 80 consists of a fine structure determiner 82 and a spectral shaper 84 in addition to the context-based entropy decoder 40. Optionally, the parametric decoder 80 consists of an inverse converter 86. The context-based entropy decoder 40 receives an entropy-encoded data stream 88 encoded according to any of the above-described embodiments of the context-based entropy encoder as outlined above. The data stream 88 thus has a spectral envelope encoded therein. The context-based entropy decoder 40 decodes sample values of the spectral envelope of the audio signal that the parametric decoder 80 is to reproduce in the manner outlined above. The fine structure determiner 82 is configured to determine the fine structure of the spectrogram of this audio signal. For this purpose, the fine structure determiner 82 can receive information from an external source, for example, also from other parts of the data stream that also consists of the data stream 88. Further variations are described below. However, in other variations, the fine structure determiner 82 can determine the fine structure alone using a probabilistic or pseudo-probabilistic process. As defined by the spectral values decoded by the context-based entropy decoder 40, the spectral shaper 84 is then configured to shape the fine structure with the spectral envelope. In other words, on the one hand, the input of the spectral shaper 84 is connected to the outputs of the context-based entropy decoder 40 and the fine structure determiner 82, respectively, to receive the spectral envelope therefrom and, on the other hand, to receive the fine structure of the spectrogram of the audio signal, and the spectral shaper 84 outputs, at its output, the fine structure of the spectrogram shaped by the spectral envelope.The inverse converter 86 can perform an inverse conversion on a fine structure shaped to output the reconstruction of the audio signal at its output.
[0052] In particular, the fine determiner 82 may be configured to determine the fine structure of a spectrogram using at least one of artificial random noise generation, spectrum reproduction, and decoding for each spectral line, which uses spectrum prediction and / or spectrum entropy context derivation. The first two possibilities are described with respect to FIG. 10. FIG. 10 illustrates the possibility that the spectral envelope 10 decoded by the context-based entropy decoder 40 forms a higher frequency extension of the lower frequency section 90, i.e., section 18 extends the lower frequency section 90 to a higher frequency, i.e., section 18 is in contact with section 19 on the higher frequency side of the latter. Thus, FIG. 10 shows the possibility that the audio signal to be actually reproduced by the parametric decoder 80 covers a frequency section 92 where section 18 simply represents the high frequency portion of the entire frequency section 92. As shown in FIG. 9, the parametric decoder 80 may additionally include a low frequency decoder 94 configured to decode a low frequency data stream 96 accompanied by a data stream 88 to obtain a low frequency version of the audio signal at its output. The spectrogram of this low frequency version is represented in FIG. 10 using reference numeral 98. In summary, this frequency version 98 of the audio signal and the fine structure shaped within section 18 result in the reproduction of the audio signal of the entire frequency section 92, i.e., of its spectrogram over the entire frequency section 92. As indicated by the dashed line in FIG. 9, the inverse converter 86 may perform an inverse conversion over the entire section 92. In this framework, the fine structure determiner 82 may receive the low frequency version 98 from the decoder 94 in the time domain or the frequency domain. In a first case, the fine structure determiner 82 causes a conversion of the received low frequency version into the spectral domain to obtain a spectrogram 98 and, as illustrated using arrow 100, to obtain the fine structure to be shaped by the spectral shaper 84 by the spectral envelope provided by the context-based entropy decoder 40 that uses spectrum reproduction.However, as already outlined above, the fine structure determiner 82 cannot even receive the low-frequency version of the audio signal from the LF decoder 94, and can merely generate a fine structure that only uses a probabilistic or pseudo-probabilistic process.
[0053] The corresponding parametric encoder adapted to the parametric decoder according to FIGS. 9 and 10 is represented in FIG. 11. The parametric encoder of FIG. 11 includes a frequency crossover 110 receiving the audio signal 112 to be encoded, a high-band encoder 114, and a low-band encoder 116. The frequency crossover 110 decomposes the inbound audio signal 112 into two components, namely a first signal 118 corresponding to the high-pass filtered version of the inbound audio signal 112, and a low-frequency signal 120 corresponding to the low-pass filtered version of the inbound audio signal 112. The frequency bands covered by the high-frequency signal 118 and the low-frequency signal 120 are adjacent to each other at several crossover frequencies (compare with 122 in FIG. 10). The low-band encoder 116 receives the low-frequency signal 120 and encodes the same into a low-frequency data stream, i.e., 96. And the high-band encoder 114 calculates sample values describing the spectral envelope of the high-frequency signal 118 within the high-frequency interval 18. The high-band encoder 114 is also equipped with the above-described context-based entropy encoder for encoding these sample values of the spectral envelope. The low-band encoder 116 may be, for example, a transform encoder, and the spectral time resolution at which the low-band encoder 116 encodes the transform or spectrogram of the low-frequency signal 120 may be greater than the spectral time resolution at which the sample values 12 decompose the spectral envelope of the high-frequency signal 118. Thus, the high-band encoder 114 outputs, in particular, the data stream 88. As indicated by the dashed line 124 in FIG. 11, the low-band encoder 116 may output information to the high-band encoder 114, for example, for controlling the high-band encoder 114 regarding this generation of sample values describing the spectral envelope, or at least regarding the selection of the spectral time resolution at which the sample values take samples of the spectral envelope.
[0054] FIG. 12 shows another possibility of implementing the parametric decoder 80 of FIG. 9 and in particular the fine structure determiner 82. In particular, according to the embodiment of FIG. 12, the fine structure determiner 82 itself receives a data stream and, based thereon, determines the fine structure of an audio signal spectrogram using decoding for each spectral line that uses spectral prediction and / or spectral entropy-context derivation. That is, the fine structure determiner 82 itself recovers the fine structure in the form of a spectrogram from the data stream, for example, from the time sequence of the spectrum of the overlap transform. However, in the case of FIG. 12, the fine structure thus determined by the fine structure 82 is related to the first frequency interval 130 and coincides with the complete frequency interval of the audio signal, i.e., 92.
[0055] In the embodiment of FIG. 12, the frequency interval 18 to which the spectral envelope 10 is related completely overlaps the interval 130. In particular, the interval 18 forms the high-frequency part of the interval 130. For example, many of the spectral lines within the spectrogram 132 are recovered by the fine structure determiner 82 and cover the frequency interval 130, and are quantized to zero, especially within the range of the high-frequency part 18. Nevertheless, in order to reproduce the audio signal with high quality, at a reasonable bit rate, even within the range of the high-frequency part 18, the parametric decoder 80 utilizes the spectral envelope 10. The spectral values 12 of the spectral envelope 10 describe the spectral envelope of the audio signal within the range of the high-frequency part 18 with a spectral time resolution coarser than the spectral time resolution of the spectrogram 132 decoded by the fine structure determiner 82. For example, the spectral time resolution of the spectral envelope 10 is coarser in terms of spectral terms, i.e., its spectral resolution is coarser than the spectral line accuracy of the fine structure 132. As described above, spectrally, the sample values 12 of the spectral envelope 10 can describe the spectral envelope 10 in the frequency bands 134 in which the spectral lines of the spectrogram 132 are classified for scaling in the direction of the scaling factor of the spectral line coefficients.
[0056] Spectrum shaper 84 then uses the sample value 12 to fill in spectral lines within the range of a spectral line group or spectral time tile corresponding to each sample value 12 using a mechanism such as spectral reproduction or artificial noise generation, and adjusts the fine structure level or energy occurring within each spectral time tile / scaling factor group according to the corresponding sample value that describes the spectral envelope. See, for example, FIG. 13. FIG. 13 illustrates the spectrum from a spectrogram 132 corresponding to one frame or that time, for example time 136 of FIG. 12. The spectrum is illustrated using reference numeral 140. As illustrated in FIG. 13, some of its portions 142 are quantized to zero. FIG. 13 shows the high-frequency portion 18 and the subdivision of the spectral lines of the spectrum 140 into scaling factor bands indicated by the parentheses. Using "x" and "b" and "e", FIG. 13 illustrates that three sample values 12 describe the spectral envelope within the high-frequency portion 18 at time 136 - one for each scaling factor band. Within the range of each scaling factor band corresponding to these sample values e, b, and x, the fine structure determiner 82 generates fine structure, for example by spectral reproduction from the lower-frequency portion 146 of a complete frequency interval 130, within at least the zero-quantized portion 142 of the spectrum 140 as indicated by the hatched region 144, and adjusts the energy resulting from the spectrum by scaling the artificial fine structure 144 according to or using the sample values e, b, and x.Interestingly, there is a non-zero quantized portion 148 of the spectrum 140 within the range of the scaling coefficient band of the intermediate or high-frequency portion 18, and thus, using intelligent gap filling according to FIG. 12, it is possible to place peaks within the range of the spectrum 140 even in the high-frequency portion 18 of the complete frequency interval 130 with spectral resolution and at any spectral line position, and nevertheless, there is an opportunity to satisfy the zero-quantized portion 142 using sample values x, b, and e for shaping the fine structure inserted within the range of these zero-quantized portions 142.
[0057] Finally, when implemented in accordance with the descriptions of FIGS. 12 and 13, FIG. 14 shows a possible parametric encoder for powering the parametric decoder of FIG. 9. In particular, in that case, the parametric encoder may include a transducer 150 configured to spectrally decompose the inbound audio signal 152 into a complete spectrogram covering the complete frequency interval 130. An overlapping transform with a variable transform length may be used. The spectral line encoder 154 encodes this spectrogram with spectral line resolution. To achieve this purpose, the spectral line encoder 154 receives both the high-frequency portion 18 and the remaining low-frequency portion from the transducer 150 so as to cover the complete frequency interval 130 without gaps and without overlapping. The parametric high-frequency encoder 156 simply receives the high-frequency portion 18 of the spectrogram 132 from the transducer 150 and generates at least the data stream 88, i.e., sample values describing the spectral envelope within the range of the high-frequency portion 18.
[0058] That is, according to the embodiments of FIGS. 12 to 14, the spectrogram 132 of the audio signal is encoded into the data stream 158 by the spectral line encoder 154. Therefore, the spectral line encoder 154 can encode one spectral line value for each spectral line of the complete section 130 for each time or frame 136. The small boxes 160 in FIG. 12 indicate these spectral line values. Along the spectral axis 16, the spectral lines can be classified into scaling coefficient bands. In other words, the frequency interval 16 can be subdivided into scaling coefficient bands consisting of groups of spectral lines. The spectral line encoder 154 can select a scaling coefficient for each scaling coefficient band at each time in order to scale the quantized spectral line values 160 encoded via the data stream 158. With a spectral time resolution that is at least coarser than the spectral time grid defined by the times and spectral lines at which the spectral line values 160 are regularly arranged and that can coincide with the raster defined by the scale coefficient resolution, the parametric high-frequency encoder 156 describes the spectral envelope within the range of the high-frequency portion 18. Interestingly, the non-zero quantized spectral line values 160 are scaled by the scaling coefficients of the scaling coefficient bands into which they fall, can be scattered at any position within the range of the high-frequency portion 18 at the spectral line resolution, and thus, for example, the fine structure determiner 82 and the spectral shaper 84 limit their fine structure synthesis and shaping, for example, to the zero-quantized portion 142 within the range of the high-frequency portion 18. Within the range of the spectral shaper 84 that uses the sample values that describe the spectral envelope within the range of the high-frequency portion, high-frequency synthesis occurs on the decoding side. Eventually, a very effective compromise occurs between the bit rate consumed on the one hand and the quality available on the other hand.
[0059] As indicated by the dashed arrow in FIG. 14, shown as 164, the spectral line encoder 154 can notify the parametric high-frequency encoder 156 regarding, for example, a reconstructable version of the spectrogram 132 as being reconstructable from the data stream 158, and the parametric high-frequency encoder 156 uses this information, for example, to control the generation of the sample value 12 by means of the spectral time resolution of the representation of the sample value 12 and / or the spectral envelope 10.
[0060] Summarizing the above, the above embodiments utilize special characteristics of sample values of the spectral envelope. Here, in contrast to [2] and [3], this type of sample value represents the average value of the spectral lines. In all the embodiments outlined above, the conversion can use the MDCT, and thus, the inverse MDCT can be used for all inverse conversions. In any case, this type of sample value of the spectral envelope is much "smoother" and is linearly related to the average value of the corresponding composite spectral lines. Additionally, according to at least some of the above embodiments, the sample values of the spectral envelope, hereinafter referred to as SFE values, are in the actual dB domain or more generally in the logarithmic function domain, which is a logarithmic function representation. This further improves the "smoothness" compared to the values in the linear or power domain for the spectral lines. For example, in AAC, the power exponent is 0.75. In contrast to [4], in at least some embodiments, the spectral envelope sample values exist in the logarithmic function domain, and the characteristics and the structure of the coding distribution are significantly different (depending on its magnitude, one logarithmic function domain value generally maps to an exponentially increasing number of linear domain values). Thus, at least some of the above-described embodiments utilize the logarithmic function representation in the quantization of the context (a smaller number of contexts generally exist) and in encoding the tails of the distribution in each context (the tail of each distribution is wider). In contrast to [2], some of the above embodiments further use a fixed or adaptive linear prediction in each context, as used when calculating the quantized context, based on the same data. Still, this method helps to significantly reduce the number of contexts while obtaining optimal performance. For example, in contrast to [4], in at least some of the embodiments, the linear prediction in the logarithmic function domain has significantly different uses and importance. For example, it is possible to completely predict both the constant energy spectral region and furthermore the fade-in and fade-out spectral regions of the signal.In contrast to [4], some of the above-described embodiments use arithmetic coding that enables optimal coding of an arbitrary distribution to use information extracted from a representative training data set. In contrast to [2] which also uses arithmetic coding, according to the embodiment, it is the prediction error value, rather than the original value, that is encoded. Further, in the embodiment, bitplane coding need not be used. Bitplane coding, however, requires several arithmetic coding steps for each integer value. In contrast, according to the embodiment, each sample value of the spectral envelope can be encoded / decoded within a range including one step that selectively uses escape coding for values outside the center of the entire sample value distribution, as described above, and it is very fast.
[0061] As described above with respect to FIGS. 9, 12, and 13, to briefly summarize an example of a parameter decoder that supports IGF again, according to this example, the fine structure determiner 82 uses spectral prediction and / or spectral entropy context derivation to derive the fine structure 132 of the spectrogram of the audio signal in the first frequency interval 130, i.e., within the complete frequency interval, for each spectral line for decoding. The per-frequency-line decoding shows the fact that the fine structure determiner 82 receives spectral line values 160 from a data stream arranged at the spectral row pitch spectrally, thereby forming a spectrum 136 for each time instance corresponding to each time portion. The use of spectral prediction may include, for example, differential coding of these spectral line values along the spectral axis 16, i.e., only the difference with respect to the spectrally immediately preceding spectral line value is decoded from the data stream and added to this preceding value. Spectral entropy-context derivation may mean the fact that the context for entropy decoding each spectral line value 160 may depend on the already decoded spectral line values in the spectral time vicinity, or at least in the spectral vicinity, of the currently decoded spectral line value 160, i.e., may be additionally selected based on the already decoded spectral line values. To fill the zero-quantized portion 142 of the fine structure, the fine structure determiner 82 may use artificial random noise generation and / or spectral reproduction. The fine structure determiner 82 may simply perform this, for example, within a second frequency interval 18 that may be limited to the high-frequency portion of the overall frequency interval 130. The spectrally reproduced portion may be obtained, for example, from the remaining frequency portion 146. The spectrum shaper then performs the shaping of the fine structure thus obtained according to the spectral envelope described by the sample values 12 in the zero-quantized portion. In particular, the contribution of the shaped fine structure of the non-zero quantized portion of the fine structure within the interval 18 to the result of the fine structure is independent of the actual spectral envelope 10.This means the following: namely, in the final fine structure spectrum, only the portions 142 are filled by spectrum reproduction using artificial random noise generation and / or spectrum envelope shaping, and the non-zero contributions 148 that remain are interspersed between the portions 142 such that the artificial random noise generation and / or spectrum reproduction, i.e., filling, is restricted to the exactly zero quantization portions 142, or all artificial random noise generation and / or spectrum generation occur alternately, i.e., form a fine structure synthesized by the spectrum envelope 10, and each synthesized fine structure is placed on the portion 148 in an additional manner. However, even in that case, the contribution as the non-zero quantized portion 148 of the original decoded fine structure is maintained.
[0062] Regarding the embodiments of FIGS. 12 - 14, one should ultimately note that the IGF (Intelligent Gap Filling) procedures or concepts described with respect to these figures significantly improve the quality of the encoded signal even at very low bitrates. An important portion of the spectrum in the high - frequency region 18 is typically quantized to zero due to insufficient bit allocation. To preserve the fine structure of the higher - frequency region 18, the IGF information, the low - frequency region is used as a source to adaptively replace the target region of the high - frequency region, which is mostly zero, i.e., region 142. An important requirement for achieving good perceptual quality is the matching of the decoded energy envelope of the spectral coefficients having that of the original signal. To achieve this, the average spectral energy is calculated over the spectral coefficients from one or more continuous AAC scaling coefficient bands. The resulting value is the sample value 12 that describes the spectral envelope. Calculating the average using the boundaries defined by the scaling coefficient bands is motivated by the existing careful tuning of those boundaries up to a part of the critical band, which is characteristic of human hearing. As described above, the average energy can be transformed into a logarithmic, e.g., dB - scale representation using an equation that can be similar to those already known in AAC scaling coefficients and uniformly quantized. In IGF, different quantization precisions can be arbitrarily used depending on the required total bitrate. The average energy constitutes an important part of the information generated by IGF, and thus its efficient representation within the data stream 88 is extremely important for the overall performance of the IGF concept.
[0063] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent descriptions of corresponding method steps, where a block or apparatus corresponds to a method step or a feature of a method step. Similarly, aspects described in the context of a method step represent descriptions of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0064] Depending on specific implementation requirements, embodiments of the present invention may be implemented in hardware or in software. The embodiments may be executed using a digital storage medium having electronically readable control signals stored thereon, such as, for example, a floppy (registered trademark) disk, a hard disk, a DVD, a Blu-Ray (registered trademark), a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory. And it cooperates (or can cooperate) with a programmable computer system so that each method is executed. Accordingly, the digital storage medium may be computer-readable.
[0065] Some embodiments according to the present invention include a data carrier having an electronically readable control signal that can cooperate with a programmable computer system to execute one of the methods described herein.
[0066] Typically, embodiments of the present invention may be implemented as a computer program product having program code, and when the computer program product runs on a computer, the program code operates to execute one of the methods. The program code may be stored, for example, on a machine-readable carrier.
[0067] Other embodiments include a computer program for performing one of the methods described in the present specification and stored on a machine-readable carrier.
[0068] In other words, an embodiment of the method of the present invention is thus a computer program having program code for performing one of the methods described in the present specification when the computer program is executed on a computer.
[0069] A further embodiment of the method of the present invention is thus a data carrier (or digital storage medium or computer-readable medium) including a computer program recorded thereon for performing one of the methods described in the present specification. The data carrier, digital storage medium or recording medium is typically tangible and / or non-transitory.
[0070] A further embodiment of the method of the present invention is thus a data stream or series of signals representing a computer program for performing one of the methods described in the present specification. The data stream or series of sequences can be configured, for example, to be transferred via a data communication connection, such as the Internet.
[0071] A further embodiment includes processing means, such as a computer or programmable logic device, configured or adapted to perform one of the methods described in the present specification.
[0072] A further embodiment includes a computer having installed thereon a computer program for performing one of the methods described in the present specification.
[0073] Further embodiments according to the present invention include an apparatus or system configured to transfer (e.g., electronically or optically) to a receiver a computer program for executing one of the methods described in the present specification. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may include, for example, a file server for transferring the computer program to the receiver.
[0074] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to execute some or all of the functions of the methods described in the present specification. In some embodiments, a field programmable gate array may cooperate with a microprocessor to execute one of the methods described in the present specification. Generally, the method is preferably executed by any hardware device.
[0075] The above-described embodiments are merely illustrative for the principles of the present invention. Modifications and details described in the present specification are understood to be obvious to those skilled in the art. Accordingly, it is intended that the present invention be limited only by the imminent scope of the patent claims and not by the specific details shown by the description of the specification and the embodiments.
[0076] References [1] International Standard ISO / IEC 14496-3:2005, Information technology - Coding of audio-visual objects - Part 3: Audio, 2005. [2] International Standard ISO / IEC 23003-3:2012, Information technology - MPEG audio technologies - Part 3: Unified Speech and Audio Coding, 2012. [3] B. Edler and N. Meine: Improved Quantization and Lossless Coding for Subband Audio Coding, AES 118th Convention, May 2005. [4] M.J. Weinberger and G. Seroussi: The LOCO-I Lossless Image Compression Algorithm: Principles and Standardization into JPEG-LS, 1999. Available online at http: / / www.hpl.hp.com / research / info_theory / loco / HPL-98-193R1.pdf
Claims
1. A context-based entropy decoder for decoded sample values (12) of a spectral envelope (10) of an audio signal, comprising: Predicting (42) the current sample value of the spectral envelope spectrally over time to obtain an estimated value of the current sample value; Determining (44) a context for the current sample value that depends on an amount for deviation between pairs of already decoded sample values of the spectral envelope in the spectral-time neighborhood of the current sample value; Entropy decoding (46) a prediction residual value of the current sample value using the determined context; A context-based entropy decoder configured to combine (48) the estimated value and the prediction residual value to obtain the current sample value.
2. The context-based entropy decoder according to claim 1, further configured to perform the spectral-time prediction by linear prediction.
3. The context-based entropy decoder according to claim 1 or 2, further configured to use a signed difference between the pairs of already decoded sample values of the spectral envelope in the spectral-time neighborhood of the current sample value to measure the deviation.
4. The context for the current sample value is further configured to depend on a first amount for deviation between a first pair of already decoded sample values of the spectral envelope in the spectral-time neighborhood of the current sample value and a second amount for deviation between a second pair of already decoded sample values of the spectral envelope in the spectral-time neighborhood of the current sample value, provided that the first pair are spectrally adjacent to each other and the second pair are temporally adjacent to each other, the context-based entropy decoder according to any of the previous claims.
5. The context-based entropy decoder according to claim 4, further configured to spectrally predict the current sample value of the spectral envelope by linearly combining the already decoded sample values of the first and second pairs.
6. If the audio signal is encoded at the bit rate greater than a predetermined threshold, the linear combination coefficients are further configured to be set such that the coefficients are the same for different contexts, and if the bit rate is less than the predetermined threshold, the coefficients are set independently for the different contexts, the context-based entropy decoder according to claim 5.
7. When decoding the sample values of the spectral envelope, at each time instant from the lowest frequency to the highest frequency, the context-based entropy decoder according to any of the previous claims, further configured to sequentially decode the sample values using a decoding order (30) across the sample values for each time instant.
8. When determining the context, the context-based entropy decoder according to any of the previous claims, further configured to quantize the amount for the deviation and determine the context using the quantized amount.
9. The context-based entropy decoder according to claim 8, further configured to be constant for the value of the amount for the deviation outside a predetermined interval (34), the predetermined interval using a quantization function (32) in the quantization of the amount for the deviation including zero.
10. The context-based entropy decoder according to claim 9, wherein the value of the spectral envelope is represented as an integer, and the length of the predetermined interval (34) is less than or equal to 1 / 16 of the number of representable states of the integer representation of the value of the spectral envelope.
11. The context-based entropy decoder according to any of the previous claims, further configured to transfer (50) the current sample value from a logarithmic domain to a linear domain so as to be derived by combination.
12. When entropy decoding the residual value, the context-adaptive entropy decoder according to any of the previous claims, further configured to sequentially decode the sample values along the decoding order and use a set of context-specific probability distributions that are constant while sequentially decoding the sample values of the spectral envelope.
13. A context-based entropy decoder according to any of the previous claims, further configured to use an escape coding mechanism when the residual value is outside a predetermined value range (68) when entropy decoding the residual value. **Claim 14** The context-based entropy decoder according to claim 13, wherein the sample value of the spectral envelope is represented as an integer, the prediction residue is represented as an integer, and the absolute value of the interval boundaries (70, 72) of the predetermined value range is less than or equal to 1 / 8 of the number of displayable states of the prediction residue value. **Claim 15** The parametric decoder comprises: a context-based entropy decoder (40) for decoding sample values of a spectral envelope of an audio signal according to any of the previous claims; a fine structure determiner (82) configured to determine a fine structure of a spectrogram of the audio signal; and a spectral shaper (84) configured to shape the fine structure according to the spectral envelope. **Claim 16** The parametric decoder according to claim 15, wherein the fine structure determiner is configured to determine the fine structure of the spectrogram using at least one of spectral line direction decoding using artificial random noise generation, spectral reproduction, and spectral prediction and / or spectral entropy-context derivation. **Claim 17** The parametric decoder according to claim 15 or 16, further comprising a low frequency interval decoder (94) configured to decode a lower frequency interval (98) of a spectrogram of the audio signal, wherein the context-based entropy encoder, the fine structure determiner, and the spectral shaper are configured such that the shaping of the fine structure by the spectral envelope is performed within a spectral high frequency extension (18) of the lower frequency interval. **Claim 18** The parametric decoder according to claim 17, wherein the low frequency interval decoder (94) is configured to determine the fine structure of the spectrogram using spectral line direction decoding using spectral prediction and / or spectral entropy-context derivation or using spectral decomposition of a time domain low frequency band audio signal decoded using the same. **Claim 19** The microstructure determiner is configured to derive the microstructure of the spectrogram of the audio signal within a first frequency interval (130), to place a zero-quantized portion (142) of the microstructure within a second frequency interval (18) overlapping with the first frequency interval, and to use spectral prediction and / or spectral entropy-context derivation for spectral line direction decoding for applying artificial random noise generation and / or spectral reproduction onto the zero-quantized portion (142), and the spectral shaper (84) is configured to perform the shaping of the microstructure according to the spectral envelope at the zero-quantized portion (142), the parametric decoder according to claim 15 or 16.
20. A context-based entropy encoder for encoding sample values of a spectral envelope of an audio signal, predicts the current sample value of the spectral envelope spectrally and temporally to obtain an estimated value of the current sample value; determines a context for the current sample value that depends on an amount for deviation between pairs of already decoded sample values of the spectral envelope in the spectral time vicinity of the current sample value; determines a prediction residual value based on the deviation between the estimated value and the current sample value; A context-based entropy encoder configured to entropy-encode the prediction residual value of the current sample value using the determined context.
21. A method using context-based entropy decoding for decoding sample values of a spectral envelope of an audio signal, predicting the current sample value of the spectral envelope spectrally and temporally to obtain an estimated value of the current sample value; determining a context for the current sample value that depends on an amount for deviation between pairs of already decoded sample values of the spectral envelope in the spectral time vicinity of the current sample value; entropy-decoding the prediction residual value of the current sample value using the determined context; combining the estimated value and the prediction residual value to obtain the current sample value.
22. A method for encoding sample values of a spectral envelope of an audio signal using context-based entropy coding, comprising: Predicting the current sample value of the spectral envelope spectrally and temporally to obtain an estimated value of the current sample value; Determining a context for the current sample value that depends on an amount for deviation between pairs of already decoded sample values of the spectral envelope in the spectral time vicinity of the current sample value; Determining a prediction residual value based on the deviation between the estimated value and the current sample value; Entropy encoding the prediction residual value of the current sample value using the determined context.
23. A computer program having program code for performing the method according to claim 21 or 22 when running on a computer.
Citation Information
Patent Citations
Efficient spectral envelope coding using variable time / frequency resolution and time / frequency switching.
JP2003529787A
Audio signal synthesizer and audio signal encoder
JP2011527447A
Audio-encoding method and apparatus, audio-decoding method and apparatus, recording medium thereof, and multimedia device employing same
WO2012165910A2