Audio encoder, audio decoder, method for encoding an audio signal, and method for decoding an encoded audio signal
Selective predictive coding in the transform domain for audio signals addresses high complexity and limited gain issues in FDP, enhancing efficiency and reducing errors by focusing on harmonic signal elements.
Patent Information
- Application Number
- JP2022082087
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-06-17
- Filing Date
- 2022-05-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2036-03-07
AI Technical Summary
The frequency-domain prediction (FDP) method in audio encoding requires high computational complexity and has limited prediction gain due to noisy elements between tonal spectral parts, leading to reconstruction errors.
Implement predictive coding selectively in the transform or filter bank domain, applying it only to spectral coefficients or groups based on an interval value, avoiding noisy elements and reducing complexity by synchronizing prediction coefficients between encoder and decoder.
Reduces computational complexity and enhances prediction efficiency by applying predictive coding only to harmonic signal elements, improving audio quality and reducing errors.
Smart Images

Figure 0007710412000001 
Figure 0007710412000002 
Figure 0007710412000003
Abstract
Description
Technical Field
[0001] Embodiments relate to methods and apparatuses for encoding an audio signal using audio encoding, specifically predictive encoding, and for decoding an encoded audio signal using predictive decoding. Preferred embodiments relate to methods and apparatuses for pitch-adaptive spectral prediction. Even more preferred embodiments relate to perceptual encoding of tonal audio signals by transform coding using an inter-frame prediction tool in the spectral domain.
Background Art
[0002] In order to improve the quality of encoded tonal signals, especially at low bitrates, recent audio transform coders have been using very long transforms and / or long-term prediction or pre / post filtering. However, long transforms imply a long algorithmic delay, which is undesirable for low-delay communication scenarios. Therefore, very low-delay predictors based on instantaneous reference pitch have recently gained popularity. The Opus codec of the IETF (Internet Engineering Task Force) utilizes pitch-applied prefiltering and postfiltering in its frequency-domain CELT (Constrained-Energy Lapped Transform) coding path (J.M. Valin, K. Vos, and T. Terriberry, "Definition of the Opus audio codec", Internet Engineering Task Force, Technical Report RFC6716, 2012, http: / / tools.ietf.org / html / rfc67161), and the EVS (Enhanced Voice Services) codec of 3GPP (3rd Generation Partnership Project) provides a long-term harmonic postfilter for the perceptual improvement of the transformed encoded signal (3GPP TS 26.443 "Codec for Enhanced Voice Services (EVS)", Release 12, December 2014). All of these approaches work in the time domain on the fully decoded signal waveform and are difficult to apply selectively in the frequency domain (none of the schemes selectively provides anything more than a simple low-pass filter for some frequencies) and / or computationally expensive. A welcomed alternative to time-domain long-term prediction (LTP) or pre / post filtering (PPF) is, as a result, provided by frequency-domain prediction (FDP) as supported in MPEG-2 AAC (ISO / IEC 13818-7 "Information technology - Part 7: Advanced Audio Coding (AAC)", 2006). This method facilitates frequency selectivity but has inherent drawbacks as described below.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] The FDP method introduced above has two drawbacks compared to other tools. First, the FDP method requires high computational complexity. Specifically, at least two linear prediction encodings are applied (i.e., from the channel conversion bins of the last two frames) to hundreds of spectral bins for each frame and channel in the worst case of prediction in all scale factor bands (ISO / IEC 13818-7 "Information technology - Part 7: Advanced Audio Coding (AAC)", 2006). Second, the FDP method includes a limited overall prediction gain. More specifically, noisy elements between the tonal spectral parts of predictable harmonics are also targeted for prediction, and these noisy parts usually cannot be predicted, thus causing errors, so the efficiency of prediction is limited.
[0005] The high complexity is due to the backward adaptability of the predictor. That is, the prediction coefficient for each bin has to be calculated based on the bins transmitted earlier. Therefore, the numerical inaccuracies between the encoder and the decoder can lead to reconstruction errors due to discrepant prediction coefficients. To overcome this problem, bit-exact identical adaptation has to be guaranteed. Further, the adaptation has to be always performed to keep the prediction coefficients up-to-date even when a group of predictors is disabled in a certain frame.
Means for Solving the Problem
[0006] Therefore, it is an object of the present invention to provide a concept for encoding an audio signal and / or decoding an encoded audio signal that avoids at least one (e.g., both) of the aforementioned problems and leads to a more efficient and computationally less expensive embodiment.
[0007] This problem is solved by the independent claims.
[0008] Advantageous embodiments are addressed by the dependent claims.
[0009] An embodiment provides an encoder for encoding an audio signal. The encoder is configured to encode the audio signal in a transform domain or a filter bank domain, the encoder is configured to determine spectral coefficients of the audio signal for a current frame and at least one previous frame, the encoder is configured to selectively apply predictive coding to a plurality of individual spectral coefficients or spectral coefficient groups, the encoder is configured to determine an interval value, and the encoder is configured to select a plurality of individual spectral coefficients or spectral coefficient groups to which predictive coding is applied based on the interval value that can be transmitted as side information together with the encoded audio signal.
[0010] A further embodiment provides a decoder for decoding an encoded audio signal (e.g., encoded by the above encoder). The decoder is configured to decode the encoded audio signal in a transform domain or a filter bank domain, and the decoder is configured to analyze the encoded audio signal to obtain the encoded spectral coefficients of the audio signal for the current frame and at least one previous frame, and the decoder is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients, and the decoder may be configured to select a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients to which predictive decoding is applied based on the transmitted interval value.
[0011] According to the concept of the present invention, predictive coding is applied to (only) the selected spectral coefficients. The spectral coefficients to which predictive coding is applied can be selected according to the signal characteristics. For example, by not applying predictive coding to noisy signal elements, the aforementioned error caused by predicting unpredictable, noisy signal elements is avoided. At the same time, since predictive coding is applied only to the selected spectral elements, the computational complexity can be reduced.
[0012] For example, perceptual coding of tonal audio signals can be performed (e.g., by an encoder) by transform coding together with an inductive / adaptive inter-frame prediction technique in the spectral domain. By applying prediction only to the spectral coefficients around the harmonic signal elements located at integer multiples of the fundamental frequency or fundamental pitch, which can be sent, for example, as interval values in an appropriate bitstream from the encoder to the decoder, the efficiency of frequency domain prediction (FDP) can be increased and the computational complexity can be reduced. Embodiments of the present invention can preferably be implemented or incorporated into an MPEG-H 3D audio codec, but are applicable to any audio transform coding system such as MPEG-2 AAC.
[0013] Further embodiments provide a method of encoding an audio signal in a transform domain or a filter bank domain, the method comprising: determining spectral coefficients of the audio signal for the current frame and at least one previous frame; determining an interval value; selectively applying predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients, wherein the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive coding is applied are selected based on the interval value; and
[0014] Further embodiments provide a method of decoding an encoded audio signal in a transform domain or a filter bank domain, the method comprising: analyzing the encoded audio signal to obtain encoded spectral coefficients of the audio signal for the current frame and at least one previous frame; obtaining an interval value; selectively applying predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients, wherein the plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients to which predictive decoding is applied are selected based on the interval value; and
[0015] Embodiments of the present invention are described herein below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
[0017] Equivalent or corresponding elements, or elements having equivalent or corresponding functionality, are indicated in the following description by equivalent or corresponding reference numerals.
[0018] In the following description, a number of details are set forth in order to provide a more thorough explanation of embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form rather than in detail in order not to obscure embodiments of the present invention. Furthermore, the features of the different embodiments described below may be combined with each other as long as there is no particular notice to the contrary.
[0019] FIG. 1 shows a schematic block diagram of an encoder 100 for encoding an audio signal 102 according to an embodiment. The encoder 100 is configured to encode the audio signal 102 in a transform domain or filter bank domain 104 (e.g., a frequency domain or a spectral domain), and the encoder 100 is configured to determine spectral coefficients 106_t0_f1 to 106_t0_f6 of the audio signal 102 for the current frame 108_t0, and spectral coefficients 106_t-1_f1 to 106_t-1_f6 of the audio signal for at least one previous frame 108_t-1. Further, the encoder 100 is configured to selectively apply predictive coding to a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5, the encoder 100 is configured to determine an interval value, and the encoder 100 is configured to select a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 to which predictive coding is applied, based on the interval value.
[0020] That is, the encoder 100 is configured to selectively apply predictive coding to a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 selected based on a single interval value transmitted as side information.
[0021] This interval value, together with its integer multiples, can correspond to the frequency that defines the center for all spectral coefficient groups to which the prediction is applied (for example, the fundamental frequency of the harmonic tone of (the audio signal 102)), that is, the first group can be around this frequency, the second group can be centered around twice this frequency, the third group can be centered around three times this frequency, and so on. Knowledge of these center frequencies enables the calculation of prediction coefficients for predicting the corresponding sine wave signal components (for example, the fundamental and harmonics of the harmonic signal). In this way, the computationally complex and error-prone backward adaptation of the prediction coefficients becomes unnecessary.
[0022] In an embodiment, the encoder 100 can be configured to determine one interval value per frame.
[0023] In an embodiment, a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 can be separated by at least one spectral coefficient 106_t0_f3.
[0024] In an embodiment, the encoder 100 can be configured to apply predictive coding to a plurality of individual spectral coefficients separated by at least one spectral coefficient, such as, for example, to two individual spectral coefficients separated by at least one spectral coefficient. Further, the encoder 100 can be configured to apply predictive coding to a plurality of groups of spectral coefficients separated by at least one spectral coefficient, such as, for example, to two groups of spectral coefficients separated by at least one spectral coefficient (each group includes at least two spectral coefficients). Further, the encoder 100 can be configured to apply predictive coding to a plurality of, individual spectral coefficients and / or groups of spectral coefficients separated by at least one spectral coefficient, such as, for example, to at least one individual spectral coefficient and at least one group of spectral coefficients separated by at least one spectral coefficient.
[0025] In the example shown in FIG. 1, the encoder 100 is configured to determine six spectral coefficients 106_t0_f1 to 106_t0_f6 for the current frame 108_t0 and six spectral coefficients 106_t-1_f1 to 106_t-1_f6 for the previous frame 108_t-1. As a result, the encoder 100 is configured to selectively apply predictive coding to the individual second spectral coefficient 106_t0_f2 of the current frame and to a group of spectral coefficients consisting of the fourth and fifth spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0. As can be seen, the individual second spectral coefficient 106_t0_f2 and the group of spectral coefficients consisting of the fourth and fifth spectral coefficients 106_t0_f4 and 106_t0_f5 are separated from each other by the third spectral coefficient 106_t0_f3.
[0026] Note that, in this specification, the term "selectively" means applying predictive coding only to the selected spectral coefficients. That is, predictive coding is not necessarily applied to all spectral coefficients, but rather is applied only to the selected, individual spectral coefficients or groups of spectral coefficients, i.e., the selected, individual spectral coefficients and / or groups of spectral coefficients that can be separated from each other by at least one spectral coefficient. That is, predictive coding can disable a plurality of selected, individual spectral coefficients or groups of spectral coefficients with respect to at least one separating spectral coefficient.
[0027] In an embodiment, the encoder 100 can be configured to selectively apply predictive coding to a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 based on at least the corresponding plurality of individual spectral coefficients 106_t-1_f2 or groups of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1.
[0028] For example, the encoder 100 can be configured to predictively code a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 by encoding the prediction error between a plurality of individual predicted spectral coefficients 110_t0_f2 or groups of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame and the plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 (or their quantized versions) of the current frame.
[0029] In FIG. 1, the encoder 100 encodes the individual predicted spectral coefficients 110_t0_f2 of the current frame 108_t0 and the individual spectral coefficients 106_t0_f2 of the current frame 108_t0, and between the predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 of the current frame and the spectral coefficient groups 106_t0_f4 and 106_t0_f5 of the current frame, by encoding the prediction error, the individual spectral coefficients 106_t0_f2, and the spectral coefficient groups consisting of the spectral coefficients 106_t0_f4 and 106_t0_f5 are encoded.
[0030] That is, the second spectral coefficient 106_t0_f2 is encoded by encoding the prediction error (or difference) between the predicted second spectral coefficient 110_t0_f2 and the (actual or determined) second spectral coefficient 106_t0_f2, and the fourth spectral coefficient 106_t0_f4 is encoded by encoding the prediction error (or difference) between the predicted fourth spectral coefficient 110_t0_f4 and the (actual or determined) fourth spectral coefficient 106_t0_f4, and the fifth spectral coefficient 106_t0_f5 is encoded by encoding the prediction error (or difference) between the predicted fifth spectral coefficient 110_t0_f5 and the (actual or determined) fifth spectral coefficient 106_t0_f5.
[0031] In one embodiment, the encoder 100 can be configured to determine a plurality of individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 by a plurality of corresponding actual versions of the previous frame 108_t-1, individual spectral coefficients 106_t-1_f2 or spectral coefficient groups 106_t-1_f4 and 106_t-1_f5.
[0032] That is, in the above determination process, the encoder 100 may directly use a plurality of individual actual spectrum coefficients 106_t-1_f2 or actual spectrum coefficient groups 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1, and 106_t-1_f2, 106_t-1_f4 and 106_t-1_f5 respectively represent the original, that is, still unquantized, spectrum coefficients or spectrum coefficient groups obtained by the encoder 100 such that the encoder can act in the conversion region or filter bank region 104.
[0033] For example, the encoder 100 can be configured to determine the predicted second spectrum coefficient 110_t0_f2 of the current frame 108_t0 based on the corresponding still unquantized version of the second spectrum coefficient 106_t-1_f2 of the previous frame 108_t-1, the predicted fourth spectrum coefficient 110_t0_f4 of the current frame 108_t0 based on the corresponding still unquantized version of the fourth spectrum coefficient 106_t-1_f4 of the previous frame 108_t-1, and also the predicted fifth spectrum coefficient 110_t0_f5 of the current frame 108_t0 based on the corresponding still unquantized version of the fifth spectrum coefficient 106_t-1_f5 of the previous frame.
[0034] The corresponding decoder, the embodiments of which will be described later in connection with FIG. 4, in the above determination step, since only a plurality of individual spectrum coefficients 106_t-1_f2 or a plurality of spectrum coefficient groups 106_t-1_f4 and 106_t-1_f5 of the transmitted quantized version of the previous frame 108_t-1 can be used for predictive decoding, by this approach, the predictive encoding and decoding scheme can exhibit a kind of harmonic shaping of quantization noise.
[0035] As it is, for example, such harmonic noise shaping that has been conventionally performed by long-term prediction (LTP) in the time domain can be subjectively advantageous for predictive coding, but in some cases, it may lead to undesirable excessive tonality being incorporated into the decoded audio signal, so it may not be preferable. For this reason, a substitute predictive coding scheme that is fully synchronized with the corresponding decoding and does not lead to quantization noise shaping while extracting any possible prediction gain itself will be described below. According to this alternative coding embodiment, the encoder 100 can be configured to determine a plurality of individual predicted spectral coefficients 110_t0_f2 or groups of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 using a plurality of individual spectral coefficients 106_t-1_f2 or groups of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the corresponding quantized version of the previous frame 108_t-1.
[0036] For example, the encoder 100 can be configured to determine the predicted second spectral coefficient 110_t0_f2 of the current frame 108_t0 based on the corresponding quantized version of the second spectral coefficient 106_t-1_f2 of the previous frame 108_t-1, the predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on the corresponding quantized version of the fourth spectral coefficient 106_t-1_f4 of the previous frame 108_t-1, and also the predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on the corresponding quantized version of the fifth spectral coefficient 106_t-1_f5 of the previous frame.
[0037] Furthermore, the encoder 100 is configured to derive prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5, and 114_f5 from the interval value, and to calculate a plurality of individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 using a plurality of corresponding quantized versions of individual spectral coefficients 106_t-1_f2 and 106_t-2_f2 or spectral coefficient groups 106_t-1_f4, 106_t-2_f4, 106_t-1_f5, and 106_t-2_f5 of at least two previous frames 108_t-1 and 108_t-2, as well as using the derived prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5, and 114_f5.
[0038] For example, the encoder 100 can be configured to derive prediction coefficients 112_f2 and 114_f2 for the second spectral coefficient 106_t0_f2 from the interval value, to derive prediction coefficients 112_f4 and 114_f4 for the fourth spectral coefficient 106_t0_f4 from the interval value, and to derive prediction coefficients 112_f5 and 114_f5 for the fifth spectral coefficient 106_t0_f5 from the interval value.
[0039] For example, the derivation of the prediction coefficients can be performed in the following manner. That is, when the interval value corresponds to the frequency f0 or its encoded version, the prediction is enabled, and the center frequency of the K-th group of spectral coefficients is fc = K*f0. When the sampling frequency is fs and the hop size of the conversion (shift between consecutive frames) is N, the ideal prediction coefficients in the K-th group assuming a sine wave signal of frequency fc are as follows.
[0040] p1 = 2*cos(N*2*pi*fc / fs) and p2 = -1.
[0041] For example, if any of the spectral coefficients 106_t0_f4 and 106_t0_f5 are within this group, the prediction coefficients are as follows.
[0042] 112_f4 = 112_f5 = 2 * cos(N * 2 * pi * fc / fs) and 114_f4 = 114_f5 = -1.
[0043] For reasons of stability, a damping factor d can be introduced, and as a result the following modified prediction coefficients are obtained.
[0044] 112_f4’ = 112_f5’ = d * 2 * cos(N * 2 * pi * fc / fs), 114_f4’ = 114_f5’ = d2.
[0045] Since the interval values are transmitted in the encoded audio signal 120, the decoder can accurately derive the same prediction coefficients 212_f4 = 212_f5 = 2 * cos(N * 2 * pi * fc / fs) and 114_f4 = 114_f5 = -1. When using the damping factor, the coefficients can be modified accordingly.
[0046] As shown in FIG. 1, the encoder 100 can be configured to provide an encoded audio signal 120. As a result, the encoder 100 can be configured to include in the encoded audio signal 120 a quantized version of the prediction error for a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 to which predictive coding is applied. Further, the encoder 100 can be configured to not include the prediction coefficients 112_f2 to 114_f5 in the encoded audio signal 120.
[0047] Thus, the encoder 100 can use only the prediction coefficients 112_f2 to 114_f5 to calculate the prediction errors between a plurality of individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5, and from them, the individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5, and the individual spectral coefficients 106_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 of the current frame. However, it will not provide the individual spectral coefficients 106_t0_f4 (or its quantized version) or spectral coefficient groups 106_t0_f4 and 106_t0_f5 (or their quantized versions) in the encoded audio signal 120, nor will it provide the prediction coefficients 112_f2 to 114_f5. Therefore, the decoder, as will be described later in connection with FIG. 4 in an embodiment, can derive the prediction coefficients 112_f2 to 114_f5 from the interval values to calculate a plurality of individual predicted spectral coefficients or predicted spectral coefficient groups for the current frame.
[0048] That is, the encoder 100 can be configured to provide the encoded audio signal 120 including the quantized version of the prediction error instead of the quantized version of the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 for the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 to which the predictive coding is applied.
[0049] Furthermore, the encoder 100 can be configured to provide an encoded audio signal 102 that includes the quantized version of the prediction error, and the quantized version of the spectral coefficients 106_t0_f3 that separates the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5, whose quantized versions are included in the encoded audio signal 120, and the spectral coefficients 106_t0_f3 or spectral coefficient groups whose quantized versions are provided without using predictive coding, in an alternating manner, with the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 being separated therefrom.
[0050] In an embodiment, the encoder 100 can be further configured to entropy encode the quantized version of the prediction error and the quantized version of the spectral coefficients 106_t0_f3 that separates the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5, and include the entropy encoded version in the encoded audio signal 120 (instead of its non-entropy encoded version).
[0051] FIG. 2 shows a plot of the amplitude of the audio signal 102 over frequency for the current frame 108_t0. Further, FIG. 2 shows the spectral coefficients in the transform domain or filter bank domain determined by the encoder 100 for the current frame 108_t0 of the audio signal 102.
[0052] As shown in FIG. 2, the encoder 100 can be configured to selectively apply predictive coding to a plurality of spectral coefficient groups 116_1 to 116_6 that are separated by at least one spectral coefficient. Specifically, in the embodiment shown in FIG. 2, the encoder 100 selectively applies predictive coding to six spectral coefficient groups 116_1 to 116_6, and each of the first five spectral coefficient groups 116_1 to 116_5 includes three spectral coefficients (for example, the second group 116_2 includes spectral coefficients 106_t0_f8, 106_t0_f9, and 106_t0_f10), and the sixth spectral coefficient group 116_6 includes two spectral coefficients. As a result, the six spectral coefficient groups 116_1 to 116_6 are separated by (five) spectral coefficient groups 118_1 to 118_5 to which predictive coding is not applied.
[0053] That is, as shown in FIG. 2, the encoder 100 can be configured to selectively apply predictive coding to the spectral coefficient groups 116_1 to 110_6 such that the spectral coefficient groups 116_1 to 116_6 to which predictive coding is applied and the spectral coefficient groups 118_1 to 118_5 to which predictive coding is not applied alternate.
[0054] In an embodiment, the encoder 100 can be configured to determine an interval value (indicated by arrows 122_1 and 122_2 in FIG. 2), and the encoder 100 can be configured to select a plurality of spectral coefficient groups 116_1 to 116_6 (or a plurality of individual spectral coefficients) to which predictive coding is applied based on the interval value.
[0055] The interval value can be, for example, the interval (or distance) between two characteristic frequencies of the audio signal 102, such as peaks 124_1 and 124_2 of the audio signal. Further, the interval value can be an integer spectral coefficient (or an index of the spectral coefficient) that approximates the interval between two characteristic frequencies of the audio signal. Of course, the interval value can also be a real number, fraction, or multiple of the integer spectral coefficient that represents the interval between two characteristic frequencies of the audio signal.
[0056] In an embodiment, the encoder 100 can be configured to determine the instantaneous fundamental frequency of the audio signal (102) and derive an interval value from the instantaneous fundamental frequency or a fraction or multiple thereof.
[0057] For example, the first peak 124_1 of the audio signal 102 can be the instantaneous fundamental frequency (or pitch, or first harmonic) of the audio signal 102. Therefore, the encoder 100 can be configured to determine the instantaneous fundamental frequency of the audio signal 102 and derive an interval value from the instantaneous fundamental frequency or a fraction or multiple thereof. In that case, the interval value can be an integer spectral coefficient (or a fraction or multiple thereof) that approximates the interval between the instantaneous fundamental frequency 124_1 and the second harmonic 124_2 of the audio signal 102.
[0058] Of course, the audio signal 102 can include more than two harmonics. For example, the audio signal 102 shown in FIG. 2 includes six spectrally distributed harmonics 124_1 to 124_6 such that the audio signal 102 includes harmonics at all integer multiples of the instantaneous fundamental frequency. Of course, it is also possible that the audio signal 102 includes only some, not all, of the harmonics, such as the first, third, and fifth harmonics.
[0059] In an embodiment, the encoder 100 can be configured to select spectral coefficient groups 116_1 to 116_6 (or individual spectral coefficients) spectrally arranged by a harmonic grid defined by an interval value for predictive coding. As a result, the harmonic grid defined by the interval value represents the periodic spectral distribution (equidistant intervals) of the harmonics in the audio signal 102. That is, the harmonic grid defined by the interval value can be a series of interval values representing the equidistant intervals of the harmonics of the audio signal.
[0060] Furthermore, the encoder 100 can be configured to select spectral coefficients (e.g., only such spectral coefficients) whose spectral index is equal to or within a peripheral range (e.g., predetermined or variable) of a plurality of spectral indexes derived based on an interval value for predictive coding.
[0061] An index (or number) of a spectral coefficient representing a harmonic of the audio signal 102 can be derived from the interval value. For example, assuming that the fourth spectral coefficient 106_t0_f4 represents the instantaneous fundamental frequency of the audio signal 102 and the interval value is 5, the spectral coefficient having index 9 can be derived based on the interval value. As can be seen in FIG. 2, the spectrally derived spectral coefficient having index 9, i.e., the ninth spectral coefficient 106_t0_f9, represents the second harmonic. Similarly, spectral coefficients having indexes 14, 19, 24, and 29 can be derived, representing the third to sixth harmonics 124_3 to 124_6. However, not only spectral coefficients having an index equal to a plurality of spectral indexes derived based on the interval value, but also spectral coefficients having an index within a predetermined range around the plurality of spectral indexes derived based on the interval value can be predictively coded. For example, as shown in FIG. 2, the range can be set to 3 so that a plurality of spectral coefficient groups rather than a plurality of individual spectral coefficients are selected for predictive coding.
[0062] Furthermore, the encoder 100 may be configured to select the spectral coefficient groups 116_1 to 116_6 (or a plurality of individual spectral coefficients) to which predictive coding is applied such that the spectral coefficient groups 116_1 to 116_6 (or a plurality of individual spectral coefficients) to which predictive coding is applied and the spectral coefficients separating the spectral coefficient groups (or a plurality of individual spectral coefficients) to which predictive coding is applied alternate periodically with a period with an allowable range of + / -1 spectral coefficients. The allowable range of + / -1 spectral coefficients may be required when the distance between two harmonics of the audio signal 102 is not equal to an integer interval value (an integer related to the index or number of spectral coefficients), but rather equal to a fraction or multiple thereof. This can also be seen in FIG. 2 in that the arrows 122_1 to 122_6 do not necessarily point exactly to the center or middle of the corresponding spectral coefficients.
[0063] That is, the audio signal 102 includes at least two harmonic signal elements 124_1 to 124_6, and the encoder 100 may be configured to selectively apply predictive coding to at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102 or a plurality of spectral coefficient groups 116_1 to 116_6 (or individual spectral coefficients) representing the spectral environment around at least two harmonic signal elements 124_1 to 124_6. The spectral environment around at least two harmonic signal elements 124_1 to 124_6 can be, for example, + / -1, 2, 3, 4, or 5 spectral elements.
[0064] As a result, the encoder 100 can be configured not to apply predictive coding to spectral coefficient groups 118_1 to 118_5 (or individual spectral coefficients) that do not represent the spectral environment of at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102, or at least two harmonic signal elements 124_1 to 124_6. That is, the encoder 100 can be configured not to apply predictive coding to a plurality of spectral coefficient groups 118_1 to 118_5 (or individual spectral coefficients) belonging to the non-tonal background noise between the signal harmonics 124_1 to 124_6.
[0065] Furthermore, the encoder 100 can be configured to determine a harmonic interval value indicating the spectral interval between at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102, the harmonic interval value indicating a plurality of individual spectral coefficients or spectral coefficient groups representing at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102.
[0066] Furthermore, the encoder 100 can be configured to provide an encoded audio signal 120 such that the encoded audio signal 120 includes an interval value (e.g., one interval value per frame), or alternatively, a parameter from which the interval value can be directly derived.
[0067] Embodiments of the present invention address the above two problems of the FDP method by introducing into the FDP process the harmonic interval values sent from the encoder (transmitter) 100 to each decoder (receiver) so that they can all operate in a fully synchronized manner. The harmonic interval value can serve as an indicator of the instantaneous fundamental frequency (or pitch) of one or more spectra associated with the frame to be encoded, specifying which spectral bins (spectral coefficients) are to be predicted. More specifically, only the spectral coefficients around the harmonic signal elements located at integer multiples of the reference pitch (as defined by the harmonic interval value) in terms of indexing are to be the subject of prediction. FIGS. 2 and 3 illustrate this pitch-adaptive prediction approach by simple examples. FIG. 3 shows the operation of a state-of-the-art predictor in MPEG-2 AAC, where not only the prediction is performed only around the harmonic grid, but all spectral bins below a certain end frequency are the subject of prediction. FIG. 2 represents a modified same predictor according to an embodiment integrated to perform prediction only on "tonal" bins close to the harmonic interval grid.
[0068] By comparing FIGS. 2 and 3, two advantages of the modification according to an embodiment become apparent, namely, (1) the number of spectral bins included in the prediction process is much smaller, reducing the complexity (in the given example, only three-fifths of the bins are predicted, i.e., 40%), and (2) the bins belonging to the non-tonal background noise between the signal harmonics are not affected by the prediction, thereby increasing the prediction efficiency.
[0069] Note that the harmonic interval value does not necessarily have to correspond to the actual instantaneous pitch of the input signal and can represent a fraction or multiple of the true pitch if it thereby results in an overall improvement in the efficiency of the prediction process. It should also be emphasized that the harmonic interval value does not have to reflect integer multiples of the bin indexing or bandwidth units and can include fractions of said units.
[0070] Next, a preferred embodiment of an MPEG-style audio coder will be described.
[0071] Preferably, pitch-adaptive prediction is incorporated into MPEG-2 AAC (ISO / IEC 13818-7 "Information technology - Part 7: Advanced Audio Coding (AAC)", 2006), or a predictor similar to that in AAC is utilized and incorporated into the MPEG-H 3D audio codec (ISO / IEC 23008-3 "Information technology - High efficiency coding, part 3: 3D audio", 2015). Specifically, a 1-bit flag can be written to and read from each bitstream for each frame and channel that is not encoded alone (for a single frame channel, the flag is not transmitted because the prediction can be disabled to ensure uniqueness). When the flag is set to 1, another 8 bits can be read and written. These 8 bits represent the quantized version of the harmonic frequency interval value (e.g., the index for the harmonic interval) for a given frame and channel. Using the interval value derived from the quantized version using either a linear or non-linear mapping function, the prediction process can be executed in the method according to an embodiment shown in FIG. 2. Preferably, only the bins located within a range of a maximum distance of 1.5 bins around the harmonic grid are subject to prediction. For example, if a harmonic line with a harmonic interval value at bin index 47.11 is shown, only the bins at indices 46, 47, and 48 are predicted. However, the maximum distance may be defined differently, either fixed a priori for all channels and frames based on the high-frequency interval value or fixed separately for each frame and channel.
[0072] FIG. 4 shows a schematic block diagram of a decoder 200 that multiplexes an encoded audio signal 120. The decoder 200 is configured to decode the encoded audio signal 120 in a transform domain or filter bank region 204, and the decoder 200 is configured to analyze the encoded audio signal 120 to obtain the encoded spectral coefficients 206_t0_f1 to 206_t0_f6 of the audio signal for the current frame 208_t0 and the encoded spectral coefficients 206_t-1_f0 to 206_t-1_f6 for at least one previous frame 208_t-1, and the decoder 200 is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients that are separated by at least one encoded spectral coefficient.
[0073] In an embodiment, the decoder 200 can be configured to apply predictive decoding to a plurality of individual encoded spectral coefficients that are separated by at least one encoded spectral coefficient, such as, for example, to two individual encoded spectral coefficients that are separated by at least one encoded spectral coefficient. Further, the decoder 200 can be configured to apply predictive decoding to a plurality of groups of encoded spectral coefficients that are separated by at least one encoded spectral coefficient, such as, for example, to two groups of encoded spectral coefficients that are separated by at least one encoded spectral coefficient (each of the groups includes at least two encoded spectral coefficients). Further, the decoder 200 can be configured to apply predictive decoding to a plurality of individual encoded spectral coefficients and / or groups of encoded spectral coefficients that are separated by at least one encoded spectral coefficient, such as, for example, to at least one individual encoded spectral coefficient and at least one group of encoded spectral coefficients that are separated by at least one encoded spectral coefficient.
[0074] In the example shown in FIG. 4, the decoder 200 is configured to determine six encoded spectral coefficients 206_t0_f1 to 206_t0_f6 for the current frame 208_t0 and six encoded spectral coefficients 206_t-1_f1 to 206_t-1_f6 for the previous frame 208_t-1. As a result, the decoder 200 is configured to selectively apply predictive decoding to the individual encoded second spectral coefficient 206_t0_f2 of the current frame and to an encoded spectral coefficient group consisting of the encoded fourth and fifth spectral coefficients 206_t0_f4 and 206_t0_f5 of the current frame 208_t0. As can be seen, the individual encoded second spectral coefficient 206_t0_f2 and the encoded spectral coefficient group consisting of the encoded fourth and fifth spectral coefficients 206_t0_f4 and 206_t0_f5 are separated from each other by the encoded third spectral coefficient 206_t0_f3.
[0075] Note that in this specification, the term "selectively" means applying predictive decoding to (only) the selected encoded spectral coefficients. That is, predictive decoding is not applied to all the encoded spectral coefficients, but rather only to the selected individual encoded spectral coefficients or encoded spectral coefficient groups, i.e., the selected individual encoded spectral coefficients and / or encoded spectral coefficient groups that are separated from each other by at least one encoded spectral coefficient. That is, predictive decoding is not applied to at least one encoded spectral coefficient that separates a plurality of selected individual encoded spectral coefficients or encoded spectral coefficient groups.
[0076] In an embodiment, the decoder 200 can be configured not to apply predictive decoding to at least one encoded spectral coefficient 206_t0_f3 that separates individual encoded spectral coefficients 206_t0_f2 or spectral coefficient groups 206_t0_f4 and 206_t0_f5.
[0077] The decoder 200 can be configured to entropy-decode the encoded spectral coefficients to obtain the quantized prediction error for the spectral coefficients 206_t0_f2, 2016_t0_f4, and 206_t0_f5 for which predictive decoding is to be applied, and the quantized spectral coefficient 206_t0_f3 for at least one spectral coefficient for which predictive encoding is not to be applied. As a result, the decoder 200 can be configured to apply the quantized prediction error to a plurality of individual predicted spectral coefficients 210_t0_f2 or predicted spectral coefficient groups 210_t0_f4 and 210_t0_f5 to obtain the decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f2, 206_t0_f4, and 206_t0_f5 for which predictive decoding is applied for the current frame 208_t0.
[0078] For example, the decoder 200 can be configured to obtain a quantized second prediction error for the quantized second spectral coefficient 206_t0_f2 in order to obtain a decoded second spectral coefficient associated with the encoded second spectral coefficient 206_t0_f2, and to apply the quantized second prediction error to the predicted second spectral coefficient 210_t0_f2. The decoder 200 can be configured to obtain a quantized fourth prediction error for the quantized fourth spectral coefficient 206_t0_f4 in order to obtain a decoded fourth spectral coefficient associated with the encoded fourth spectral coefficient 206_t0_f4, and to apply the quantized fourth prediction error to the predicted fourth spectral coefficient 210_t0_f4. The decoder 200 can be configured to obtain a quantized fifth prediction error for the quantized fifth spectral coefficient 206_t0_f5 in order to obtain a decoded fifth spectral coefficient associated with the encoded fifth spectral coefficient 206_t0_f5, and to apply the quantized fifth prediction error to the predicted fifth spectral coefficient 210_t0_f5.
[0079] Furthermore, the decoder 200 can be configured to determine a plurality of individual predicted spectral coefficients 210_t0_f2 or groups of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 for the current frame 208_t0 based on a corresponding plurality of individual encoded spectral coefficients 206_t-1_f2 (e.g., using a plurality of previously decoded spectral coefficients associated with the plurality of individual encoded spectral coefficients 206_t-1_f2) or groups of encoded spectral coefficients 206_t-1_f4 and 206_t-1_f5 (e.g., using a previously decoded group of spectral coefficients associated with the 206_t-1_f4 and 206_t-1_f5 of the encoded spectral coefficients) of the previous frame 208_t-1.
[0080] For example, the decoder 200 can be configured to determine the predicted second spectral coefficient 210_t0_f2 of the current frame 208_t0 using the previously decoded (quantized) second spectral coefficient associated with the encoded second spectral coefficient 206_t-1_f2 of the previous frame 208_t-1, the predicted fourth spectral coefficient 210_t0_f4 of the current frame 208_t0 using the previously decoded (quantized) fourth spectral coefficient associated with the encoded fourth spectral coefficient 206_t-1_f4 of the previous frame 208_t-1, and also the predicted fifth spectral coefficient 210_t0_f5 of the current frame 208_t0 using the previously decoded (quantized) fifth spectral coefficient associated with the encoded fifth spectral coefficient 206_t-1_f5 of the previous frame 208_t-1.
[0081] Furthermore, the decoder 200 can be configured to derive prediction coefficients from the interval values, and the decoder 200 can calculate a plurality of individual predicted spectral coefficients 210_t0_f2 or groups of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 for the current frame 208_t0 using corresponding pluralities of the previously multiplexed individual spectral coefficients or previously multiplexed groups of spectral coefficients of at least two previous frames 208_t-1 and 208_t-2, and using the derived prediction coefficients.
[0082] For example, the decoder 200 can be configured to derive prediction coefficients 212_f2 and 214_f2 for the encoded second spectral coefficient 206_t0_f2 from the interval values, to derive prediction coefficients 212_f4 and 214_f4 for the encoded fourth spectral coefficient 206_t0_f4 from the interval values, and to derive prediction coefficients 212_f5 and 214_f5 for the encoded fifth spectral coefficient 206_t0_f5 from the interval values.
[0083] Note that the decoder 200 can be configured to multiplex the encoded audio signal 120 in order to obtain a quantized prediction error, instead of a plurality of individual quantized spectral coefficients or groups of quantized spectral coefficients for a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients to which prediction multiplexing is applied.
[0084] Furthermore, the decoder 200 can be configured to decode the encoded audio signal 120 in order to obtain quantized spectral coefficients that separate a plurality of individual spectral coefficients or groups of spectral coefficients such that the encoded spectral coefficient 206_t0_f2 or groups of encoded spectral coefficients 206_t0_f4 and 206_t0_f5 for which a quantized prediction error is obtained and the encoded spectral coefficient 206_t0_f3 or group of encoded spectral coefficients for which quantized spectral coefficients are obtained alternate.
[0085] The decoder 200 can be configured to provide the decoded audio signal 220 using the decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f2, 206_t0_f4, and 206_t0_f5 to which prediction decoding is applied, and using the entropy-decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f1, 206_t0_f3, and 206_t0_f6 to which prediction decoding is not applied.
[0086] In an embodiment, the decoder 200 can be configured to obtain an interval value, and the decoder 200 can be configured to select a plurality of individual encoded spectral coefficients 206_t0_f2 or groups of encoded spectral coefficients 206_t0_f4 and 206_t0_f5 to which prediction decoding is applied based on the interval value.
[0087] As already described above in connection with the description of the corresponding encoder 100, the interval value can be, for example, the interval (or distance) between two characteristic frequencies of the audio signal. Further, the interval value can be an integer spectral coefficient (or an index of the spectral coefficient) that approximates the interval between two characteristic frequencies of the audio signal. Of course, the interval value can also be a fraction or multiple of an integer spectral coefficient that represents the interval between two characteristic frequencies of the audio signal.
[0088] The decoder 200 can be configured to select individual spectral coefficients or groups of spectral coefficients spectrally arranged by a harmonic grid defined by an interval value for predictive decoding. The harmonic grid defined by the interval value can represent a periodic spectral distribution (equidistant intervals) of the harmonics in the audio signal 102. That is, the harmonic grid defined by the interval value can be a series of interval values that represent the equidistant intervals of the harmonics of the audio signal 102.
[0089] Furthermore, the decoder 200 can be configured to select spectral coefficients (e.g., only such spectral coefficients) whose spectral index is equal to or within a surrounding range (e.g., a predetermined or variable range) of a plurality of spectral indexes derived based on an interval value for predictive encoding. As a result, the decoder 200 can be configured to set the width of the range according to the interval value.
[0090] In an embodiment, the encoded audio signal includes an interval value or an encoded version thereof (e.g., a parameter from which the interval value can be directly derived), and the decoder 200 can be configured to extract the interval value or an encoded version thereof from the encoded audio signal to obtain the interval value.
[0091] As an alternative, the decoder 200 can be configured to determine the spacing value itself, i.e., such that the encoded audio signal does not include the spacing value. In that case, the decoder 200 can be configured to determine the instantaneous fundamental frequency (of the encoded audio signal 120 representing the audio signal 102) and to derive the spacing value from the instantaneous fundamental frequency or a fraction or multiple thereof.
[0092] In an embodiment, the decoder 200 can be configured to select a plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied such that the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied and the spectral coefficients separating the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied alternate periodically with a period with an allowable range of + / -1 spectral coefficients.
[0093] In an embodiment, the audio signal 102 represented by the encoded audio signal 120 includes at least two harmonic signal components, and the decoder 200 is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients 206_t0_f2 or groups of encoded spectral coefficients 206_t0_f4 and 206_t0_f5 representing at least two harmonic signal components of the audio signal 102 or the spectral environment around the at least two harmonic signal components. The spectral environment around the at least two harmonic signal components can be, for example, + / -1, 2, 3, 4, or 5 spectral elements.
[0094] As a result, the decoder 200 can be configured to selectively apply predictive decoding to at least two harmonic signal elements and a plurality of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficient groups 206_t0_f4 and 206_t0_f5 that are associated with the identified harmonic signal elements (e.g., represent the identified harmonic signal elements or enclose the identified harmonic signal elements).
[0095] Alternatively, the encoded audio signal 120 may include information (e.g., a spacing value) that identifies at least two harmonic signal elements. In that case, the decoder 200 can be configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficient groups 206_t0_f4 and 206_t0_f5 that are associated with the identified harmonic signal elements (e.g., represent the identified harmonic signal elements or enclose the identified harmonic signal elements).
[0096] In both of the above alternative methods, the decoder 200 can be configured not to apply predictive decoding to a plurality of individual encoded spectral coefficients 206_t0_f3, 206_t0_f1, and 206_t0_f6, or encoded spectral coefficient groups that do not represent at least two harmonic signal elements or the spectral environment of at least two harmonic signal elements of the audio signal 102.
[0097] That is, the decoder 200 can be configured not to apply predictive decoding to a plurality of individual encoded spectral coefficients 206_t0_f3, 206_t0_f1, 206_t0_f6, or encoded spectral coefficient groups that belong to the non-tonal background noise between the signal harmonics of the audio signal 102.
[0098] FIG. 5 shows a flowchart of a method 300 for encoding an audio signal according to an embodiment. The method 300 includes a step 302 of determining spectral coefficients of the audio signal for the current frame and at least one previous frame, and a step 304 of selectively applying predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients separated by at least one spectral coefficient.
[0099] FIG. 6 shows a flowchart of a method 400 for decoding an encoded audio signal according to an embodiment. The method 400 includes a step 402 of analyzing the encoded audio signal to obtain the encoded spectral coefficients of the audio signal for the current frame and at least one previous frame, and a step 404 of selectively applying predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients separated by at least one encoded spectral coefficient.
[0100] Although some aspects have been described in connection with an apparatus, it is clear that these aspects also represent a description of corresponding methods, where a block or device corresponds to a method step or a feature of a method step. Similarly, aspects described in connection with method steps also represent a description of corresponding blocks or items or features of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware device such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such a device.
[0101] The encoded audio signal according to the present invention can be stored in a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0102] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. For example, it can be implemented using a digital storage medium having electronically readable control signals stored thereon, such as a floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory, which cooperates with (or is capable of cooperating with) a programmable computer system so that each method is executed. Therefore, the digital storage medium can be computer-readable.
[0103] Some embodiments according to the present invention include a data carrier having electronically readable control signals, and the data carrier can cooperate with a programmable computer system so that one of the methods described herein is executed.
[0104] Generally, embodiments of the present invention can be implemented as a computer program product with program code, and the program code functions to execute one of the methods when the computer program product is executed on a computer. The program code can be stored, for example, on a machine-readable carrier.
[0105] Another embodiment includes a computer program stored on a machine-readable carrier for executing one of the methods described herein.
[0106] That is, an embodiment of the method according to the present invention is, as a result, a computer program having program code for executing one of the methods described herein when the computer program is executed on a computer.
[0107] A further embodiment of the method according to the invention is, as a result, a data carrier (or digital storage medium, or computer-readable medium) containing thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is usually tangible and / or non-transitory.
[0108] A further embodiment of the method according to the invention is, as a result, a data stream or a series of signals representing a computer program for performing one of the methods described herein. The data stream or series of signals can be configured, for example, to be transmitted via a data communication connection, such as via the Internet.
[0109] A further embodiment includes processing means, such as a computer, or a programmable logic device, configured or adapted to perform one of the methods described herein.
[0110] A further embodiment includes a computer having installed thereon a computer program for performing one of the methods described herein.
[0111] A further embodiment according to the invention includes an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver can be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system can include, for example, a file server for transmitting the computer program to the receiver.
[0112] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware device.
[0113] The apparatuses described herein may be implemented using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0114] The methods described herein may be performed using a hardware device, or using a computer, or using a combination of a hardware device and a computer.
[0115] The above embodiments are merely illustrative of the principles of the present invention. It will be appreciated that modifications and variations of the configurations and details described herein will be apparent to other persons skilled in the art. As a result, it is intended that the present invention be limited only by the appended claims, rather than by the specific details represented by the descriptions and explanations of the embodiments herein.
Claims
Decoder (200) for decoding an encoded audio signal (120), wherein the decoder (200) is configured to decode the encoded audio signal (120) in a transform domain or filter bank domain (204), the decoder (200) is configured to obtain the encoded spectral coefficients (206_t0_f1:206_t0_f6; 206_t-1_f1:206_t-1_f6) of the audio signal (120) for the current frame (208_t0) and at least one previous frame (208_t-1) by analyzing the encoded audio signal (120), the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients (206_t0_f2) or encoded spectral coefficient groups (206_t0_f4, 206_t0_f5), the decoder (200) is configured to obtain an interval value, and the decoder (200) is configured to select the plurality of individual encoded spectral coefficients (206_t0_f2) or encoded spectral coefficient groups (206_t0_f4, 206_t0_f5) to which predictive decoding is applied based on the interval value. The decoder (200) is configured to select, for predictive decoding, individual spectral coefficients (206_t0_f2) or spectral coefficient groups (206_t0_f4, 206_t0_f5) spectrally arranged by a harmonic grid defined by the interval value. The interval value is a harmonic interval value representing the interval between harmonics, decoder (200). **Claim 2** The decoder (200) according to claim 1, wherein the plurality of individual encoded spectral coefficients (206_t0_f2) or encoded spectral coefficient groups (206_t0_f4, 206_t0_f5) are separated by at least one encoded spectral coefficient (206_t0_f3). **Claim 3** The decoder (200) according to claim 2, wherein the predictive decoding is not applied to at least one spectral coefficient (206_t0_f3) separating the individual spectral coefficients (206_t0_f2) or the spectral coefficient groups (206_t0_f4, 206_t0_f5). **Claim 4** The decoder (200) is configured to entropy-decode the encoded spectral coefficients to obtain a quantized prediction error for the spectral coefficients (206_t0_f2, 206_t0_f4, 206_t0_f5) for which predictive decoding is to be applied, and a quantized spectral coefficient for the spectral coefficient (206_t0_f3) for which predictive decoding is not to be applied, The decoder (200) according to any one of claims 1 to 3, wherein, for the current frame (208_t0), the decoder (200) is configured to apply the quantized prediction error to a plurality of individual predicted spectral coefficients (210_t0_f2) or predicted spectral coefficient groups (210_t0_f4, 210_t0_f5) to obtain a decoded spectral coefficient associated with the encoded spectral coefficients (206_t0_f2, 206_t0_f4, 206_t0_f5) for which predictive decoding is to be applied.
5. The decoder (200) according to claim 4, wherein the decoder (200) is configured to determine the plurality of individual predicted spectral coefficients (210_t0_f2) or predicted spectral coefficient groups (210_t0_f4, 210_t0_f5) for the current frame (208_t0) based on a corresponding plurality of the individual encoded spectral coefficients (206_t-1_f2) or encoded spectral coefficient groups (206_t-1_f4, 206_t-1_f5) of the previous frame (208_t-1).
6. The decoder (200) according to claim 5, wherein the decoder (200) is configured to derive a prediction coefficient from the interval value, and the decoder (200) is configured to calculate the plurality of individual predicted spectral coefficients (210_t0_f2) or predicted spectral coefficient groups (210_t0_f4, 210_t0_f5) for the current frame (208_t0) using a corresponding plurality of previously decoded individual spectral coefficients or previously decoded spectral coefficient groups of at least two previous frames, and using the derived prediction coefficient.
7. The decoder (200) according to any one of claims 1 to 6, configured to decode the encoded audio signal (120) to obtain a quantized prediction error instead of a plurality of individual quantized spectral coefficients or groups of quantized spectral coefficients for the plurality of individual encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5) to which predictive decoding is applied.
8. The decoder according to claim 7, configured to decode the encoded audio signal (120) to obtain quantized spectral coefficients for encoded spectral coefficients (206_t0_f3) to which predictive coding is not applied, such that the encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5) for which a quantized prediction error is obtained and the encoded spectral coefficients (206_t0_f3) for which quantized spectral coefficients are obtained alternate.
9. The decoder (200) according to any one of claims 1 to 8, configured to select spectral coefficients whose spectral index is equal to or within a peripheral range of a plurality of spectral indices derived based on the interval value for predictive decoding.
10. The decoder (200) according to claim 9, configured to set the width of the range according to the interval value.
11. The encoded audio signal (120) includes the interval value or an encoded version thereof, and the decoder (200) is configured to extract the interval value or the encoded version thereof from the encoded audio signal (120) to obtain the interval value. The decoder (200) according to any one of claims 1 to 10.
12. The decoder (200) according to any one of claims 1 to 10, configured to determine the interval value.
13. The decoder (200) according to claim 12, configured to determine an instantaneous fundamental frequency and to derive the interval value from the instantaneous fundamental frequency or a fraction or multiple thereof.
14. The decoder (200) according to any one of claims 1 to 13, configured to select the plurality of individual spectral coefficients (206_t0_f2) or spectral coefficient groups (206_t0_f4, 206_t0_f5) to which predictive decoding is applied and the spectral coefficients (206_t0_f3) to which predictive decoding is not applied such that they periodically alternate in a period with an allowable range of + / −1 spectral coefficient.
15. The audio signal (102) represented by the encoded audio signal (120) includes at least two harmonic signal components (124_1:124_6), and the decoder (200) is configured to selectively apply predictive decoding to the at least two harmonic signal components (124_1:124_6) of the audio signal (102) or a plurality of individual encoded spectral coefficients or encoded spectral coefficient groups representing the spectral environment around the at least two harmonic signal components (124_1:124_6). The decoder (200) according to any one of claims 1 to 14.
16. The decoder (200) according to claim 15, configured to identify the at least two harmonic signal components (124_1:124_6) and to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or encoded spectral coefficient groups associated with the identified harmonic signal components (124_1:124_6).
17. The encoded audio signal (120) includes the interval value or its encoded version, the interval value identifies the at least two harmonic signal elements (124_1:124_6), and the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or encoded spectral coefficient groups associated with the identified harmonic signal elements (124_1:124_6). The decoder (200) according to claim 15.
18. The decoder (200) is configured not to apply predictive decoding to a plurality of individual encoded spectral coefficients or encoded spectral coefficient groups that do not represent the at least two harmonic signal elements (124_1:124_6) of the audio signal or the spectral environment of the at least two harmonic signal elements (124_1:124_6). The decoder (200) according to any one of claims 15 to 17.
19. The decoder (200) is configured not to apply predictive decoding to a plurality of individual encoded spectral coefficients or encoded spectral coefficient groups belonging to non-tonal background noise between the signal harmonics (124_1:124_6) of the audio signal. The decoder (200) according to any one of claims 15 to 18.
20. The encoded audio signal (120) includes the interval value or its encoded version, the interval value is a harmonic interval value, and the harmonic interval value indicates a plurality of individual encoded spectral coefficients or encoded spectral coefficient groups representing at least two harmonic signal elements (124_1:124_6) of the audio signal (102). The decoder (200) according to any one of claims 1 to 19.
21. The spectral coefficients are spectral bins. The decoder (200) according to any one of claims 1 to 20.
22. A method (400) for decoding an encoded audio signal in a transform domain or a filter bank domain, the method comprising: Analyzing the encoded audio signal (402) to obtain the encoded spectral coefficients of the audio signal for the current frame and at least one previous frame; Obtaining an interval value; Selectively applying predictive decoding (404) to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients, wherein the plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients to which predictive decoding is applied are selected based on the interval value; Selecting individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) spectrally arranged by a harmonic grid defined by the interval value for predictive decoding; comprising; The method, wherein the interval value is a harmonic interval value representing an interval between harmonics. **Claim 23** A computer program for performing the method according to claim 22.
Citation Information
Patent Citations
Encoding of forecast residual signal
JP1985031198A
JPP6666356B
JPP7078592B
Prediction of spectral coefficients in waveform coding and decoding
US20070016415A1