Audio encoder, audio decoder, method for encoding audio signals, and method for decoding encoded audio signal
Selective predictive coding of spectral coefficients based on harmonic signal elements addresses the complexity and efficiency issues in FDP, enhancing audio coding performance.
Patent Information
- Application Number
- JP2025114947
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-06-17
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-25
AI Technical Summary
Frequency-domain prediction (FDP) methods in audio coding require high computational complexity and have limited prediction efficiency due to noisy components being subjected to prediction, leading to reconstruction errors.
Apply predictive coding selectively to spectral coefficients around harmonic signal elements, using a spacing value to determine which coefficients are predicted, reducing complexity and improving efficiency by avoiding noisy regions.
Reduces computational complexity and enhances prediction efficiency by applying predictive coding only to harmonic signal elements, minimizing errors and improving audio quality.
Smart Images

Figure 2025138865000001_ABST
Abstract
Description
[Technical Field]
[0001] Embodiments relate to audio coding, in particular to methods and apparatus for encoding audio signals using predictive coding and for decoding encoded audio signals using predictive decoding. Preferred embodiments relate to methods and apparatus for pitch-adaptive spectral prediction. Further preferred embodiments relate to perceptual coding of tonal audio signals by transform coding using inter-frame prediction tools in the spectral domain. [Background technology]
[0002] To improve the quality of coded tonal signals, especially at low bit rates, modern audio transform coders use very long transforms and / or long-term prediction or pre- / post-filtering. However, long transforms imply long algorithmic delays, which are undesirable for low-delay communication scenarios. Therefore, very low-delay predictors based on instantaneous reference pitch have recently gained popularity. The Internet Engineering Task Force (IETF) Opus codec utilizes pitch-adaptive pre- and post-filtering in its frequency-domain Constrained-Energy Lapped Transform (CELT) coding path (J.M. Balin, K. Vos, and T. Terriberry, "Definition of the Opus audio codec," Internet Engineering Task Force, Technical Report RFC6716, 2012, http: / / tools.ietf.org / html / rfc67161), and the 3rd Generation Partnership Project (3GPP) Enhanced Voice Services (EVS) codec provides a long-term harmonic post-filter for perceptual improvement of transform-coded signals (3GPP TS 26.443 "Codec for Enhanced Voice Services (EVS)," Release 12, December 2014). Both of these approaches operate in the time domain on the fully decoded signal waveform and are difficult and / or computationally expensive to apply frequency-selectively (each scheme only provides a simple low-pass filter selectively for some frequencies). A welcome alternative to time-domain long-term prediction (LTP) or pre / post filtering (PPF) is consequently provided by frequency-domain prediction (FDP), as supported in MPEG-2 AAC (ISO / IEC 13818-7 "Information technology - Part 7: Advanced Audio Coding (AAC)", 2006). Although this method facilitates frequency selectivity, it has its own disadvantages, as described below. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] J.M. Balin, K. Vos, and T. Terriberry, "Definition of the Opus audio codec," Internet Engineering Task Force, Technical Report RFC6716, 2012, http: / / tools.ietf.org / html / rfc67161 [Non-patent document 2] 3GPP TS 26.443 "Codec for Enhanced Voice Services (EVS)", Release 12, December 2014 Summary of the Invention [Problem to be solved by the invention]
[0004] The FDP method introduced above has two drawbacks compared to other tools. First, it requires high computational complexity. Specifically, at least two linear predictive coding passes (i.e., from the channel transform bins of the last two frames) are applied to hundreds of spectral bins for each frame and channel, with worst-case prediction across all scale factor bands (ISO / IEC 13818-7 "Information technology - Part 7: Advanced Audio Coding (AAC)", 2006). Second, the FDP method has limited overall prediction gain. More specifically, the prediction efficiency is limited because noisy components between predictable harmonic tonal spectral regions are also subject to prediction, and these noisy regions are usually not predictable and therefore introduce errors.
[0005] The high complexity is due to the backward adaptivity of the predictors, which means that the prediction coefficients for each bin must be calculated based on previously transmitted bins. Therefore, numerical inaccuracies between the encoder and decoder can lead to reconstruction errors due to discrepant prediction coefficients. To overcome this problem, bit-exact identical adaptation must be guaranteed. Furthermore, even if a group of predictors is disabled in a frame, adaptation must always be performed to keep the prediction coefficients up to date. [Means for solving the problem]
[0006] It is therefore an object of the present invention to provide a concept for encoding an audio signal and / or decoding an encoded audio signal that avoids at least one (e.g. both) of the aforementioned problems and leads to more efficient and less computationally expensive implementations.
[0007] The independent claims solve this problem.
[0008] The dependent claims deal with advantageous embodiments.
[0009] An embodiment provides an encoder for encoding an audio signal, the encoder being configured to encode the audio signal in a transform domain or a filter bank domain, the encoder being configured to determine spectral coefficients of the audio signal for a current frame and at least one previous frame, the encoder being configured to selectively apply predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients, the encoder being configured to determine a spacing value, and the encoder being configured to select a plurality of individual spectral coefficients or groups of spectral coefficients to which the predictive coding is applied based on the spacing value, which may be transmitted as side information together with the encoded audio signal.
[0010] A further embodiment provides a decoder for decoding an encoded audio signal (e.g., encoded by the above-mentioned encoder), wherein the decoder is configured to decode the encoded audio signal in a transform domain or a filter bank domain, the decoder is configured to analyze the encoded audio signal to obtain coded spectral coefficients of the audio signal for a current frame and at least one previous frame, and the decoder is configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients, and the decoder may be configured to select a plurality of individual coded spectral coefficients or groups of coded spectral coefficients to which predictive decoding is applied based on a transmitted spacing value.
[0011] According to the inventive concept, predictive coding is applied to (only) selected spectral coefficients. The spectral coefficients to which predictive coding is applied can be selected according to signal characteristics. For example, by not applying predictive coding to noisy signal elements, the aforementioned errors caused by predicting unpredictable noisy signal elements are avoided. At the same time, since predictive coding is applied only to selected spectral elements, computational complexity can be reduced.
[0012] For example, perceptual coding of tonal audio signals can be performed (e.g., by an encoder) by transform coding in conjunction with guided / adaptive spectral domain inter-frame prediction techniques. The efficiency of frequency domain prediction (FDP) can be increased and computational complexity reduced by applying the prediction only to spectral coefficients around harmonic signal elements located at integer multiples of a fundamental frequency or fundamental pitch, which can be sent, for example, as interval values, in an appropriate bitstream from the encoder to the decoder. Embodiments of the present invention can be preferably implemented or incorporated in an MPEG-H 3D audio codec, but are applicable to any audio transform coding system, such as, for example, MPEG-2 AAC.
[0013] A further embodiment provides a method for encoding an audio signal in a transform or filterbank domain, the method comprising: determining spectral coefficients of the audio signal for a current frame and at least one previous frame; determining an interval value; selectively applying predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients, the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive coding is applied being selected based on an interval value; Includes:
[0014] A further embodiment provides a method for decoding an encoded audio signal in a transform or filter bank domain, the method comprising: analyzing the encoded audio signal to obtain encoded spectral coefficients of the audio signal for a current frame and at least one previous frame; Obtaining an interval value; selectively applying predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients, wherein the plurality of individual coded spectral coefficients or groups of coded spectral coefficients to which predictive decoding is applied are selected based on an interval value; Includes:
[0015] Embodiments of the present invention are described herein below with reference to the accompanying drawings. [Brief explanation of the drawings]
[0016] [Figure 1] 1 shows a schematic block diagram of an encoder for encoding an audio signal according to one embodiment; [Figure 2]1 illustrates the amplitude of an audio signal plotted over frequency for a current frame and corresponding selected spectral coefficients to which predictive coding is applied, according to one embodiment. [Figure 3] The figure shows the amplitude of the audio signal plotted over frequency for the current frame and the corresponding spectral coefficients that are subject to prediction by MPEG-2 AAC. [Figure 4] 1 shows a schematic block diagram of a decoder for decoding an encoded audio signal according to one embodiment; [Figure 5] 2 shows a flowchart of a method for encoding an audio signal according to an embodiment; [Figure 6] 2 shows a flowchart of a method for decoding an encoded audio signal according to one embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0017] Equivalent or corresponding elements, or elements with equivalent or corresponding functionality, are indicated in the following description by equivalent or corresponding reference numerals.
[0018] In the following description, numerous details are set forth to more fully describe embodiments of the present invention. However, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without these specific details. In other instances, well-known structures and devices are shown in block diagram form, rather than in detail, in order to avoid obscuring the embodiments of the present invention. Furthermore, features of different embodiments described below may be combined with each other unless otherwise noted.
[0019] 1 shows a schematic block diagram of an encoder 100 for encoding an audio signal 102, according to one embodiment. The encoder 100 is configured to encode the audio signal 102 in a transform or filter bank domain 104 (e.g., the frequency domain or the spectral domain), and the encoder 100 is configured to determine spectral coefficients 106_t0_f1 to 106_t0_f6 of the audio signal 102 for a current frame 108_t0 and spectral coefficients 106_t-1_f1 to 106_t-1_f6 of the audio signal for at least one previous frame 108_t-1. Further, the encoder 100 is configured to selectively apply predictive coding to a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5, the encoder 100 is configured to determine an interval value, and the encoder 100 is configured to select a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 to which the predictive coding is applied based on the interval value.
[0020] That is, the encoder 100 is configured to selectively apply predictive coding to a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 selected based on a single spacing value transmitted as side information.
[0021] This spacing value, together with its integer multiples, may correspond to a frequency (e.g., the fundamental frequency of a harmonic tone (of the audio signal 102)) that defines the center for all groups of spectral coefficients to which prediction is applied; i.e., the first group may be around this frequency, the second group may be centered around this frequency times 2, the third group may be centered around this frequency times 3, etc. Knowledge of these center frequencies allows the calculation of prediction coefficients for predicting the corresponding sinusoidal signal components (e.g., the fundamental and overtones of a harmonic signal). In this way, complex and error-prone backward adaptation of the prediction coefficients is not necessary.
[0022] In an embodiment, the encoder 100 may be configured to determine one spacing value per frame.
[0023] In an embodiment, the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 may be separated by at least one spectral coefficient 106_t0_f3.
[0024] In embodiments, encoder 100 may be configured to apply predictive coding to a plurality of individual spectral coefficients that are separated by at least one spectral coefficient, such as, for example, to two individual spectral coefficients that are separated by at least one spectral coefficient. Furthermore, encoder 100 may be configured to apply predictive coding to a plurality of spectral coefficient groups (each group including at least two spectral coefficients) that are separated by at least one spectral coefficient, such as, for example, to two spectral coefficient groups that are separated by at least one spectral coefficient. Furthermore, encoder 100 may be configured to apply predictive coding to a plurality of individual spectral coefficients and / or spectral coefficient groups that are separated by at least one spectral coefficient, such as, for example, to at least one individual spectral coefficient and at least one spectral coefficient group that are separated by at least one spectral coefficient.
[0025] In the example shown in Figure 1, the encoder 100 is configured to determine six spectral coefficients 106_t0_f1 through 106_t0_f6 for the current frame 108_t0 and six spectral coefficients 106_t-1_f1 through 106_t-1_f6 for the previous frame 108_t-1. As a result, the encoder 100 is configured to selectively apply predictive coding to each second spectral coefficient 106_t0_f2 of the current frame and to a group of spectral coefficients consisting of the fourth and fifth spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0. As can be seen, each second spectral coefficient 106_t0_f2 and each group of spectral coefficients consisting of the fourth and fifth spectral coefficients 106_t0_f4 and 106_t0_f5 are separated from each other by the third spectral coefficient 106_t0_f3.
[0026] It should be noted that the term "selectively" in this specification refers to applying predictive coding to (only) selected spectral coefficients. That is, predictive coding is not necessarily applied to all spectral coefficients, but rather only to selected individual spectral coefficients or groups of spectral coefficients, i.e., selected individual spectral coefficients and / or groups of spectral coefficients that can be separated from each other by at least one spectral coefficient. That is, predictive coding can be disabled for at least one spectral coefficient that separates selected individual spectral coefficients or groups of spectral coefficients.
[0027] In an embodiment, the encoder 100 may be configured to selectively apply predictive coding to a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 of a current frame 108_t0 based on at least a corresponding plurality of individual spectral coefficients 106_t-1_f2 or spectral coefficient groups 106_t-1_f4 and 106_t-1_f5 of a previous frame 108_t-1.
[0028] For example, the encoder 100 may be configured to predictively encode a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 by encoding a prediction error between a plurality of individual predicted spectral coefficients 110_t0_f2 or groups of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame 108_t0 and a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame (or a quantized version thereof).
[0029] In FIG. 1, the encoder 100 encodes the individual spectral coefficient 106_t0_f2 and the group of spectral coefficients 106_t0_f4 and 106_t0_f5 by encoding the prediction errors between the individual predicted spectral coefficient 110_t0_f2 of the current frame 108_t0 and the individual spectral coefficient 106_t0_f2 of the current frame 108_t0, and between the predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 of the current frame and the spectral coefficient groups 106_t0_f4 and 106_t0_f5 of the current frame.
[0030] That is, the second spectral coefficient 106_t0_f2 is coded by coding the prediction error (or difference) between the predicted second spectral coefficient 110_t0_f2 and the (actual or determined) second spectral coefficient 106_t0_f2, the fourth spectral coefficient 106_t0_f4 is coded by coding the prediction error (or difference) between the predicted fourth spectral coefficient 110_t0_f4 and the (actual or determined) fourth spectral coefficient 106_t0_f4, and the fifth spectral coefficient 106_t0_f5 is coded by coding the prediction error (or difference) between the predicted fifth spectral coefficient 110_t0_f5 and the (actual or determined) fifth spectral coefficient 106_t0_f5.
[0031] In one embodiment, the encoder 100 may be configured to determine a plurality of individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 by a plurality of individual spectral coefficients 106_t-1_f2 or spectral coefficient groups 106_t-1_f4 and 106_t-1_f5 of the corresponding actual version of the previous frame 108_t-1.
[0032] That is, the encoder 100 may directly use multiple individual actual spectral coefficients 106_t-1_f2 or actual spectral coefficient groups 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1 in the above decision process, with 106_t-1_f2, 106_t-1_f4 and 106_t-1_f5 representing the original, i.e., not yet quantized, spectral coefficients or spectral coefficient groups, respectively, as obtained by the encoder 100 such that the encoder may work in the transform domain or filter bank domain 104.
[0033] For example, the encoder 100 may be configured to determine a predicted second spectral coefficient 110_t0_f2 of the current frame 108_t0 based on the second spectral coefficient 106_t-1_f2 of the corresponding yet-to-be-quantized version of the previous frame 108_t-1, a predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on the fourth spectral coefficient 106_t-1_f4 of the corresponding yet-to-be-quantized version of the previous frame 108_t-1, and a predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on the fifth spectral coefficient 106_t-1_f5 of the corresponding yet-to-be-quantized version of the previous frame.
[0034] A corresponding decoder, an embodiment of which will be described later in connection with Figure 4, can use only a number of individual spectral coefficients 106_t-1_f2 or a number of groups of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the transmitted quantized version of the previous frame 108_t-1 for predictive decoding in the above decision step, so that this approach allows the predictive encoding and decoding scheme to exhibit a kind of harmonic shaping of the quantization noise.
[0035] While such harmonic noise shaping, traditionally performed, for example, by long-term prediction (LTP) in the time domain, can be subjectively advantageous for predictive coding, it may be undesirable in some cases because it can lead to the introduction of undesirable excessive tonality into the decoded audio signal. For this reason, an alternative predictive coding scheme is described below that is fully synchronized with the corresponding decoding and, as such, extracts any possible prediction gain but does not lead to quantization noise shaping. According to this alternative coding embodiment, the encoder 100 may be configured to determine a plurality of individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 using a plurality of individual spectral coefficients 106_t-1_f2 or spectral coefficient groups 106_t-1_f4 and 106_t-1_f5 of a corresponding quantized version of the previous frame 108_t-1.
[0036] For example, the encoder 100 may be configured to determine a predicted second spectral coefficient 110_t0_f2 of the current frame 108_t0 based on the second spectral coefficient 106_t-1_f2 of the corresponding quantized version of the previous frame 108_t-1, a predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on the fourth spectral coefficient 106_t-1_f4 of the corresponding quantized version of the previous frame 108_t-1, and a predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on the fifth spectral coefficient 106_t-1_f5 of the corresponding quantized version of the previous frame.
[0037] Further, the encoder 100 is configured to derive prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5, and 114_f5 from the interval values, and to derive a plurality of individual predicted spectral coefficients 110_t0_f2 or groups of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 from at least two previous frames 108_t-1 and 108_t t-1_f2 and 106_t-2_f2 or spectral coefficient groups 106_t-1_f4, 106_t-2_f4, 106_t-1_f5 and 106_t-2_f5, and using the derived prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5 and 114_f5.
[0038] For example, the encoder 100 may be configured to derive prediction coefficients 112_f2 and 114_f2 for the second spectral coefficient 106_t0_f2 from the interval value, to derive prediction coefficients 112_f4 and 114_f4 for the fourth spectral coefficient 106_t0_f4 from the interval value, and to derive prediction coefficients 112_f5 and 114_f5 for the fifth spectral coefficient 106_t0_f5 from the interval value.
[0039] For example, the derivation of the prediction coefficients can be derived in the following way: if the spacing value corresponds to a frequency f0 or its coded version, then the center frequency of the Kth group of spectral coefficients, for which prediction is enabled, is fc = K * f0. If the sampling frequency is fs and the hop size (shift between successive frames) of the transform is N, then the ideal prediction coefficients in the Kth group, assuming a sinusoidal signal of frequency fc, are:
[0040] p1=2*cos(N*2*pi*fc / fs) and p2=-1.
[0041] For example, if both spectral coefficients 106_t0_f4 and 106_t0_f5 are in this group, the prediction coefficients are:
[0042] 112_f4=112_f5=2*cos(N*2*pi*fc / fs) and 114_f4=114_f5=-1.
[0043] For stability reasons, a damping factor d can be introduced, resulting in the following modified prediction coefficients:
[0044] 112_f4'=112_f5'=d*2*cos(N*2*pi*fc / fs), 114_f4'=114_f5'=d2.
[0045] Since the spacing values are transmitted in the coded audio signal 120, the decoder can derive exactly the same prediction coefficients 212_f4=212_f5=2*cos(N*2*pi*fc / fs) and 114_f4=114_f5=-1. If damping factors are used, the coefficients can be modified accordingly.
[0046] 1, the encoder 100 may be configured to provide an encoded audio signal 120. As a result, the encoder 100 may be configured to include, in the encoded audio signal 120, quantized versions of prediction errors for a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 to which predictive coding is applied. Furthermore, the encoder 100 may be configured to exclude, in the encoded audio signal 120, the prediction coefficients 112_f2 to 114_f5.
[0047] In this way, the encoder 100 may use only the prediction coefficients 112_f2 to 114_f5 to calculate a plurality of individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5, and therefrom a prediction error between the individual predicted spectral coefficients 110_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 and the individual spectral coefficients 106_t0_f2 or predicted spectral coefficient groups 110_t0_f4 and 110_t0_f5 of the current frame, but will not provide in the encoded audio signal 120 either the individual spectral coefficient 106_t0_f4 (or a quantized version thereof) or the spectral coefficient groups 106_t0_f4 and 106_t0_f5 (or a quantized version thereof), nor the prediction coefficients 112_f2 to 114_f5. Thus, the decoder, an embodiment of which is described below in connection with FIG. 4, may derive prediction coefficients 112_f2 to 114_f5 from the interval value to calculate multiple, individual predicted spectral coefficients or groups of predicted spectral coefficients for the current frame.
[0048] That is, the encoder 100 may be configured to provide an encoded audio signal 120 that includes a quantized version of the prediction error for a plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5 to which predictive coding is applied, instead of a quantized version of the plurality of individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5.
[0049] Furthermore, the encoder 100 may be configured to provide an encoded audio signal 102 including quantized versions of spectral coefficients 106_t0_f3 separating a plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 such that the spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5, whose quantized versions of the prediction errors are included in the encoded audio signal 120, alternate with the spectral coefficients 106_t0_f3 or groups of spectral coefficients, whose quantized versions are provided without predictive coding.
[0050] In an embodiment, the encoder 100 may be further configured to entropy encode the quantized version of the prediction error and the quantized version of the spectral coefficient 106_t0_f3 separating the multiple individual spectral coefficients 106_t0_f2 or spectral coefficient groups 106_t0_f4 and 106_t0_f5, and to include the entropy encoded version in the encoded audio signal 120 (instead of its non-entropy encoded version).
[0051] 2 illustrates the amplitude of the audio signal 102 plotted over frequency for a current frame 108_t0. Additionally, FIG. 2 illustrates the spectral coefficients in the transform or filter bank domain determined by the encoder 100 for the current frame 108_t0 of the audio signal 102.
[0052] As shown in Figure 2, the encoder 100 can be configured to selectively apply predictive coding to multiple spectral coefficient groups 116_1 to 116_6, which are separated by at least one spectral coefficient. Specifically, in the embodiment shown in Figure 2, the encoder 100 selectively applies predictive coding to six spectral coefficient groups 116_1 to 116_6, with the first five spectral coefficient groups 116_1 to 116_5 each including three spectral coefficients (e.g., the second group 116_2 includes spectral coefficients 106_t0_f8, 106_t0_f9, and 106_t0_f10), and the sixth spectral coefficient group 116_6 including two spectral coefficients. As a result, the six spectral coefficient groups 116_1 to 116_6 are separated by (five) spectral coefficient groups 118_1 to 118_5, to which predictive coding is not applied.
[0053] That is, as shown in FIG. 2, the encoder 100 can be configured to selectively apply predictive coding to the spectral coefficient groups 116_1 to 116_6 such that the spectral coefficient groups 116_1 to 116_6 to which predictive coding is applied alternate with the spectral coefficient groups 118_1 to 118_5 to which predictive coding is not applied.
[0054] In an embodiment, the encoder 100 may be configured to determine an interval value (indicated by arrows 122_1 and 122_2 in FIG. 2), and the encoder 100 may be configured to select, based on the interval value, a plurality of spectral coefficient groups 116_1 to 116_6 (or a plurality of individual spectral coefficients) to which predictive coding is applied.
[0055] The spacing value may be, for example, the spacing (or distance) between two characteristic frequencies of the audio signal 102, such as peaks 124_1 and 124_2 of the audio signal. Furthermore, the spacing value may be an integer spectral coefficient (or an index of spectral coefficients) that approximates the spacing between the two characteristic frequencies of the audio signal. Of course, the spacing value may also be a real number or a fraction or multiple of the integer spectral coefficient that represents the spacing between the two characteristic frequencies of the audio signal.
[0056] In an embodiment, the encoder 100 may be configured to determine an instantaneous fundamental frequency of the audio signal (102) and to derive the interval value from the instantaneous fundamental frequency or a fraction or multiple thereof.
[0057] For example, the first peak 124_1 of the audio signal 102 may be the instantaneous fundamental frequency (or pitch, or first harmonic) of the audio signal 102. As such, the encoder 100 may be configured to determine the instantaneous fundamental frequency of the audio signal 102 and to derive the spacing value from the instantaneous fundamental frequency or a fraction or multiple thereof. The spacing value may then be an integer spectral coefficient (or a fraction or multiple thereof) that approximates the spacing between the instantaneous fundamental frequency 124_1 and the second harmonic 124_2 of the audio signal 102.
[0058] Of course, audio signal 102 may contain more than two harmonics. For example, audio signal 102 shown in Figure 2 contains six harmonics 124_1 through 124_6 spectrally distributed such that audio signal 102 contains harmonics at all integer multiples of the instantaneous fundamental frequency. Of course, it is also possible for audio signal 102 to contain only some but not all of the harmonics, such as the first, third, and fifth harmonics.
[0059] In an embodiment, the encoder 100 may be configured to select, for predictive coding, spectral coefficient groups 116_1 to 116_6 (or individual spectral coefficients) that are spectrally arranged according to a harmonic grid defined by spacing values. As a result, the harmonic grid defined by the spacing values represents a periodic spectral distribution (equidistant spacing) of the harmonics in the audio signal 102. That is, the harmonic grid defined by the spacing values may be a series of spacing values that represent equidistant spacing of the harmonics of the audio signal.
[0060] Furthermore, the encoder 100 can be configured to select, for predictive coding, spectral coefficients (e.g., only those spectral coefficients) whose spectral indices are equal to or fall within a range (e.g., a predetermined and variable) around a plurality of spectral indices derived based on the spacing value.
[0061] From the interval value, an index (or number) of a spectral coefficient representing a harmonic of the audio signal 102 can be derived. For example, assuming that the fourth spectral coefficient 106_t0_f4 represents the instantaneous fundamental frequency of the audio signal 102 and the interval value is 5, a spectral coefficient having index 9 can be derived based on the interval value. As can be seen in FIG. 2, the spectral coefficient thus derived with index 9, i.e., the ninth spectral coefficient 106_t0_f9, represents the second harmonic. Similarly, spectral coefficients having indexes 14, 19, 24, and 29 can be derived to represent the third through sixth harmonics 124_3 through 124_6. However, not only spectral coefficients having indices equal to the spectral indices derived based on the interval value, but also spectral coefficients having indices within a predetermined range around the spectral indices derived based on the interval value can be predictively coded. For example, as shown in FIG. 2, the range can be 3 so that groups of spectral coefficients, rather than individual spectral coefficients, are selected for predictive coding.
[0062] Furthermore, the encoder 100 can be configured to select the spectral coefficient groups 116_1 to 116_6 (or a plurality of individual spectral coefficients) to which predictive coding is applied such that the spectral coefficient groups 116_1 to 116_6 (or a plurality of individual spectral coefficients) to which predictive coding is applied alternate periodically with a periodicity of + / -1 spectral coefficient between the spectral coefficients separating the spectral coefficient groups (or a plurality of individual spectral coefficients) to which predictive coding is applied. A tolerance of + / -1 spectral coefficient may be required when the distance between two harmonics of the audio signal 102 is not equal to an integer spacing value (an integer related to the index or number of the spectral coefficients) but rather a fraction or multiple thereof. This can also be seen in FIG. 2, in that the arrows 122_1 to 122_6 do not necessarily point exactly to the center or central portion of the corresponding spectral coefficients.
[0063] That is, the audio signal 102 includes at least two harmonic signal elements 124_1 to 124_6, and the encoder 100 can be configured to selectively apply predictive coding to a plurality of spectral coefficient groups 116_1 to 116_6 (or individual spectral coefficients) that represent the at least two harmonic signal elements 124_1 to 124_6, or the spectral environment surrounding the at least two harmonic signal elements 124_1 to 124_6, of the audio signal 102. The spectral environment surrounding the at least two harmonic signal elements 124_1 to 124_6 can be, for example, + / - 1, 2, 3, 4, or 5 spectral elements.
[0064] As a result, the encoder 100 can be configured to not apply predictive coding to spectral coefficient groups 118_1 to 118_5 (or multiple individual spectral coefficients) that do not represent the spectral environment of at least two harmonic signal elements 124_1 to 124_6 or at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102. That is, the encoder 100 can be configured to not apply predictive coding to multiple spectral coefficient groups 118_1 to 118_5 (or individual spectral coefficients) that belong to non-tonal background noise between the signal harmonics 124_1 to 124_6.
[0065] Further, the encoder 100 may be configured to determine harmonic spacing values indicative of a spectral spacing between at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102, the harmonic spacing values indicative of a plurality of individual spectral coefficients or groups of spectral coefficients representing the at least two harmonic signal elements 124_1 to 124_6 of the audio signal 102.
[0066] Furthermore, the encoder 100 may be configured to provide an encoded audio signal 120 such that the encoded audio signal 120 includes interval values (e.g., one interval value per frame) or (alternatively) parameters from which the interval values can be directly derived.
[0067] An embodiment of the present invention addresses these two challenges of FDP techniques by incorporating a harmonic spacing value into the FDP process, sent from the encoder (transmitter) 100 to each decoder (receiver) so that both can work in perfect synchronization. The harmonic spacing value can serve as an indicator of one or more spectral instantaneous fundamental frequencies (or pitches) associated with the frame to be encoded, specifying which spectral bins (spectral coefficients) are to be predicted. More specifically, only spectral coefficients around harmonic signal elements located (in terms of indexing) at integer multiples of the reference pitch (as defined by the harmonic spacing value) are to be considered for prediction. Figures 2 and 3 illustrate this pitch-adaptive prediction approach with a simple example. Figure 3 shows the operation of a state-of-the-art predictor in MPEG-2 AAC, which considers all spectral bins below a certain end frequency rather than predicting only around the harmonic grid, and Figure 2 shows the same predictor modified in one embodiment, integrated to predict only "tonal" bins close to the harmonic spacing grid.
[0068] Comparing Figures 2 and 3 reveals two advantages of the modification according to one embodiment: (1) much fewer spectral bins are involved in the prediction process, reducing complexity (40% in the given example, since only three-fifths of the bins are predicted); and (2) bins belonging to non-tonal background noise between signal harmonics are not affected by the prediction, which should increase prediction efficiency.
[0069] It should be emphasized that the harmonic spacing values do not necessarily correspond to the actual instantaneous pitch of the input signal, but can represent fractions or multiples of the true pitch if doing so results in an overall improvement in the efficiency of the prediction process. It should also be emphasized that the harmonic spacing values do not necessarily reflect integer multiples of bin indexing or bandwidth units, but can include fractions of said units.
[0070] Subsequently, a preferred implementation in an MPEG-style audio coder is described.
[0071] Preferably, pitch-adaptive prediction is incorporated into MPEG-2 AAC (ISO / IEC 13818-7 "Information technology—Part 7: Advanced Audio Coding (AAC)," 2006) or, utilizing a predictor similar to that in AAC, into the MPEG-H 3D audio codec (ISO / IEC 23008-3 "Information technology—High efficiency coding, part 3: 3D audio," 2015). Specifically, a one-bit flag can be written to and read from the respective bitstream for each frame and channel that is not independently coded (no flag is sent for independent frame channels, as prediction can be disabled to ensure independence). If the flag is set to 1, eight more bits can be written and read. These eight bits represent the quantized version of the harmonic frequency spacing value (e.g., an index to the harmonic spacing) for the given frame and channel. Using the spacing values derived from the quantized version using either a linear or nonlinear mapping function, the prediction process can be performed using the method shown in Figure 2 according to one embodiment. Preferably, only bins located within a maximum distance of 1.5 bins around the harmonic grid are considered for prediction. For example, if the harmonic spacing value indicates a harmonic line at bin index 47.11, only bins at indices 46, 47, and 48 are predicted. However, the maximum distance can be defined differently, either a priori fixed for all channels and frames or separately for each frame and channel based on the high-frequency spacing value.
[0072] 4 shows a schematic block diagram of a decoder 200 for decoding the coded audio signal 120. The decoder 200 is configured to decode the coded audio signal 120 in a transform or filter bank domain 204, the decoder 200 is configured to analyze the coded audio signal 120 to obtain coded spectral coefficients 206_t0_f1 to 206_t0_f6 of the audio signal for a current frame 208_t0 and coded spectral coefficients 206_t-1_f0 to 206_t-1_f6 for at least one previous frame 208_t-1, and the decoder 200 is configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients separated by at least one coded spectral coefficient.
[0073] In an embodiment, decoder 200 may be configured to apply predictive decoding to a plurality of individual coded spectral coefficients that are separated by at least one coded spectral coefficient, such as, for example, to two individual coded spectral coefficients that are separated by at least one coded spectral coefficient. Furthermore, decoder 200 may be configured to apply predictive decoding to a plurality of coded spectral coefficient groups (each group including at least two coded spectral coefficients) that are separated by at least one coded spectral coefficient, such as, for example, to two coded spectral coefficient groups that are separated by at least one coded spectral coefficient. Furthermore, decoder 200 may be configured to apply predictive decoding to a plurality of individual coded spectral coefficients and / or coded spectral coefficient groups that are separated by at least one coded spectral coefficient, such as, for example, to at least one individual coded spectral coefficient and at least one coded spectral coefficient group that are separated by at least one coded spectral coefficient.
[0074] 4, the decoder 200 is configured to determine six coded spectral coefficients 206_t0_f1 to 206_t0_f6 for the current frame 208_t0 and six coded spectral coefficients 206_t-1_f1 to 206_t-1_f6 for the previous frame 208_t-1. As a result, the decoder 200 is configured to selectively apply predictive decoding to each coded second spectral coefficient 206_t0_f2 of the current frame and to a coded spectral coefficient group consisting of the coded fourth and fifth spectral coefficients 206_t0_f4 and 206_t0_f5 of the current frame 208_t0. As can be seen, the coded spectral coefficient groups consisting of the individual coded second spectral coefficients 206_t0_f2 and the coded fourth and fifth spectral coefficients 206_t0_f4 and 206_t0_f5 are separated from each other by the coded third spectral coefficient 206_t0_f3.
[0075] Note that the term "selectively" in this specification refers to applying predictive decoding to (only) selected coded spectral coefficients. That is, predictive decoding is not applied to all coded spectral coefficients, but rather only to selected individual coded spectral coefficients or groups of coded spectral coefficients, i.e., selected individual coded spectral coefficients and / or groups of coded spectral coefficients that are separated from each other by at least one coded spectral coefficient. That is, predictive decoding is not applied to at least one coded spectral coefficient that separates the selected individual coded spectral coefficients or groups of coded spectral coefficients.
[0076] In an embodiment, the decoder 200 may be configured not to apply predictive decoding to at least one coded spectral coefficient 206_t0_f3 that separates the individual coded spectral coefficients 206_t0_f2 or the spectral coefficient groups 206_t0_f4 and 206_t0_f5.
[0077] The decoder 200 may be configured to entropy decode the coded spectral coefficients to obtain quantized prediction errors for the spectral coefficients 206_t0_f2, 2016_t0_f4, and 206_t0_f5 to which predictive decoding is to be applied, and a quantized spectral coefficient 206_t0_f3 for at least one spectral coefficient to which predictive coding is not to be applied. As a result, the decoder 200 may be configured to apply the quantized prediction errors to a plurality of individual predicted spectral coefficients 210_t0_f2 or groups of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 to obtain, for the current frame 208_t0, decoded spectral coefficients that are associated with the coded spectral coefficients 206_t0_f2, 206_t0_f4, and 206_t0_f5 to which predictive decoding is to be applied.
[0078] For example, the decoder 200 may be configured to obtain a quantized second prediction error for the quantized second spectral coefficient 206_t0_f2 and apply the quantized second prediction error to the predicted second spectral coefficient 210_t0_f2 to obtain a decoded second spectral coefficient associated with the coded second spectral coefficient 206_t0_f4, and the decoder 200 may be configured to apply the quantized second prediction error to the predicted second spectral coefficient 210_t0_f2 to obtain a decoded fourth spectral coefficient associated with the coded fourth spectral coefficient 206_t0_f4. The decoder 200 may be configured to obtain a quantized fourth prediction error for the quantized fifth spectral coefficient 206_t0_f5 and to apply the quantized fifth prediction error to the predicted fifth spectral coefficient 210_t0_f5 to obtain a decoded fifth spectral coefficient associated with the coded fifth spectral coefficient 206_t0_f5.
[0079] Furthermore, the decoder 200 may be configured to determine a plurality of individual predicted spectral coefficients 210_t0_f2 or predicted spectral coefficient groups 210_t0_f4 and 210_t0_f5 for the current frame 208_t0 based on a corresponding plurality of individual coded spectral coefficients 206_t-1_f2 (e.g., using a plurality of previously decoded spectral coefficients associated with the plurality of individual coded spectral coefficients 206_t-1_f2) or coded spectral coefficient groups 206_t-1_f4 and 206_t-1_f5 (e.g., using a previously decoded spectral coefficient group associated with the coded spectral coefficients 206_t-1_f4 and 206_t-1_f5) of the previous frame 208_t-1.
[0080] For example, the decoder 200 can be configured to determine the predicted second spectral coefficient 210_t0_f2 of the current frame 208_t0 using a previously decoded (quantized) second spectral coefficient associated with the coded second spectral coefficient 206_t-1_f2 of the previous frame 208_t-1, the predicted fourth spectral coefficient 210_t0_f4 of the current frame 208_t0 using a previously decoded (quantized) fourth spectral coefficient associated with the coded fourth spectral coefficient 206_t-1_f4 of the previous frame 208_t-1, and the predicted fifth spectral coefficient 210_t0_f5 of the current frame 208_t0 using a previously decoded (quantized) fifth spectral coefficient associated with the coded fifth spectral coefficient 206_t-1_f5 of the previous frame 208_t-1.
[0081] Further, the decoder 200 may be configured to derive prediction coefficients from the interval values, and the decoder 200 may be configured to calculate a plurality of individual predicted spectral coefficients 210_t0_f2 or predicted spectral coefficient groups 210_t0_f4 and 210_t0_f5 for the current frame 208_t0 using a corresponding plurality of previously decoded individual spectral coefficients or previously decoded spectral coefficient groups of at least two previous frames 208_t-1 and 208_t-2 and using the derived prediction coefficients.
[0082] For example, the decoder 200 may be configured to derive prediction coefficients 212_f2 and 214_f2 for the coded second spectral coefficient 206_t0_f2 from the interval value, prediction coefficients 212_f4 and 214_f4 for the coded fourth spectral coefficient 206_t0_f4 from the interval value, and prediction coefficients 212_f5 and 214_f5 for the coded fifth spectral coefficient 206_t0_f5 from the interval value.
[0083] It should be noted that the decoder 200 may be configured to decode the coded audio signal 120 to obtain quantized prediction errors instead of a plurality of individual quantized spectral coefficients or groups of quantized spectral coefficients for a plurality of individual coded spectral coefficients or groups of coded spectral coefficients to which predictive decoding is applied.
[0084] Furthermore, the decoder 200 may be configured to decode the coded audio signal 120 to obtain quantized spectral coefficients separating a plurality of individual spectral coefficients or groups of spectral coefficients such that the coded spectral coefficients 206_t0_f2 or the coded spectral coefficient groups 206_t0_f4 and 206_t0_f5 for which the quantized prediction errors are obtained alternate with the coded spectral coefficients 206_t0_f3 or the coded spectral coefficient groups for which the quantized spectral coefficients are obtained.
[0085] The decoder 200 can be configured to provide a decoded audio signal 220 using decoded spectral coefficients associated with the coded spectral coefficients 206_t0_f2, 206_t0_f4 and 206_t0_f5 to which predictive decoding is applied, and using entropy decoded spectral coefficients associated with the coded spectral coefficients 206_t0_f1, 206_t0_f3 and 206_t0_f6 to which predictive decoding is not applied.
[0086] In an embodiment, the decoder 200 may be configured to obtain an interval value, and the decoder 200 may be configured to select, based on the interval value, a plurality of individual coded spectral coefficients 206_t0_f2 or groups of coded spectral coefficients 206_t0_f4 and 206_t0_f5 to which predictive decoding is applied.
[0087] As already mentioned above in connection with the description of the corresponding encoder 100, the spacing value may, for example, be the spacing (or distance) between two characteristic frequencies of the audio signal. Furthermore, the spacing value may be an integer spectral coefficient (or an index of spectral coefficients) that approximates the spacing between two characteristic frequencies of the audio signal. Of course, the spacing value may also be a fraction or multiple of the integer spectral coefficient that represents the spacing between two characteristic frequencies of the audio signal.
[0088] The decoder 200 may be configured to select, for predictive decoding, individual spectral coefficients or groups of spectral coefficients that are spectrally arranged according to a harmonic grid defined by spacing values. The harmonic grid defined by spacing values may represent a periodic spectral distribution (equidistant spacing) of harmonics in the audio signal 102. That is, the harmonic grid defined by spacing values may be a series of spacing values that represent equidistant spacing of harmonics of the audio signal 102.
[0089] Furthermore, the decoder 200 can be configured to select, for predictive coding, spectral coefficients (e.g., only those spectral coefficients) whose spectral indices fall within a range (e.g., a predetermined and variable range) equal to or around a plurality of spectral indices derived based on the interval value. As a result, the decoder 200 can be configured to set the width of the range depending on the interval value.
[0090] In an embodiment, the encoded audio signal includes the interval value or an encoded version thereof (e.g., a parameter from which the interval value can be directly derived), and the decoder 200 may be configured to extract the interval value or an encoded version thereof from the encoded audio signal to obtain the interval value.
[0091] Alternatively, the decoder 200 may be configured to determine the interval values itself, i.e., the encoded audio signal does not include the interval values, in which case the decoder 200 may be configured to determine the instantaneous fundamental frequency (of the encoded audio signal 120 representing the audio signal 102) and to derive the interval values from the instantaneous fundamental frequency or a fraction or multiple thereof.
[0092] In an embodiment, the decoder 200 may be configured to select the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied such that the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied alternates periodically with a period with a tolerance of + / -1 spectral coefficient between the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied and the spectral coefficients separating the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive decoding is applied.
[0093] In an embodiment, the audio signal 102 represented by the encoded audio signal 120 includes at least two harmonic signal elements, and the decoder 200 is configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients 206_t0_f2 or groups of coded spectral coefficients 206_t0_f4 and 206_t0_f5 representing the at least two harmonic signal elements or the spectral environment around the at least two harmonic signal elements of the audio signal 102. The spectral environment around the at least two harmonic signal elements may be, for example, + / - 1, 2, 3, 4, or 5 spectral elements.
[0094] As a result, the decoder 200 can be configured to identify at least two harmonic signal elements and to selectively apply predictive decoding to a plurality of individual coded spectral coefficients 206_t0_f2 or coded spectral coefficient groups 206_t0_f4 and 206_t0_f5 associated with the identified harmonic signal elements (e.g., representing or surrounding the identified harmonic signal elements).
[0095] Alternatively, the encoded audio signal 120 may include information (e.g., spacing values) identifying at least two harmonic signal elements, in which case the decoder 200 may be configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients 206_t0_f2 or coded spectral coefficient groups 206_t0_f4 and 206_t0_f5 associated with, e.g., representing, or surrounding, the identified harmonic signal element.
[0096] In both of the above alternative methods, the decoder 200 may be configured not to apply predictive decoding to a plurality of individual coded spectral coefficients 206_t0_f3, 206_t0_f1 and 206_t0_f6 or groups of coded spectral coefficients that do not represent the spectral environment of at least two harmonic signal components or at least two harmonic signal components of the audio signal 102.
[0097] That is, the decoder 200 can be configured not to apply predictive decoding to a plurality of individual coded spectral coefficients 206_t0_f3, 206_t0_f1, 206_t0_f6 or groups of coded spectral coefficients that belong to non-tonal background noise between signal harmonics of the audio signal 102.
[0098] 5 shows a flowchart of a method 300 for encoding an audio signal according to one embodiment, comprising a step 302 of determining spectral coefficients of the audio signal for a current frame and at least one previous frame, and a step 304 of selectively applying predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients separated by at least one spectral coefficient.
[0099] 6 shows a flowchart of a method 400 for decoding an encoded audio signal according to one embodiment, comprising a step 402 of analyzing the encoded audio signal to obtain coded spectral coefficients of the audio signal for a current frame and at least one previous frame, and a step 404 of selectively applying predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients separated by at least one coded spectral coefficient.
[0100] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent a description of a corresponding method, with blocks or devices corresponding to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps may be performed by (or using) a hardware apparatus, such as, for example, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the most important method steps may be performed by such an apparatus.
[0101] The encoded audio signals according to the present invention can be stored on a digital storage medium or can be transmitted over a transmission medium, such as a wireless or wired transmission medium, such as the Internet.
[0102] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. For example, they can be implemented using digital storage media, such as floppy disks, DVDs, Blu-rays, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memories, having electronically readable control signals stored thereon and cooperating (or capable of cooperating) with a programmable computer system to execute the respective methods. Thus, the digital storage media can be computer-readable.
[0103] Some embodiments of the present invention include a data carrier having electronically readable control signals, the data carrier being capable of interfacing with a programmable computer system such that one of the methods described herein is performed.
[0104] In general, embodiments of the present invention can be implemented as a computer program product with program code that operates to perform one of the methods when the computer program product is run on a computer, and that can be stored on a machine-readable carrier, for example.
[0105] Further embodiments comprise the computer program for performing one of the methods described herein, stored on a machine readable carrier.
[0106] Thus, a method embodiment of the present invention is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.
[0107] A further embodiment of the method according to the invention is consequently a data carrier (or digital storage medium, or computer-readable medium) comprising thereon a computer program for performing one of the methods described herein. The data carrier, digital storage medium or recording medium is typically tangible and / or non-transitory.
[0108] A further embodiment of the inventive method is, consequently, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or sequence of signals may for example be adapted to be transmitted over a data communication connection, for example via the Internet.
[0109] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.
[0110] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.
[0111] Further embodiments according to the invention include an apparatus or system configured to transmit (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may be, for example, a computer, a mobile device, a memory device, etc. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.
[0112] In some embodiments, a programmable logic device (e.g., a field programmable gate array) may be used to perform some or all of the functionality of the methods described herein. In some embodiments, a field programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus.
[0113] The apparatus described herein may be implemented using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0114] The methods described herein may be performed using a hardware apparatus, or using a computer, or using a combination of a hardware apparatus and a computer.
[0115] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Consequently, it is the intention to be limited only by the scope of the appended claims, and not by the specific details represented by the description and illustration of the embodiments herein.
Claims
1. 1. An encoder (100) for encoding an audio signal (102), the encoder (100) being configured to encode the audio signal (102) in a transform domain or a filter bank domain (104), the encoder being configured to determine spectral coefficients (106_t0_f1:106_t0_f6; 106_t-1_f1:106_t-1_f6) of the audio signal (102) for a current frame (108_t0) and at least one previous frame (108_t-1), the encoder 1. An encoder (100) configured to selectively apply predictive coding to a plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5), the encoder (100) configured to determine a spacing value, and the encoder (100) configured to select the plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5) to which predictive coding is applied based on the spacing value.
2. The encoder (100) of claim 1, wherein the spacing value is a harmonic spacing value representing the spacing between harmonics.
3. 3. The encoder (100) of claim 1 or 2, wherein the plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5) are separated by at least one spectral coefficient (106_t0_f3).
4. 4. The encoder (100) of claim 3, wherein the predictive coding is not applied to the individual spectral coefficients (106_t0_f2) or to the at least one spectral coefficient (106_t0_f3) separating the groups of spectral coefficients (106_t0_f4, 106_t0_f5).
5. 5. The encoder (100) of claim 1, wherein the encoder (100) is configured to predictively encode the plurality of individual spectral coefficients (106_t0_f2) or the plurality of groups of spectral coefficients (106_t0_f4, 106_t0_f5) of the current frame (108_t0) by encoding a prediction error between the plurality of individual predicted spectral coefficients (110_t0_f2) or the plurality of groups of predicted spectral coefficients (110_t0_f4, 110_t0_f5) of the current frame (108_t0) and the plurality of individual spectral coefficients (106_t0_f2) or the plurality of groups of spectral coefficients (106_t0_f4, 106_t0_f5) of the current frame (108_t0).
6. 6. The encoder (100) of claim 5, wherein the encoder (100) is configured to derive prediction coefficients from the interval values, and wherein the encoder (100) is configured to calculate the plurality of individual predicted spectral coefficients (110_t0_f2) or predicted spectral coefficient groups (110_t0_f4, 110_t0_f5) for the current frame (108_t0) using a corresponding plurality of individual spectral coefficients (106_t-2_f2, 106_t-1_f2) or corresponding spectral coefficient groups (106_t-2_f4, 106_t-1_f4; 106_t-2_f5, 106_t-1_f5) of at least two previous frames (108_t-2, 108_t-1) and using the derived prediction coefficients.
7. 6. The encoder (100) of claim 5, wherein the encoder (100) is configured to determine the plurality of individual predicted spectral coefficients (110_t0_f2) or predicted groups of spectral coefficients (110_t0_f4, 110_t0_f4) for the current frame (108_t0) using the plurality of individual spectral coefficients (106_t-1_f2) or the groups of spectral coefficients (106_t-1_f4, 106_t-1_f5) of a corresponding quantized version of the previous frame (108_t-1).
8. 8. The encoder of claim 7, wherein the encoder is configured to derive prediction coefficients from the interval values, and wherein the encoder is configured to calculate the plurality of individual predicted spectral coefficients or groups of predicted spectral coefficients for the current frame using the plurality of individual spectral coefficients or groups of spectral coefficients of corresponding quantized versions of at least two previous frames and using the derived prediction coefficients.
9. 9. The encoder (100) of claim 6 or 8, wherein the encoder (100) is configured to provide an encoded audio signal (120), the encoded audio signal (120) not including the prediction coefficients or an encoded version thereof.
10. 10. The encoder (100) of claim 5, configured to provide an encoded audio signal (120), the encoded audio signal (120) comprising a quantized version of the prediction errors instead of a quantized version of the plurality of individual spectral coefficients (106_t0_f2) or the plurality of groups of spectral coefficients (106_t0_f4, 106_t0_f5) for the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive coding is applied.
11. 11. The encoder of claim 10, wherein the encoded audio signal comprises quantized versions of the spectral coefficients to which no predictive coding is applied, such that spectral coefficients or groups of spectral coefficients, the quantized versions of which are included in the encoded audio signal, alternate with spectral coefficients or groups of spectral coefficients, the quantized versions of which are provided without predictive coding.
12. 12. The encoder (100) of claim 1, configured to determine an instantaneous fundamental frequency of the audio signal (102) and to derive the interval value from the instantaneous fundamental frequency or a fraction or multiple thereof.
13. 13. The encoder (100) of claim 1, configured to select, for predictive coding, individual spectral coefficients or groups of spectral coefficients (116_1:116_6) spectrally arranged according to a harmonic grid defined by the spacing values.
14. 13. The encoder (100) of claim 1, configured to select spectral coefficients whose spectral indices are equal to or fall within a range around a plurality of spectral indices derived based on the spacing value for predictive coding.
15. The encoder (100) of claim 14, wherein the encoder (100) is configured to set the width of the range depending on the interval value.
16. 16. The encoder (100) of claim 1, configured to select the plurality of individual spectral coefficients or groups of spectral coefficients (116_1:116_6) to which predictive coding is applied such that the plurality of individual spectral coefficients or groups of spectral coefficients (116_1:116_6) to which predictive coding is applied alternates cyclically with a periodicity with a tolerance of + / -1 spectral coefficient.
17. 17. The encoder (100) of claim 1, wherein the audio signal (102) comprises at least two harmonic signal elements (124_1:124_6), and the encoder (100) is configured to selectively apply predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients (116_1:116_6) representing the at least two harmonic signal elements (124_1:124_6) of the audio signal (102) or a spectral environment surrounding the at least two harmonic signal elements (124_1:124_6).
18. 18. The encoder (100) of claim 17, wherein the encoder (100) is configured to not apply predictive coding to the at least two harmonic signal elements (124_1:124_6) of the audio signal (102) or to a plurality of individual spectral coefficients or groups of spectral coefficients (118_1:118_5) that do not represent the spectral environment of the at least two harmonic signal elements (124_1:124_6).
19. 19. The encoder (100) of claim 17 or 18, wherein the encoder (100) is configured to not apply predictive coding to individual spectral coefficients or groups of spectral coefficients (118_1:118_5) belonging to non-tonal background noise between the signal harmonics (124_1:124_6).
20. 20. The encoder (100) of claim 17, wherein the spacing value is a harmonic spacing value indicative of a spectral spacing between the at least two harmonic signal elements (124_1:124_6) of the audio signal (102), the harmonic spacing value indicative of a plurality of individual spectral coefficients or groups of spectral coefficients (116_1:116_6) representing the at least two harmonic signal elements (124_1:124_6) of the audio signal (102).
21. 21. The encoder (100) of claim 1, configured to provide an encoded audio signal (120), the encoder (100) configured to include the interval value or an encoded version thereof in the encoded audio signal (120).
22. The encoder (100) of any one of claims 1 to 21, wherein the spectral coefficients are spectral bins.
23. A decoder (200) for decoding an encoded audio signal (120), the decoder (200) being configured to decode the encoded audio signal (120) in a transform domain or a filter bank domain (204), the decoder (200) being configured to analyze the encoded audio signal (120) to obtain coded spectral coefficients (206_t0_f1:206_t0_f6;206_t-1_f1:206_t-1_f6) of the audio signal (120) for a current frame (208_t0) and at least one previous frame (208_t-1).
1. A decoder (200) comprising: a decoder (200) configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients (206_t0_f2) or groups of coded spectral coefficients (206_t0_f4, 206_t0_f5); the decoder (200) configured to obtain a spacing value; and the decoder (200) configured to select the plurality of individual coded spectral coefficients (206_t0_f2) or groups of coded spectral coefficients (206_t0_f4, 206_t0_f5) to which predictive decoding is applied based on the spacing value.
24. 24. The decoder (200) of claim 23, wherein the spacing values are harmonic spacing values representing spacing between harmonics.
25. 25. The decoder (200) of claim 24, wherein the plurality of individual coded spectral coefficients (206_t0_f2) or groups of coded spectral coefficients (206_t0_f4, 206_t0_f5) are separated by at least one coded spectral coefficient (206_t0_f3).
26. 26. The decoder (200) of claim 25, wherein the predictive decoding is not applied to at least one spectral coefficient (206_t0_f3) separating the individual spectral coefficients (206_t0_f2) or the groups of spectral coefficients (206_t0_f4, 206_t0_f5).
27. the decoder (200) is configured to entropy decode the coded spectral coefficients to obtain quantized prediction errors for the spectral coefficients (206_t0_f2, 206_t0_f4, 206_t0_f5) to which predictive combining is to be applied, and quantized spectral coefficients for the spectral coefficients (206_t0_f3) to which predictive combining is not to be applied, 27. A decoder (200) as claimed in any one of claims 24 to 26, wherein the decoder (200) is configured to apply the quantized prediction error to a plurality of individual predicted spectral coefficients (210_t0_f2) or groups of predicted spectral coefficients (210_t0_f4, 210_t0_f5) to obtain, for the current frame (208_t0), decoded spectral coefficients that are associated with the coded spectral coefficients (206_t0_f2, 206_t0_f4, 206_t0_f5) to which predictive decoding is applied.
28. 28. The decoder (200) of claim 27, wherein the decoder (200) is configured to determine the plurality of individual predicted spectral coefficients (210_t0_f2) or predicted spectral coefficient groups (210_t0_f4, 210_t0_f5) for the current frame (208_t0) based on a corresponding plurality of the individual coded spectral coefficients (206_t-1_f2) or coded spectral coefficient groups (206_t-1_f4, 206_t-1_f5) of the previous frame (208_t-1).
29. 29. The decoder of claim 28, wherein the decoder is configured to derive prediction coefficients from the interval values, and wherein the decoder is configured to calculate the plurality of predicted individual spectral coefficients or predicted groups of spectral coefficients for the current frame using a corresponding plurality of previously decoded individual spectral coefficients or previously decoded groups of spectral coefficients of at least two previous frames and using the derived prediction coefficients.
30. 30. The decoder (200) according to any one of claims 24 to 29, wherein the decoder (200) is configured to decode the coded audio signal (120) to obtain quantized prediction errors instead of a plurality of individual quantized spectral coefficients or groups of quantized spectral coefficients for the plurality of individual coded spectral coefficients (206_t0_f2) or groups of coded spectral coefficients (206_t0_f4, 206_t0_f5) to which predictive decoding is applied.
31. 31. The decoder (200) of claim 30, configured to decode the coded audio signal (120) to obtain quantized spectral coefficients for coded spectral coefficients (206_t0_f3) to which no predictive coding is applied, such that coded spectral coefficients (206_t0_f2) or coded spectral coefficient groups (206_t0_f4, 206_t0_f5) for which quantized prediction errors are obtained alternate with coded spectral coefficients (206_t0_f3) or coded spectral coefficient groups for which quantized spectral coefficients are obtained.
32. 30. The decoder (200) of any one of claims 22 to 29, wherein the decoder (200) is configured to select, for predictive coding, individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) spectrally arranged according to a harmonic grid defined by the spacing value.
33. 33. The decoder (200) of any one of claims 24 to 32, wherein the decoder (200) is configured to select spectral coefficients whose spectral indices are equal to or fall within a range around a plurality of spectral indices derived based on the spacing value for predictive decoding.
34. 34. The decoder (200) of claim 33, wherein the decoder (200) sets the width of the range depending on the interval value.
35. 35. A decoder (200) according to any one of claims 24 to 34, wherein the encoded audio signal (120) comprises the interval value or an encoded version thereof, and the decoder (200) is configured to extract the interval value or the encoded version thereof from the encoded audio signal (120) to obtain the interval value.
36. The decoder (200) of any one of claims 24 to 34, wherein the decoder (200) is configured to determine the interval value.
37. 37. The decoder (200) of claim 36, wherein the decoder (200) is configured to determine an instantaneous fundamental frequency and to derive the interval value from the instantaneous fundamental frequency or a fraction or multiple thereof.
38. 38. The decoder (200) according to any one of claims 24 to 37, wherein the decoder (200) is configured to select the plurality of individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) to which predictive decoding is applied such that the plurality of individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) to which predictive decoding is applied alternate periodically with a periodicity with a tolerance of + / - 1 spectral coefficient.
39. 39. The decoder (200) of any one of claims 24 to 38, wherein the audio signal (102) represented by the encoded audio signal (120) comprises at least two harmonic signal elements (124_1:124_6), and the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients representing the at least two harmonic signal elements (124_1:124_6) of the audio signal (102) or a spectral environment surrounding the at least two harmonic signal elements (124_1:124_6).
40. 40. The decoder (200) of claim 39, configured to identify the at least two harmonic signal elements (124_1:124_6) and to selectively apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients associated with the identified harmonic signal elements (124_1:124_6).
41. 40. The decoder (200) of claim 39, wherein the encoded audio signal (120) comprises the spacing value or a coded version thereof, the spacing value identifying the at least two harmonic signal elements (124_1:124_6), and the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients associated with the identified harmonic signal elements (124_1:124_6).
42. 42. The decoder (200) of claim 39, wherein the decoder (200) is configured to not apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients that are not representative of the spectral environment of the at least two harmonic signal elements (124_1:124_6) or the at least two harmonic signal elements (124_1:124_6) of the audio signal.
43. 43. The decoder (200) of any one of claims 39 to 42, wherein the decoder (200) is configured to not apply predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients belonging to non-tonal background noise between signal harmonics (124_1: 124_6) of the audio signal.
44. 44. The decoder (200) of any one of claims 24 to 43, wherein the encoded audio signal (120) comprises the spacing values or coded versions thereof, the spacing values being harmonic spacing values, the harmonic spacing values indicating a plurality of individual coded spectral coefficients or groups of coded spectral coefficients representing at least two harmonic signal elements (124_1:124_6) of the audio signal (102).
45. 45. A decoder (200) according to any one of claims 24 to 44, wherein the spectral coefficients are spectral bins.
46. A method (300) for encoding an audio signal in a transform or filterbank domain, the method comprising: determining (302) spectral coefficients of the audio signal for a current frame and at least one previous frame; determining an interval value; Selectively applying predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients (304), wherein the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive coding is applied are selected based on the spacing value. A method comprising:
47. A method (400) for decoding an encoded audio signal in a transform or filterbank domain, the method comprising: analyzing (402) the encoded audio signal to obtain encoded spectral coefficients of the audio signal for a current frame and at least one previous frame; Obtaining an interval value; Selectively applying predictive decoding to a plurality of individual coded spectral coefficients or groups of coded spectral coefficients (404), wherein the plurality of individual coded spectral coefficients or groups of coded spectral coefficients to which predictive decoding is applied are selected based on the spacing value. A method comprising:
48. 48. A computer program for carrying out the method of claim 46 or 47.
49. An encoder (100) for encoding an audio signal (102), the encoder (100) being configured to encode the audio signal (102) in a transform domain or a filter bank domain (104), the encoder being configured to determine spectral coefficients (106_t0_f1:106_t0_f6; 106_t-1_f1:106_t-1_f6) of the audio signal (102) for a current frame (108_t0) and at least one previous frame (108_t-1), a coder (100) configured to selectively apply predictive coding to a plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5), the encoder (100) configured to determine a spacing value, and the encoder (100) configured to select the plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5) to which predictive coding is applied based on the spacing value; The encoder (100) is configured to select, for predictive coding, individual spectral coefficients or groups of spectral coefficients (116_1:116_6) spectrally arranged according to a harmonic grid defined by the spacing values.
50. A decoder (200) for decoding an encoded audio signal (120), the decoder (200) being configured to decode the encoded audio signal (120) in a transform domain or a filter bank domain (204), the decoder (200) analyzing the encoded audio signal (120) to obtain coded spectral coefficients (206_t0_f1:206_t0_f6;206_t-1_f1:206_t-1_f6) of the audio signal (120) for a current frame (208_t0) and at least one previous frame (208_t-1). the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual coded spectral coefficients (206_t0_f2) or groups of coded spectral coefficients (206_t0_f4:206_t0_f5), the decoder (200) is configured to obtain a spacing value, and the decoder (200) is configured to select the plurality of individual coded spectral coefficients (206_t0_f2) or groups of coded spectral coefficients (206_t0_f4, 206_t0_f5) to which predictive decoding is applied based on the spacing value, The decoder (200) is configured to select, for predictive decoding, individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) spectrally arranged according to a harmonic grid defined by the spacing values.