Audio encoder and audio decoder and corresponding methods
By using selective predictive coding, predictive coding is applied only to the spectral coefficients at integer multiples of pitch, solving the problems of high computational latency and complexity in existing technologies, and achieving more efficient audio encoding and decoding.
Patent Information
- Application Number
- CN202110984953.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2015-06-17
- Filing Date
- 2016-03-07
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2036-05-28
AI Technical Summary
Existing audio coding techniques suffer from long computational latency and high complexity at low bit rates. In particular, in frequency domain prediction, the backward adaptability of the predictor leads to inaccurate values between the encoder and decoder, and the prediction efficiency of noise components is limited.
Selective predictive coding is adopted, which applies predictive coding only to spectral coefficients located at integer multiples of pitch and transmits prediction information through spacing values, thereby avoiding the prediction of noise signal components and reducing computational complexity.
It improves the efficiency of frequency domain prediction, reduces computational complexity, and reduces prediction errors, making it suitable for low-latency communication scenarios.
Smart Images

Figure CN114067812B_ABST
Abstract
Description
[0001] This application is a divisional application of Fraunhofer Association for the Promotion of Applied Scientific Research, filed on March 7, 2016, with application number 201680015022.2, entitled "Audio Encoder and Audio Decoder and Corresponding Method". Technical Field
[0002] The embodiments relate to audio coding, and more particularly to a method and apparatus for encoding audio signals using predictive coding, and a method and apparatus for decoding encoded audio signals using predictive decoding. Preferred embodiments relate to a method and apparatus for pitch-adaptive spectral prediction. More preferred embodiments relate to perceptual coding of pitch audio signals using frequency-domain inter-frame prediction tools via transform coding. Background Technology
[0003] To improve the quality of encoded tone signals, especially at low bit rates, modern audio transform encoders employ very long transforms and / or long predictions or pre-filtering / post-filtering. However, long transforms imply long computational delays, which are undesirable for low-latency communication scenarios. Therefore, predictors with very low latency based on instantaneous fundamental pitch have recently gained popularity. The IETF (Internet Engineering Task Force) Opus codec utilizes pitch-adaptive pre- and post-filtering in its frequency-domain CELT (Confined Energy Overlap Transform) coding path [J.M. Valin, K. Vos, and T. Terriberry, "Definition of the Opus audio codec," 2012, IETF RFC 6716. http: / / tools.ietf.org / html / rfc67161.], and the 3GPP (3rd Generation Partnership Project) EVS (Enhanced Voice Services) codec provides a perceptually improved long harmonic post-filter for the transformed decoded signal [3GPP TS 26.443, "Codec for Enhanced Voice Services (EVS)," Release 12, Dec. 2014.]. Both approaches operate on the fully decoded signal waveform in the time domain, making frequency-selective application difficult and / or computationally expensive (both schemes only provide simple low-pass filters for some frequency selectivity). Therefore, a popular alternative to Long Time Prediction (LTP) or Pre-Filtering / Post-Filtering (PPF) is provided through Frequency Domain Prediction (FDP), as supported in MPEG-2 AAC [ISO / IEC 13818-7, "Information technology – Part 7: Advanced Audio Coding (AAC)" 2006.]. While promoting frequency selectivity, this approach has its own drawbacks, as described below.
[0004] The FDP method described above has two drawbacks compared to other tools. First, the FDP method requires high computational complexity. Specifically, in the worst case of prediction across all scaling factor bands, at least second-order (i.e., channel transform bins from the previous two frames) linear prediction coding is applied to hundreds of spectral bins for each frame and channel [ISO / IEC 13818-7, "Information technology – Part 7: Advanced Audio Coding (AAC)," 2006.]. Second, the FDP method involves a limited overall prediction gain. More precisely, because noise components between predictable harmonic and tonal spectral parts are also predicted, the prediction efficiency is limited, and errors are introduced because these noise components are often unpredictable.
[0005] This high complexity stems from the predictor's backward adaptation. This means that the prediction coefficients for each bin must be calculated based on previously transmitted bins. Therefore, numerical inaccuracies between the encoder and decoder can lead to reconstruction errors attributable to divergent prediction coefficients. To overcome this problem, bit-to-bit adaptation must be guaranteed to be identical. Furthermore, adaptation must be continuously performed to keep the prediction coefficients up-to-date, even if predictor groups are deactivated in some frames. Summary of the Invention
[0006] Therefore, the object of the present invention is to provide a concept for encoding audio signals and / or decoding encoded audio signals that avoids at least one of the aforementioned problems (e.g., both) and results in a more efficient and computationally inexpensive implementation.
[0007] This objective is achieved through independent claims.
[0008] Beneficial implementations are proposed by the dependent claims.
[0009] An embodiment provides an encoder for encoding audio signals. The encoder is used to encode audio signals in a transform domain or filter bank domain, wherein the encoder is used to determine spectral coefficients of the audio signal for a current frame and at least one previous frame, wherein the encoder is used to selectively apply predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients, wherein the encoder is used to determine a spacing value, and wherein the encoder is used to select the plurality of individual spectral coefficients or groups of spectral coefficients to which predictive coding is applied based on the spacing value, which can be transmitted as side information along with the encoded audio signal.
[0010] Other embodiments provide a decoder for decoding an encoded audio signal (e.g., encoded by the encoder described above). The decoder is used to decode the encoded audio signal in a transform domain or filter bank domain, wherein the decoder is used to parse the encoded audio signal to obtain encoded spectral coefficients of the audio signal for the current frame and at least one previous frame, and wherein the decoder is used to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or a group of encoded spectral coefficients, wherein the decoder can be used to select the plurality of individual encoded spectral coefficients or the group of encoded spectral coefficients to which predictive decoding is applied based on a transmitted spacing value.
[0011] According to the concept of the present invention, predictive coding is applied (only) to selected spectral coefficients. The spectral coefficients to which predictive coding is applied can be selected based on signal characteristics. For example, by not applying predictive coding to noise signal components, the errors introduced by predicting unpredictable noise signal components are avoided. Simultaneously, computational complexity can be reduced because predictive coding is applied only to selected spectral components.
[0012] For example, guided / adaptive spectral domain inter-frame prediction methods can be used to perform perceptual coding of tonal audio signals by means of transform coding (e.g., by an encoder). The efficiency of frequency domain prediction (FDP) can be increased and computational complexity reduced by applying prediction only to spectral coefficients around harmonic signal components located at, for example, integer multiples of the fundamental frequency or pitch (which can be signaled from the encoder to the decoder (e.g., as spacing values) in a suitable bitstream). Embodiments of the present invention can preferably be implemented or integrated into MPEG-H 3D audio codecs, but can also be applied to any audio transform coding system such as, for example, MPEG-2 AAC.
[0013] Other embodiments provide a method for encoding an audio signal in a transform domain or filter bank domain, the method comprising:
[0014] - Determine the spectral coefficients of the audio signal for the current frame and at least one previous frame;
[0015] - Determine the spacing value; and
[0016] - Predictive coding is selectively applied to multiple individual spectral coefficients or groups of spectral coefficients, wherein the selection of multiple individual spectral coefficients or groups of spectral coefficients to which predictive coding is applied is based on a spacing value.
[0017] Other embodiments provide a method for decoding an encoded audio signal in a transform domain or filter bank domain, the method comprising:
[0018] - The encoded audio signal is parsed to obtain the encoded spectral coefficients of the audio signal for the current frame and at least one previous frame;
[0019] - Obtain the spacing value; and
[0020] - Predictive decoding is selectively applied to a plurality of individual encoded spectral coefficients or a group of encoded spectral coefficients, wherein the plurality of individual encoded spectral coefficients or the group of encoded spectral coefficients to which predictive decoding is applied are selected based on a spacing value. Attached Figure Description
[0021] Hereinafter, embodiments of the present invention are described with reference to the accompanying drawings, wherein:
[0022] Figure 1 A schematic block diagram of an encoder for encoding audio signals according to an embodiment is shown;
[0023] Figure 2 The figure shows the amplitude of the audio signal plotted with respect to frequency for the current frame according to an embodiment and the selected spectral coefficients of the corresponding applied predictive coding;
[0024] Figure 3 The figure shows the amplitude of the audio signal plotted with respect to frequency for the current frame according to an embodiment and the corresponding spectral coefficients predicted according to MPEG-2 AAC.
[0025] Figure 4 A schematic block diagram of a decoder for decoding encoded audio signals according to an embodiment is shown;
[0026] Figure 5 A flowchart illustrating a method for encoding audio signals according to an embodiment;
[0027] Figure 6 A flowchart illustrating a method for decoding an encoded audio signal according to an embodiment is provided. Detailed Implementation
[0028] In the following description, equivalent or related elements, or elements having equivalent or related functions, are marked with equivalent or related reference numerals.
[0029] In the following description, numerous details are set forth to provide a more detailed explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other examples, well-known structures and devices are shown in block diagram form rather than in detail in order to avoid obscuring embodiments of the invention. Furthermore, unless otherwise specifically noted, features of the different embodiments described below may be combined with each other.
[0030] Figure 1 A schematic block diagram of an encoder 100 for encoding an audio signal 102 according to an embodiment is shown. The encoder 100 is used to encode the audio signal 102 in a transform domain or filter bank domain 104 (e.g., frequency domain or spectral domain), wherein the encoder 100 is used to determine spectral coefficients 106_t0_f1 to 106_t0_f6 of the audio signal 102 for the current frame 108_t0 and to determine spectral coefficients 106_t-1_f1 to 106_t-1_f6 of the audio signal for at least one previous frame 108_t-1. Additionally, encoder 100 is used to selectively apply predictive coding to a group of multiple individual spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5, wherein encoder 100 is used to determine a spacing value, wherein encoder 100 is used to select the group of multiple individual spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5 to which predictive coding is applied based on the spacing value.
[0031] In other words, encoder 100 is used to selectively apply predictive coding to a group of individual spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5 selected based on the spacing value transmitted as side information.
[0032] The spacing values can correspond to frequencies (e.g., the fundamental frequency of the harmonic modulation of (audio signal 102)), which, along with their integer multiples, define the center of the groups of all spectral coefficients for which predictions are applied: the first group can be centered at this frequency, the second group at twice this frequency, the third group at three times this frequency, and so on. Knowing these center frequencies enables the calculation of prediction coefficients used to predict the corresponding sinusoidal signal components (e.g., the fundamental and overtones of the harmonic signal). Therefore, complex and error-prone backward adaptation of prediction coefficients is no longer needed.
[0033] In this embodiment, encoder 100 can be used to determine a spacing value for each frame.
[0034] In an embodiment, a group of multiple individual spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5 may be separated by at least one spectral coefficient 106_t0_f3.
[0035] In an embodiment, encoder 100 can be used to apply predictive coding to a plurality of individual spectral coefficients separated by at least one spectral coefficient, such as to two individual spectral coefficients separated by at least one spectral coefficient. Additionally, encoder 100 can be used to apply predictive coding to a plurality of groups of spectral coefficients separated by at least one spectral coefficient (each group comprising at least two spectral coefficients), such as to two groups of spectral coefficients separated by at least one spectral coefficient. Furthermore, encoder 100 can be used to apply predictive coding to a plurality of individual spectral coefficients and / or groups of spectral coefficients separated by at least one spectral coefficient, such as to at least one individual spectral coefficient and at least one group of spectral coefficients separated by at least one spectral coefficient.
[0036] exist Figure 1 In the illustrated example, encoder 100 determines six spectral coefficients 106_t0_f1 to 106_t0_f6 for the current frame 108_t0 and six spectral coefficients 106_t-1_f1 to 106_t-1_f6 for the previous frame 108_t-1. Thus, encoder 100 selectively applies predictive coding to the individual second spectral coefficients 106_t0_f2 of the current frame and to a group of spectral coefficients consisting of the fourth and fifth spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0. As can be seen, the individual second spectral coefficients 106_t0_f2 and the group of spectral coefficients consisting of the fourth and fifth spectral coefficients 106_t0_f4 and 106_t0_f5 are separated from each other by a third spectral coefficient 106_t0_f3.
[0037] It should be noted that the term "selectivity" as used herein refers to applying predictive coding (only) to selected spectral coefficients. In other words, predictive coding is not required to be applied to all spectral coefficients, but only to selected individual spectral coefficients or groups of spectral coefficients, which may be separated by at least one spectral coefficient. In other words, predictive coding may be deactivated for at least one spectral coefficient that separates the selected individual spectral coefficients or groups of spectral coefficients.
[0038] In an embodiment, encoder 100 may be used to selectively apply predictive coding to a group of individual spectral coefficients 106_t0_f2 or 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 based on at least a group of corresponding individual spectral coefficients 106_t-1_f2 or 106_t-1_f4 and 106_t0_f5 of the previous frame 108_t-1.
[0039] For example, encoder 100 can be used to predictively encode the group of individual spectral coefficients 106_t0_f2 or 106_t0_f4 and 110_t0_f5 of the current frame 108_t0 by encoding the prediction error between the group of individual spectral coefficients 106_t0_f2 or 106_t0_f4 and 106_t0_f5 of the current frame (or their quantized versions).
[0040] exist Figure 1 In this process, encoder 100 encodes the prediction error between the predicted individual spectral coefficients 110_t0_f2 of the current frame 108_t0 and the individual spectral coefficients 106_t0_f2 of the current frame 108_t0, as well as the prediction error between the group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame and the group of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame. It also encodes the individual spectral coefficient 106_t0_f2 and the group of spectral coefficients composed of spectral coefficients 106_t0_f4 and 106_t0_f5.
[0041] In other words, the second spectral coefficient 106_t0_f2 is encoded by encoding the prediction error (or difference) between the predicted second spectral coefficient 110_t0_f2 and the (actual or determined) second spectral coefficient 106_t0_f2, wherein the fourth spectral coefficient 106_t0_f4 is encoded by encoding the prediction error (or difference) between the predicted fourth spectral coefficient 110_t0_f4 and the (actual or determined) fourth spectral coefficient 106_t0_f4, and wherein the fifth spectral coefficient 106_t0_f5 is encoded by encoding the prediction error (or difference) between the predicted fifth spectral coefficient 110_t0_f5 and the (actual or determined) fifth spectral coefficient 106_t0_f5.
[0042] In an embodiment, encoder 100 may be used to determine a plurality of predicted individual spectral coefficients 110_t0_f2 or a group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 by means of the corresponding actual version of a plurality of individual spectral coefficients 106_t-1_f2 or a group of spectral coefficients 106_t-1_f4 and 110_t0_f5 of the previous frame 108_t-1.
[0043] In other words, during the determination process described above, encoder 100 can directly use multiple actual individual spectral coefficients 106_t-1_f2 or groups of actual spectral coefficients 106_t-1_f4 and 106_t-1_f5 from the previous frame 108_t-1 (where 106_t-1_f2, 106_t-1_f4 and 106_t-1_f5 represent the original, unquantized spectral coefficients or groups of spectral coefficients, respectively), because they are obtained by encoder 100 so that the encoder can operate in the transform domain or filter bank domain 104.
[0044] For example, encoder 100 can be used to determine the second predicted spectral coefficient 110_t0_f2 of the current frame 108_t0 based on the unquantized version of the second spectral coefficient 106_t-1_f2 of the previous frame 108_t-1, the predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on the unquantized version of the fourth spectral coefficient 106_t-1_f4 of the previous frame 108_t-1, and the predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on the unquantized version of the fifth spectral coefficient 106_t-1_f5 of the previous frame.
[0045] Using this method, the predictive encoding and decoding scheme can exhibit a harmonic shaping of quantized noise, because the corresponding decoder (about...) Figure 4 (As described below) In the above determination step, only transmitted quantized versions of multiple individual spectral coefficients 106_t-1_f2 or multiple groups of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1 can be used for predictive decoding.
[0046] While this harmonic noise shaping, because it is, for example, performed conventionally in the time domain by long-term prediction (LTP) and can subjectively benefit predictive coding, may be undesirable in some cases as it can lead to unwanted, excessive tones being introduced into the decoded audio signal. For this reason, an alternative predictive coding scheme is described below that is fully synchronized with the corresponding decoding and similarly utilizes only any possible prediction gain without causing quantization noise shaping. According to this alternative coding embodiment, encoder 100 can be used to determine a plurality of predicted individual spectral coefficients 110_t0_f2 or a group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 for the current frame 108_t0 using the corresponding quantized version of a group of a plurality of individual spectral coefficients 106_t-1_f2 or 106_t-1_f4 and 110_t0_f5 of the previous frame 108_t-1.
[0047] For example, encoder 100 can be used to determine the second predicted spectral coefficient 110_t0_f2 of the current frame 108_t0 based on the corresponding quantization version of the second spectral coefficient 106_t-1_f2 of the previous frame 108_t-1, to determine the predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on the corresponding quantization version of the fourth spectral coefficient 106_t-1_f4 of the previous frame 108_t-1, and to determine the predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on the corresponding quantization version of the fifth spectral coefficient 106_t-1_f5 of the previous frame 108_t-1.
[0048] Additionally, encoder 100 can be used to derive prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5, and 114_f5 from the spacing value, and uses multiple individual spectral coefficients 106_t-1_f2 and 106_t-2_f2 or spectral coefficients 106_t-1_f4, 106_t-2_f4, and 114_f5 from at least two previous frames 108_t-1 and 108_t-2. The corresponding quantized versions of the groups 06_t-1_f5 and 106_t-2_f5, and the predicted coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5 and 114_f5, are used to calculate multiple predicted individual spectral coefficients 110_t0_f2 or groups of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 for the current frame 108_t0.
[0049] For example, encoder 100 can be used to: derive prediction coefficients 112_f2 and 114_f2 from the spacing value for the second spectral coefficient 106_t0_f2, derive prediction coefficients 112_f4 and 114_f4 from the spacing value for the fourth spectral coefficient 106_t0_f4, and derive prediction coefficients 112_f5 and 114_f5 from the spacing value for the fifth spectral coefficient 106_t0_f5.
[0050] For example, the prediction coefficients can be derived as follows: if the spacing value or its encoded version corresponds to frequency f0, then the center frequency of the Kth group of spectral coefficients enabling prediction is fc = K * f0. If the sampling frequency is fs and the transform jump size (shift between consecutive frames) is N, then the ideal predictor coefficients for a sinusoidal signal with frequency fc in the Kth group are assumed to be:
[0051] p1 = 2*cos(N*2*pi*fc / fs) and p2 = -1.
[0052] If, for example, the spectral coefficients 106_t0_f4 and 106_t0_f5 are within this group, then the prediction coefficients are:
[0053] 112_f4=112_f5=2*cos(N*2*pi*fc / fs) and 114_f4=114_f5=-1
[0054] For stability reasons, a damping factor d can be introduced to modify the prediction coefficients:
[0055] 112_f4'=112_f5'=d*2*cos(N*2*pi*fc / fs), 114_f4'=114_f5'=d 2 .
[0056] Since the spacing values are transmitted in the encoded audio signal 120, the decoder can obtain the exact same prediction coefficients 212_f4=212_f5=2*cos(N*2*pi*fc / fs) and 114_f4=114_f5=-1. If a damping factor is used, the coefficients can be modified accordingly.
[0057] as Figure 1 As indicated, encoder 100 can be used to provide an encoded audio signal 120. Thus, encoder 100 can be configured to include a quantized version of the prediction error in the encoded audio signal 120 for a group of individual spectral coefficients 106_t0_f2 or 106_t0_f4 and 106_t0_f5 to which predictive coding is applied. Additionally, encoder 100 can be configured not to include prediction coefficients 112_f2 to 114_f5 in the encoded audio signal 120.
[0058] Therefore, encoder 100 can use only prediction coefficients 112_f2 to 114_f5 to calculate the prediction error between a group of predicted individual spectral coefficients 110_t0_f2 or predicted spectral coefficients 110_t0_f4 and 110_t0_f5 and the group of predicted individual spectral coefficients 110_t0_f2 or predicted spectral coefficients 110_t0_f4 and 110_t0_f5 from the current frame therefrom, and the group of individual spectral coefficients 106_t0_f2 or predicted spectral coefficients 110_t0_f4 and 110_t0_f5, but the encoded audio signal 120 will not provide individual spectral coefficients 106_t0_f4 (or its quantized version) or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 (or their quantized version), nor will it provide prediction coefficients 112_f2 to 114_f5. Therefore, decoder (hereinafter regarding...) Figure 4 (Describing its embodiments) Prediction coefficients 112_f2 to 114_f5 can be derived from the spacing value for calculating multiple predicted individual spectral coefficients or groups of predicted spectral coefficients for the current frame.
[0059] In other words, encoder 100 can be configured to provide an encoded audio signal 120 that includes a quantized version of the prediction error, rather than a quantized version of the group of individual spectral coefficients 106_t0_f2 or 106_t0_f4 and 106_t0_f5, for which predictive coding is applied.
[0060] Additionally, encoder 100 can be used to provide an encoded audio signal 120 comprising a quantized version of spectral coefficient 106_t0_f3 separated from groups of individual spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5, such that there is an alternation between groups of spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5 (for which a quantized version of the prediction error is included in the encoded audio signal 120) and groups of spectral coefficients 106_t0_f3 or spectral coefficients (for which a quantized version is provided without predictive coding).
[0061] In an embodiment, encoder 100 may also be used to entropy encode a quantized version of the prediction error and a quantized version of spectral coefficient 106_t0_f3 that separates a group of individual spectral coefficients 106_t0_f2 or spectral coefficients 106_t0_f4 and 106_t0_f5, and to include the entropy-encoded version (instead of its unentropy-encoded version) in the encoded audio signal 120.
[0062] Figure 2 The graph shows the amplitude of the audio signal 102 plotted with respect to frequency for the current frame 108_t0. Additionally, in Figure 2 In the figure, the spectral coefficients in the transform domain or filter bank domain are determined by the encoder 100 for the current frame 108_t0 of the audio signal 102.
[0063] like Figure 2 As shown, encoder 100 can be used to selectively apply predictive coding to multiple groups 116_1 to 116_6 of spectral coefficients separated by at least one spectral coefficient. Specifically, in... Figure 2In the illustrated embodiment, encoder 100 selectively applies predictive coding to six groups 116_1 to 116_6 of the spectral coefficients, wherein each of the first five groups 116_1 to 116_5 comprises three spectral coefficients (e.g., the second group 116_2 comprises spectral coefficients 106_t0_f8, 106_t0_f9, and 106_t0_f10), and the sixth group 116_6 comprises two spectral coefficients. Therefore, these six groups 116_1 to 116_6 of the spectral coefficients are separated by five groups 118_1 to 118_5 of the spectral coefficients to which predictive coding is not applied.
[0064] In other words, such as Figure 2 As indicated, encoder 100 can be used to selectively apply predictive coding to groups 116_1 to 116_6 of spectral coefficients, such that there is an alternation between groups 116_1 to 116_6 of spectral coefficients to which predictive coding is applied and groups 118_1 to 118_5 of spectral coefficients to which predictive coding is not applied.
[0065] In this embodiment, encoder 100 can be used to determine the spacing value (indicated by arrows 122_1 and 122_2). Figure 2 (in the middle), wherein encoder 100 can be used to select multiple groups 116_1 to 116_6 (or multiple individual spectral coefficients) of spectral coefficients for which predictive coding is applied based on the spacing value.
[0066] The spacing value can be, for example, the spacing (or distance) between two characteristic frequencies of the audio signal 102, such as the peaks 124_1 and 124_2 of the audio signal. Alternatively, the spacing value can be an integer number (or index) of the spectral coefficients that approximate the spacing between two characteristic frequencies of the audio signal. Naturally, the spacing value can also be a real value, fraction, or multiple of the integer number of the spectral coefficients describing the spacing between two characteristic frequencies of the audio signal.
[0067] In an embodiment, encoder 100 can be used to determine the instantaneous fundamental frequency of an audio signal (102) and derive the spacing value from the instantaneous fundamental frequency or a fraction or multiple thereof.
[0068] For example, the first peak 124_1 of the audio signal 102 can be the instantaneous fundamental frequency (or pitch, or first harmonic) of the audio signal 102. Therefore, the encoder 100 can be used to determine the instantaneous fundamental frequency of the audio signal 102 and derive the spacing value from the instantaneous fundamental frequency or its fraction or multiple. In this case, the spacing value can be an integer number (or its fraction or multiple) of the spectral coefficients that approximate the spacing between the instantaneous fundamental frequency 124_1 and the second harmonic 124_2 of the audio signal 102.
[0069] Naturally, the audio signal 102 may include more than two harmonics. For example, shown in Figure 2 The audio signal 102 includes six harmonics 124_1 to 124_6 distributed in the spectrum such that the audio signal 102 includes harmonics at every integer multiple of the instantaneous fundamental frequency. Naturally, the audio signal 102 may also exclude all but some of the harmonics, such as the first, third, and fifth harmonics.
[0070] In an embodiment, encoder 100 can be used to select groups 116_1 to 116_6 (or individual spectral coefficients) of spectral coefficients arranged according to a harmonic grid defined by spacing values for predictive coding. Thus, the harmonic grid defined by spacing values describes the periodic spectral distribution (equidistant spacing) of harmonics in the audio signal 102. In other words, the harmonic grid defined by spacing values can be a sequence of spacing values describing the equidistant spacing of harmonics in the audio signal.
[0071] Additionally, encoder 100 can be used to select spectral coefficients (e.g., only those spectral coefficients) whose spectral indices are equal to or within a range (e.g., predetermined or variable) of the multiple spectral indices derived based on the spacing value for predictive coding.
[0072] The indices (or numbers) of the spectral coefficients representing the harmonics of audio signal 102 can be derived from the spacing values. For example, assuming the fourth spectral coefficient 106_t0_f4 represents the instantaneous fundamental frequency of audio signal 102 and assuming a spacing value of five, then a spectral coefficient with index nine can be derived based on the spacing value. (The last sentence appears to be incomplete and possibly refers to a different context.) Figure 2 As seen in the diagram, the resulting spectral coefficient with index nine (i.e., the ninth spectral coefficient 106_t0_f9) represents the second harmonic. Similarly, spectral coefficients with indices 14, 19, 24, and 29 can be derived to represent the third to sixth harmonics 124_3 to 124_6. However, not only can spectral coefficients with indices equal to the multiple spectral indices derived based on the spacing value be predicted and encoded, but spectral coefficients with indices within a given range surrounding the multiple spectral indices derived based on the spacing value can also be predicted and encoded. For example, as... Figure 2 As shown, the range can be three, so that rather than multiple individual spectral coefficients, multiple groups of spectral coefficients are selected for predictive coding.
[0073] Additionally, encoder 100 can be used to select groups 116_1 to 116_6 (or multiple individual spectral coefficients) of applied predictive coding spectral coefficients such that there is a periodic alternation between the groups 116_1 to 116_6 (or multiple individual spectral coefficients) of applied predictive coding spectral coefficients and the spectral coefficients separating the groups (or multiple individual spectral coefficients) of applied predictive coding spectral coefficients, with a period of + / -1 spectral coefficient tolerance. A + / -1 spectral coefficient tolerance may be required when the distance between two harmonics of audio signal 102 is not equal to an integer spacing value (an integer with respect to the index or number of the spectral coefficient) but rather a fraction or multiple thereof. This can also be used in... Figure 2 As can be seen, arrows 122_1 to 122_6 do not always point exactly to the center or middle of the corresponding spectral coefficients.
[0074] In other words, the audio signal 102 may include at least two harmonic signal components 124_1 to 124_6, wherein the encoder 100 may be used to selectively apply predictive coding to a plurality of groups 116_1 to 116_6 (or individual spectral coefficients) of spectral coefficients representing the spectral environment surrounding the at least two harmonic signal components 124_1 to 124_6 of the audio signal 102 or the spectral environment surrounding the at least two harmonic signal components 124_1 to 124_6. The spectral environment surrounding the at least two harmonic signal components 124_1 to 124_6 may be, for example, + / - 1, 2, 3, 4, or 5 spectral components.
[0075] Therefore, encoder 100 can be used to exclude predictive coding from applying predictive coding to those groups 118_1 to 118_5 (or multiple individual spectral coefficients) of spectral coefficients that do not represent the spectral environment of at least two harmonic signal components 124_1 to 124_6 of the audio signal 102. In other words, encoder 100 can be used to exclude predictive coding from applying predictive coding to multiple groups 118_1 to 118_5 (or individual spectral coefficients) of spectral coefficients belonging to non-tonal background noise between the signal harmonics 124_1 to 124_6.
[0076] Additionally, encoder 100 can be used to determine a harmonic spacing value that indicates the spectral spacing between at least two harmonic signal components 124_1 to 124_6 of audio signal 102, the harmonic spacing value indicating a plurality of individual spectral coefficients or a plurality of groups of spectral coefficients representing at least two harmonic signal components 124_1 to 124_6 of audio signal 102.
[0077] In addition, encoder 100 can be used to provide encoded audio signal 120 such that encoded audio signal 120 includes spacing values (e.g., one spacing value per frame) or (optionally) parameters from which spacing values can be directly derived.
[0078] This invention addresses two problems of the aforementioned FDP method by introducing a harmonic spacing value into the FDP process. This harmonic spacing value is signaled from the encoder (transmitter) 100 to each decoder (receiver), enabling them to operate in complete synchronization. The harmonic spacing value serves as an indicator of the instantaneous fundamental frequency (or pitch) of one or more spectra associated with the frame to be encoded, and identifies which spectral bins (spectral coefficients) should be predicted. More specifically, only those spectral coefficients surrounding harmonic signal components located at (with respect to their index) integer multiples of the fundamental pitch (as defined by the harmonic spacing value) should be predicted. Figure 2 and 3 This pitch adaptive prediction method is illustrated with a simple example, where Figure 3 The operation of the predictor, demonstrating the current level of technology in MPEG-2 AAC, not only predicts around the harmonic grating but also predicts every spectral block below a certain stop frequency, and in which... Figure 2 The illustration shows the same predictor with modifications according to the embodiment integrated to perform predictions only on those "tone" bins closest to the harmonic spacing grid.
[0079] Compare Figure 2 and Figure 3 Two advantages of the modifications according to the embodiments are revealed: (1) very few spectral bins are included in the prediction process, reducing complexity (by approximately 40% in the given example due to the prediction of only three-fifths of the bins), and (2) bins belonging to non-tonal background noise between harmonic signals are not affected by the prediction, which should increase the efficiency of the prediction.
[0080] It should be noted that the harmonic spacing value does not necessarily need to correspond to the actual instantaneous pitch of the input signal; it can also represent a fraction or multiple of the true pitch, as long as it improves the overall efficiency of the prediction process. Furthermore, it must be emphasized that the harmonic spacing value does not necessarily reflect an integer multiple of the bin index or bandwidth unit, but can include fractions of such units.
[0081] A preferred implementation of the MPEG-style audio encoder will then be described.
[0082] Pitch adaptive prediction is preferably integrated into MPEG-2 AAC [ISO / IEC 13818-7, "Information technology – Part 7: Advanced Audio Coding (AAC)," 2006.] or integrated into MPEG-H 3D audio codecs using a similar predictor as in AAC [ISO / IEC 23008-3, "Information technology – High efficiency coding, part 3: 3D audio" 2015.]. Specifically, for each frame and channel that is not independently coded, a one-bit flag can be written to and read from the individual bitstreams (for independent frame channels, the flag may not be transmitted because prediction can be disabled to ensure independence). If the flag is set to one, the other 8 bits can be written to and read from. These 8 bits represent a quantized version (e.g., an index) of the harmonic spacing value for a given frame and channel. The harmonic spacing value, derived from the quantized version using a linear or nonlinear mapping function, can be determined according to... Figure 2 The prediction process is implemented in the manner shown in the embodiment. Preferably, only cells within a maximum distance of 1.5 cells around the harmonic grid are predicted. For example, if the harmonic distance value indicates a harmonic line at cell index 47.11, then only cells at indices 46, 47, and 48 will be predicted. However, the maximum distance can be specified differently, either as a fixed prior for all channels and frames or based on harmonic spacing values for each frame and channel separately.
[0083] Figure 4 A schematic block diagram of a decoder 200 for decoding an encoded signal 120 is shown. The decoder 200 is used to decode the encoded audio signal 120 in a transform domain or filter bank domain 204, wherein the decoder 200 is used to parse the encoded audio signal 120 to obtain encoded spectral coefficients 206_t0_f1 to 206_t0_f6 of the audio signal for the current frame 208_t0 and to obtain encoded spectral coefficients 206_t-1_f0 to 206_t-1_f6 for at least one previous frame 208_t-1, and wherein the decoder 200 is used to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or a group of encoded spectral coefficients separated by at least one encoded spectral coefficient.
[0084] In an embodiment, the decoder 200 can be used to apply predictive decoding to a plurality of individual coded spectral coefficients separated by at least one coded spectral coefficient, such as to two individual coded spectral coefficients separated by at least one coded spectral coefficient. Additionally, the decoder 200 can be used to apply predictive decoding to a plurality of groups of coded spectral coefficients separated by at least one coded spectral coefficient (each group comprising at least two coded spectral coefficients), such as to two groups of coded spectral coefficients separated by at least one coded spectral coefficient. Furthermore, the decoder 200 can be used to apply predictive decoding to a plurality of individual coded spectral coefficients and / or groups of coded spectral coefficients separated by at least one coded spectral coefficient, such as to at least one individual coded spectral coefficient and at least one group of coded spectral coefficients separated by at least one coded spectral coefficient.
[0085] exist Figure 4 In the illustrated example, decoder 200 can be used to determine six encoded spectral coefficients 206_t0_f1 to 206_t0_f6 for the current frame 208_t0 and six encoded spectral coefficients 206_t-1_f1 to 206_t-1_f6 for the previous frame 208_t-1. Thus, decoder 200 is used to selectively apply predictive decoding to the individual second encoded spectral coefficients 206_t0_f2 of the current frame and to the group of encoded spectral coefficients consisting of the fourth and fifth encoded spectral coefficients 206_t0_f4 and 206_t0_f5 of the current frame 208_t0. As can be seen, the individual second encoded spectral coefficients 206_t0_f2 and the group of encoded spectral coefficients consisting of the fourth and fifth encoded spectral coefficients 206_t0_f4 and 206_t0_f5 are separated from each other by a third encoded spectral coefficient 206_t0_f3.
[0086] It should be noted that the term "selectivity" as used herein refers to applying predictive decoding (only) to selected coded spectral coefficients. In other words, predictive decoding is not necessarily applied to all coded spectral coefficients, but only to selected individual coded spectral coefficients or groups of coded spectral coefficients that are separated from each other by at least one coded spectral coefficient. In other words, predictive decoding is not applied to at least one coded spectral coefficient that separates the selected individual coded spectral coefficients or groups of coded spectral coefficients.
[0087] In an embodiment, the decoder 200 may be used to prevent predictive decoding from being applied to at least one encoded spectral coefficient 206_t0_f3 that separates the groups of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5.
[0088] Decoder 200 can be used to perform entropy decoding on the encoded spectral coefficients to obtain quantization prediction errors for spectral coefficients 206_t0_f2, 206_t0_f4, and 206_t0_f5 to be applied for predictive decoding, and to obtain quantized spectral coefficient 206_t0_f3 for at least one spectral coefficient to which predictive decoding will not be applied. Thus, decoder 200 can be used to apply the quantization prediction errors to a group of multiple individually predicted spectral coefficients 210_t0_f2 or predicted spectral coefficients 210_t0_f4 and 210_t0_f5, to obtain decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f2, 206_t0_f4, and 206_t0_f5 to which predictive decoding is applied for the current frame 208_t0.
[0089] For example, decoder 200 can be used to obtain a second quantization prediction error for a second quantized spectral coefficient 206_t0_f2 and apply the second quantization prediction error to the predicted second spectral coefficient 210_t0_f2 to obtain a second decoded spectral coefficient associated with the second encoded spectral coefficient 206_t0_f2, wherein decoder 200 can be used to obtain a fourth quantization prediction error for a fourth quantized spectral coefficient 206_t0_f4 and apply the fourth quantization prediction error to the predicted fourth spectral coefficient 210_t0_f4 to obtain a fourth decoded spectral coefficient associated with the fourth encoded spectral coefficient 206_t0_f4, and wherein decoder 200 can be used to obtain a fifth quantization prediction error for a fifth quantized spectral coefficient 206_t0_f5 and apply the fifth quantization prediction error to the predicted fifth spectral coefficient 210_t0_f5 to obtain a fifth decoded spectral coefficient associated with the fifth encoded spectral coefficient 206_t0_f5.
[0090] Additionally, the decoder 200 can be used to determine, for the current frame 208_t0, a plurality of predicted individual spectral coefficients 210_t0_f2 or a group of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 based on a plurality of individual encoded spectral coefficients 206_t-1_f2 corresponding to the previous frame 208_t-1 (e.g., using a plurality of previously decoded spectral coefficients associated with a plurality of individual encoded spectral coefficients 206_t-1_f2) or a group of encoded spectral coefficients 206_t-1_f4 and 206_t-1_f5 (e.g., using a group of previously decoded spectral coefficients associated with a group of encoded spectral coefficients 206_t-1_f4 and 206_t-1_f5).
[0091] For example, decoder 200 can be used to determine the second predicted spectral coefficient 210_t0_f2 of the current frame 208_t0 using the second spectral coefficient of the previously decoded (quantized) process associated with the second encoded spectral coefficient 206_t-1_f2 of the previous frame 208_t-1, to determine the fourth predicted spectral coefficient 210_t0_f4 of the current frame 208_t0 using the fourth spectral coefficient of the previously decoded (quantized) process associated with the fourth encoded spectral coefficient 206_t-1_f4 of the previous frame 208_t-1, and to determine the fifth predicted spectral coefficient 210_t0_f5 of the current frame 208_t0 using the fifth spectral coefficient of the previously decoded (quantized) process associated with the fifth encoded spectral coefficient 206_t-1_f5 of the previous frame 208_t-1.
[0092] Furthermore, decoder 200 can be used to derive prediction coefficients from the spacing value, and wherein decoder 200 can be used to calculate a plurality of predicted individual spectral coefficients 210_t0_f2 or a group of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 for the current frame 208_t0 using at least two previous frames 208_t-1 and 208_t-2 corresponding to a plurality of previously decoded individual spectral coefficients or a group of previously decoded spectral coefficients and the derived prediction coefficients.
[0093] For example, the decoder 200 can be used to: derive prediction coefficients 212_f2 and 214_f2 from the spacing value for the second encoded spectral coefficients 206_t0_f2, prediction coefficients 212_f4 and 214_f4 from the spacing value for the fourth encoded spectral coefficients 206_t0_f4, and prediction coefficients 212_f5 and 214_f5 from the spacing value for the fifth encoded spectral coefficients 206_t0_f5.
[0094] It should be noted that decoder 200 can be used to decode encoded audio signal 120 to obtain quantization prediction error for multiple individual encoded spectral coefficients or groups of encoded spectral coefficients for application prediction decoding, rather than multiple individual quantized spectral coefficients or groups of quantized spectral coefficients.
[0095] Additionally, the decoder 200 can be used to decode the encoded audio signal 120 to obtain quantized spectral coefficients that separate multiple individual spectral coefficients or groups of spectral coefficients, such that there is an alternation of groups of encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 (for which quantization prediction error is obtained) and groups of encoded spectral coefficients 206_t0_f3 or encoded spectral coefficients (for which quantization spectral coefficients are obtained).
[0096] Decoder 200 can be used to provide a decoded audio signal 220 using decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f2, 206_t0_f4 and 206_t0_f5 with applied predictive decoding and entropy-decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f1, 206_t0_f3 and 206_t0_f6 without applied predictive decoding.
[0097] In an embodiment, decoder 200 can be used to obtain a spacing value, wherein decoder 200 can be used to select, based on the spacing value, a group of multiple individually encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 for which predictive decoding is applied.
[0098] As mentioned in the above description of the corresponding encoder 100, the spacing value can be, for example, the spacing (or distance) between two characteristic frequencies of the audio signal. Alternatively, the spacing value can be an integer number (or index) of the spectral coefficients that approximates the spacing between two characteristic frequencies of the audio signal. Naturally, the spacing value can also be a fraction or multiple of the integer number of the spectral coefficients describing the spacing between two characteristic frequencies of the audio signal.
[0099] Decoder 200 can be used to select individual spectral coefficients or groups of spectral coefficients arranged spectrally according to a harmonic grating defined by spacing values for predictive decoding. The harmonic grating defined by spacing values can describe the periodic spectral distribution (equidistant spacing) of harmonics in audio signal 102. In other words, the harmonic grating defined by spacing values can be a sequence of spacing values describing the equidistant spacing of harmonics in audio signal 102.
[0100] Additionally, decoder 200 can be used to select spectral coefficients (e.g., only those spectral coefficients) whose spectral indices are equal to or lie within a range (e.g., a predetermined or variable range) surrounding the multiple spectral indices derived based on the spacing values, for predictive decoding. Thus, decoder 200 can be used to set the width of this range based on the spacing values.
[0101] In an embodiment, the encoded audio signal may include a spacing value or an encoded version thereof (e.g., a parameter from which the spacing value can be directly derived), wherein the decoder 200 may be used to extract the spacing value or an encoded version thereof from the encoded audio signal to obtain the spacing value.
[0102] Alternatively, decoder 200 can be used to determine the spacing value by itself, i.e., the encoded audio signal does not include the spacing value. In this case, decoder 200 can be used to determine the instantaneous fundamental frequency (representing the encoded audio signal 120 of audio signal 102) and derive the spacing value from the instantaneous fundamental frequency or its fractions or multiples.
[0103] In an embodiment, decoder 200 may be used to select a plurality of individual spectral coefficients or groups of spectral coefficients for applied predictive decoding such that there is a periodic alternation between the plurality of individual spectral coefficients or groups of spectral coefficients for applied predictive decoding and the spectral coefficients separating the plurality of individual spectral coefficients or groups of spectral coefficients for applied predictive decoding, with a period of + / -1 spectral coefficient tolerance.
[0104] In an embodiment, the audio signal 102 represented by the encoded audio signal 120 includes at least two harmonic signal components, wherein the decoder 200 is used to selectively apply predictive decoding to a group of a plurality of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 representing at least two harmonic signal components of the audio signal 102 or the spectral environment surrounding at least two harmonic signal components. The spectral environment surrounding at least two harmonic signal components may be, for example, + / - 1, 2, 3, 4 or 5 spectral components.
[0105] Thus, decoder 200 can be used to identify at least two harmonic signal components and selectively apply predictive decoding to a group of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 associated with the identified harmonic signal component (e.g., representing or surrounding the identified harmonic signal component).
[0106] Optionally, the encoded audio signal 120 may include information identifying at least two harmonic signal components (e.g., spacing values). In this case, the decoder 200 may be used to selectively apply predictive decoding to the group of those individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 associated with the identified harmonic signal component (e.g., representing or surrounding the identified harmonic signal component).
[0107] In the aforementioned alternative, decoder 200 may be used to avoid applying predictive decoding to a plurality of individually encoded spectral coefficients 206_t0_f3, 206_t0_f1, and 206_t0_f6 or a group of encoded spectral coefficients that do not represent at least two harmonic signal components or the spectral environment of at least two harmonic signal components of audio signal 102.
[0108] In other words, the decoder 200 can be used to avoid applying predictive decoding to those individual encoded spectral coefficients 206_t0_f3, 206_t0_f1, 206_t0_f6 or groups of encoded spectral coefficients between signal harmonics belonging to the audio signal 102.
[0109] Figure 5 A flowchart illustrating a method 300 for encoding an audio signal according to an embodiment is provided. Method 300 includes: a step 302 of determining spectral coefficients of an audio signal for a current frame or at least one previous frame; and a step 304 of selectively applying predictive coding to a plurality of individual spectral coefficients or groups of spectral coefficients separated by at least one spectral coefficient.
[0110] Figure 6 A flowchart illustrating a method 400 for decoding an encoded audio signal according to an embodiment is provided. Method 400 includes: a step 402 of parsing the encoded audio signal to obtain encoded spectral coefficients of the audio signal for a current frame and at least one previous frame; and a step 404 of selectively applying predictive decoding to a plurality of individual encoded spectral coefficients or a group of encoded spectral coefficients separated by at least one encoded spectral coefficient.
[0111] Although some aspects have been described in the context of the apparatus, it is clear that these aspects also represent a description of the corresponding method, where blocks or devices correspond to method steps or features of method steps. Similarly, aspects described in the context of method steps also represent a description of corresponding blocks or entries or features of the corresponding apparatus. Some or all of the method steps may be performed by (or using) hardware devices (e.g., microprocessors, programmable computers, or electronic circuits). In some embodiments, one or more of the most important method steps may be performed by such devices.
[0112] The encoded audio signals of this invention can be stored on a digital storage medium or transmitted on a transmission medium such as a wireless transmission medium or a wired transmission medium such as the Internet.
[0113] Depending on certain implementation requirements, embodiments of the present invention can be implemented in hardware or software. This implementation can be performed using a digital storage medium (e.g., floppy disk, DVD, Blu-ray, CD, ROM, PROM, EPROM, EEPROM, or flash memory) having electronically readable control signals stored thereon, which cooperates (or is capable of cooperating with) a programmable computer system to perform its respective methods. Therefore, the digital storage medium can be computer-readable.
[0114] Some embodiments of the invention include a data carrier having electronically readable control signals, which is capable of cooperating with a programmable computer system to perform one of the methods described herein.
[0115] Generally, embodiments of the present invention can be implemented as a computer program product containing program code, which, when run on a computer, can operate to perform a method. The program code may, for example, be stored on a machine-readable medium.
[0116] Other embodiments include a computer program for performing one of the methods described herein, which is stored on a machine-readable medium.
[0117] In other words, an embodiment of the inventive method is therefore a computer program with program code that, when run on a computer, performs one of the methods described herein.
[0118] Other embodiments of the method of the present invention are therefore data carriers (or digital storage media or computer-readable media) comprising, on which a computer program for performing one of the methods described herein is recorded. Data carriers, digital storage media, or computer-readable media are typically tangible and / or non-transitory.
[0119] Other embodiments of the method of the present invention are therefore data streams or signal sequences representing a computer program for performing one of the methods described herein. The data streams or signal sequences can be transmitted, for example, via a data communication connection (e.g., via the Internet).
[0120] Other embodiments include a computational component, such as a computer or a programmable logic device, for or suitable for performing one of the methods described herein.
[0121] Other embodiments include a computer for performing the methods described herein, the computer having a computer program installed thereon.
[0122] Other embodiments of the invention include means or systems for transmitting (e.g., electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device, or a similar device. The means or system may, for example, include a file server for transmitting the computer program to the receiver.
[0123] In some embodiments, a programmable logic device (e.g., a field-programmable gate array) may be used to perform some or all of the functions of the methods described herein. In some embodiments, the field-programmable gate array may cooperate with a microprocessor to perform one of the methods described herein. Generally, this method is preferably performed by any hardware device.
[0124] The apparatus described herein may be implemented using hardware devices, computers, or a combination of hardware devices and computers.
[0125] The methods described herein can be performed using hardware devices, computers, or a combination of hardware devices and computers.
[0126] The embodiments described above are merely illustrative of the principles of the invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, this is intended to limit the scope of the invention solely by the scope of the appended claims and not by the manner in which the embodiments are described herein are presented.
Claims
1. A decoder (200) for decoding an encoded audio signal (120), wherein the decoder (200) is configured to decode the encoded audio signal (120) in a transform domain or a filter bank domain (204), wherein the decoder (200) is configured to parse the encoded audio signal (120) to obtain encoded spectral coefficients (206_t0_f1:206_t0_f6; 206_t-1_f1:206_t-1_f6) of the audio signal (120) for a current frame (208_t0) and at least one previous frame (208_t-1), and wherein the decoder (200) is configured to selectively apply a prediction decoding to a plurality of individual encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5), wherein the decoder (200) is configured to obtain a pitch value, wherein the decoder (200) is configured to select the plurality of individual encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5) to which a prediction decoding is applied based on the pitch value, wherein the decoder (200) is configured to select individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) that are spectrally arranged according to a harmonic grid defined by the pitch value for the prediction decoding.
2. The decoder (200) according to claim 1, wherein the pitch value is a harmonic pitch value that describes a pitch between harmonics.
3. The decoder (200) according to claim 1, wherein the plurality of individual encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5) are separated by at least one encoded spectral coefficient (206_t0_f3).
4. The decoder (200) according to claim 3, wherein a prediction decoding is not applied to the at least one spectral coefficient (206_t0_f3) that separates the individual spectral coefficients (206_t0_f2) or the groups of spectral coefficients (206_t0_f4, 206_t0_f5).
5. The decoder (200) according to claim 1, wherein the decoder (200) is configured to entropy decode the encoded spectral coefficients to obtain quantized prediction errors for spectral coefficients (206_t0_f2, 206_t0_f4, 206_t0_f5) to which a prediction decoding is to be applied and to obtain quantized spectral coefficients for spectral coefficients (206_t0_f3) to which a prediction decoding is not to be applied; and wherein the decoder (200) is configured to apply the quantized prediction error to a plurality of predicted individual spectral coefficients (210_t0_f2) or a group of predicted spectral coefficients (210_t0_f4, 210_t0_f5) to obtain, for the current frame (208_t0), decoded spectral coefficients associated with the encoded spectral coefficients (206_t0_f2, 206_t0_f4, 206_t0_f5) to which the prediction decoding is applied.
6. The decoder (200) according to claim 5, wherein the decoder (200) is configured to determine the plurality of predicted individual spectral coefficients (210_t0_f2) or the group of predicted spectral coefficients (210_t0_f4, 210_t0_f5) for the current frame (208_t0) based on a corresponding plurality of individual encoded spectral coefficients (206_t-1_f2) or a group of encoded spectral coefficients (206_t-1_f4, 206_t-1_f5) of a previous frame (208_t-1).
7. The decoder (200) according to claim 6, wherein the decoder (200) is configured to derive prediction coefficients from the pitch value, and wherein the decoder (200) is configured to calculate the plurality of predicted individual spectral coefficients (210_t0_f2) or the group of predicted spectral coefficients (210_t0_f4, 210_t0_f5) for the current frame (208_t0) using a corresponding plurality of previously decoded individual spectral coefficients or a group of previously decoded spectral coefficients of at least two previous frames and using the derived prediction coefficients.
8. The decoder (200) according to claim 1, wherein the decoder (200) is configured to decode the encoded audio signal (120) such that a quantized prediction error is obtained for a plurality of individual encoded spectral coefficients (206_t0_f2) or a group of encoded spectral coefficients (206_t0_f4, 206_t0_f5) to which the prediction decoding is applied instead of a plurality of individual quantized spectral coefficients or a group of quantized spectral coefficients.
9. The decoder (200) according to claim 1, wherein the decoder (200) is configured to select spectral coefficients for the prediction decoding whose spectral indices are equal to or within a range around a plurality of spectral indices derived based on the pitch value.
10. The decoder (200) according to claim 9, wherein the decoder (200) is configured to set a width of the range in dependence on the pitch value.
11. The decoder (200) according to claim 1, wherein the encoded audio signal (120) comprises the pitch value or an encoded version of the pitch value, wherein the decoder (200) is configured to extract the pitch value or the encoded version of the pitch value from the encoded audio signal (120) to obtain the pitch value.
12. The decoder (200) of claim 1, wherein the decoder (200) is configured to determine the pitch value.
13. The decoder (200) of claim 12, wherein the decoder (200) is configured to determine an instantaneous fundamental frequency and derive the pitch value from the instantaneous fundamental frequency or a fraction or multiple of the instantaneous fundamental frequency.
14. The decoder (200) of claim 1, wherein the audio signal (102) represented by the encoded audio signal (120) comprises at least two harmonic signal components (124_1:124_6), wherein the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients associated with at least two harmonic signal components (124_1:124_6) or a spectral environment surrounding the at least two harmonic signal components (124_1:124_6) of the audio signal (102).
15. The decoder (200) of claim 14, wherein the decoder (200) is configured to identify the at least two harmonic signal components (124_1:124_6) and selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients associated with the identified harmonic signal components (124_1:124_6).
16. The decoder (200) of claim 14, wherein the encoded audio signal (120) comprises the pitch value or an encoded version of the pitch value, wherein the pitch value identifies the at least two harmonic signal components (124_1:124_6), wherein the decoder (200) is configured to selectively apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients associated with the identified harmonic signal components (124_1:124_6).
17. The decoder (200) of claim 14, wherein the decoder (200) is configured to not apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients that do not represent at least two harmonic signal components (124_1:124_6) of the audio signal or a spectral environment of the at least two harmonic signal components (124_1:124_6).
18. The decoder (200) of claim 14, wherein the decoder (200) is configured to not apply predictive decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients that belong to a non-tonal background noise between signal harmonics (124_1:124_6) of the audio signal.
19. The decoder (200) of claim 1, wherein the encoded audio signal (120) comprises the pitch value or an encoded version of the pitch value, wherein the pitch value is a harmonic pitch value, the harmonic pitch value indicating a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients representing at least two harmonic signal components (124_1:124_6) of the audio signal (102).
20. The decoder (200) of claim 1, wherein a spectral coefficient is a spectral bin.
21. A method (400) for decoding an encoded audio signal in a transform domain or a filter bank domain, the method comprising: parsing (402) the encoded audio signal to obtain encoded spectral coefficients of an audio signal for a current frame and at least one previous frame; obtaining a pitch value; selectively applying (404) a prediction decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients, wherein the plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients to which the prediction decoding is applied is selected based on the pitch value; and selecting individual spectral coefficients (206_t0_f2) or groups of spectral coefficients (206_t0_f4, 206_t0_f5) that are spectrally arranged according to a harmonic grid defined by the pitch value for the prediction decoding.
22. A computer program product for performing the method of claim 21.
Citation Information
Patent Citations
Audio encoder and audio decoder and corresponding methods
CN114067813B
Voice transcoder
US20040153316A1
Scan patterns for interlaced video content
US20050078754A1
Systems, methods, apparatus, and computer-readable media for dynamic bit allocation
US20120029925A1
Model based prediction in a critically sampled filterbank
WO2014108393A1