Encoder, decoder, encoding method and decoding method for frequency domain long-term prediction of tonal signals for audio encoding

By using the FDLMSP concept in the MDCT domain to model and predict the harmonic components of audio signals, the problems of harmonic component overlap and insufficient stability in frequency domain prediction are solved, improving coding efficiency and prediction accuracy, especially in low-latency audio coding.

CN115004298BActive Publication Date: 2026-01-09FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980103473.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-27
Publication Date
2026-01-09
Estimated Expiration
2039-11-27

AI Technical Summary

Technical Problem

Existing audio coding techniques suffer from severe harmonic component overlap and insufficient prediction stability in frequency domain prediction, especially in high fundamental frequency signals, resulting in low coding efficiency.

Method used

The concept of Frequency Domain Least Mean Square Prediction (FDLMSP) is adopted to directly model and predict the harmonic components of the audio signal in the MDCT domain. The linear equations are solved by the least mean square algorithm to estimate the harmonic parameters and encode them.

Benefits of technology

It improves the coding efficiency of audio encoding, reduces the bit rate, performs exceptionally well in low-latency audio encoding scenarios, and enhances the stability and prediction accuracy of the encoder.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115004298B_ABST
    Figure CN115004298B_ABST
Patent Text Reader

Abstract

An encoder (100) for encoding a current frame of an audio signal from one or more previous frames of the audio signal is provided according to embodiments. The one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. To generate an encoding of the current frame, the encoder (100) will determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. Further, the encoder (100) will determine the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of spectral coefficients constituted by three or more of the plurality of spectral coefficients of each of the one or more previous frames of the audio signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to audio signal encoding, audio signal processing, and audio signal decoding, and more particularly, to an apparatus and method for long-term frequency domain prediction of audio-encoded tone signals. Background Technology

[0002] In the field of audio coding, prediction is used to remove redundancy from audio signals. By subtracting the predicted data from the original data and then quantizing and encoding the residual, which typically exhibits low entropy, the bit rate used for transmitting and storing audio signals can be reduced [1]. Long-term prediction (LTP) is a prediction method designed to remove periodic components from audio signals [2].

[0003] In the Moving Picture Experts Group (MPEG)-2 Advanced Audio Coding (AAC) standard, the Improved Discrete Cosine Transform (MDCT) is used as the time-frequency transform of the perceptual audio encoder with backward adaptive LTP[3].

[0004] Figure 4 The structure of a transform-aware audio encoder with backward adaptive LTP is shown. Figure 4 The audio encoder includes an MDCT unit 410, a psychoacoustic model unit 420, a pitch estimation unit 430, a long-term prediction unit 440, a quantizer 450, and a quantizer reconstruction unit 460.

[0005] like Figure 4 As shown, the prediction unit takes the reconstructed MDCT frame as input. To perform conventional temporal long-term prediction (TDLTP), the MDCT coefficients of the reconstructed signal need to be transformed into the time domain first. Then, the predicted time-domain segment is transformed back into the MDCT domain for residual calculation.

[0006] MDCT uses an overlap analysis window that reduces blocking effects and still provides perfect reconstruction at the synthesis step in the inverse transform via an overlap-addition (OLA) process [4]. Since alias-free reconstruction of the second half of the current frame requires the first half of the future frame [4], the prediction lag needs to be carefully selected [2].

[0007] If only fully reconstructed samples from the buffer are used for prediction, a delay of an integer multiple of the pitch period can exist between the selected previous pitch lag and the pitch lag to be predicted. Due to the non-stationarity of audio signals, a longer delay reduces the stability of the prediction. For signals with high fundamental frequencies and short pitch periods, the negative impact of this additional delay on prediction is even more pronounced.

[0008] The concept of frequency-domain prediction (FDP) operating directly in the MDCT domain (see also

[13] ) is presented in [5]. In this approach, each harmonic component of a tonal signal is processed individually during prediction. A prediction for a bin in the current frame is obtained by computing the sinusoidal progression of the spectral neighboring bins in the previous frame.

[0009] However, when the frequency resolution of these MDCT coefficients is relatively low with respect to the fundamental frequency of a tonal signal, the harmonic components can severely overlap each other over bins, resulting in poor performance of this frequency-domain approach. SUMMARY

[0010] An encoder according to an embodiment is provided for encoding a current frame of an audio signal from one or more previous frames of the audio signal. The one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. To generate an encoding of the current frame, the encoder will determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. Further, the encoder will determine the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first set of spectral coefficients consisting of three or more of the plurality of spectral coefficients of each of the one or more previous frames of the audio signal.

[0011] Further, a decoder according to an embodiment is provided for reconstructing a current frame of an audio signal. One or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The decoder will receive an encoding of the current frame. The decoder will determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. The two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first set of spectral coefficients consisting of three or more of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal. Further, the decoder will reconstruct the current frame from the encoding of the current frame and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.

[0012] Furthermore, an apparatus for frame loss concealment according to an embodiment is provided. One or more previous frames of an audio signal precede a current frame of the audio signal. The current frame and each of the one or more previous frames comprises one or more harmonic components of the audio signal, wherein the current frame and each of the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The apparatus is to determine an estimate of two harmonic parameters for each of the one or more harmonic components of a first previous frame of the one or more previous frames, wherein the two harmonic parameters for each of the one or more harmonic components of the first previous frame depend on a first set of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal. If the apparatus does not receive the current frame, or if the apparatus receives the current frame in a corrupted state, the apparatus is to reconstruct the current frame from the estimate of the two harmonic parameters for each of the one or more harmonic components of the first previous frame.

[0013] Furthermore, a method for encoding a current frame of an audio signal from one or more previous frames of the audio signal according to an embodiment is provided. The one or more previous frames precede the current frame. The current frame and each of the one or more previous frames comprises one or more harmonic components of the audio signal. The current frame and each of the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. To generate an encoding of the current frame, the method comprises determining an estimate of two harmonic parameters for each of the one or more harmonic components of a first previous frame of the one or more previous frames. Determining the estimate of the two harmonic parameters for each of the one or more harmonic components of the first previous frame is performed using a first set of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames of the audio signal.

[0014] Furthermore, a method for reconstructing a current frame of an audio signal according to an embodiment is provided. One or more previous frames of the audio signal precede the current frame. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal. Each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The method comprises receiving an encoding of the current frame. Furthermore, the method comprises determining an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, wherein the two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first set of spectral coefficients consisting of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal. Furthermore, the method comprises reconstructing the current frame from the encoding of the current frame and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.

[0015] Furthermore, a method for frame loss concealment according to an embodiment is provided. One or more previous frames of an audio signal precede a current frame of the audio signal, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain. The method comprises determining an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, wherein the two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first set of spectral coefficients consisting of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal. Furthermore, the method comprises reconstructing the current frame from the two harmonic parameters for each of the one or more harmonic components of the most previous frame, if the current frame is not received or if the current frame is received in a corrupted state.

[0016] Furthermore, a computer program according to an embodiment for implementing one of the above methods when the computer program is executed by a computer or signal processor is provided.

[0017] Long-term prediction (LTP) is traditionally used to predict signals that have a certain periodicity in the time domain. In the case of transform coding with backward adaptation in an audio encoder, the decoder unit has only the frequency coefficients at hand, so an inverse transform is needed before prediction. Embodiments provide a frequency domain least mean square prediction (FDLMSP) concept that operates directly in the modified discrete cosine transform (MDCT) domain and that, for example, significantly reduces the bit rate of audio coding, even at very low frequency resolution. Thus, some embodiments can be used, for example, in a transform codec to enhance coding efficiency, especially in low-delay audio coding scenarios.

[0018] Some embodiments provide a frequency domain least mean square prediction (FDLMSP) concept that performs LTP directly in the MDCT domain. However, this new concept does not predict individually on each frequency bin, but models the harmonic components of a tonal signal in the transform domain using a system of real-valued linear equations. The prediction is done after solving the linear equations in least mean square (LMS). Then, based on the phase progression properties of the harmonics, the parameters of the harmonics are used to predict the current frame. It should be noted that this prediction concept can also be applied to other real-valued linear transforms or filter banks, such as different types of discrete cosine transform (DCT) or polyphase quadrature filter (PQF) [6].

[0019] A signal model is introduced, the harmonic component estimation and prediction process is explained in detail, experiments to evaluate the FDLMSP concept compared to TD LTP and FD P are described, and results are shown and discussed. BRIEF DESCRIPTION OF DRAWINGS

[0020] In the following, embodiments of the present application will be described in more detail with reference to the accompanying drawings, in which:

[0021] Figure 1 An encoder for encoding a current frame of an audio signal from one or more previous frames of the audio signal according to an embodiment is shown.

[0022] Figure 2 A decoder for decoding an encoding of a current frame of an audio signal according to an embodiment is shown.

[0023] Figure 3 A system according to an embodiment is shown.

[0024] Figure 4 The structure of a transform perceptual audio encoder with backward adaptive LTP is shown.

[0025] Figure 5 The bit rate saved on single tone prediction with different prediction bandwidths and MDCT lengths using the three prediction concepts is shown.

[0026] Figure 6 Bitrates saved on six different items with bandwidth limited to 4 kHz, MDCT frame length of 64 and 512 in four different working modes are shown.

[0027] Figure 7 An apparatus for frame loss concealment according to an embodiment is shown.

[0028] Figure 8 A schematic block diagram of an encoder for encoding an audio signal according to the FDP prediction concept according to an example is shown.

[0029] Figure 9 A schematic block diagram of a decoder 201 for decoding an encoded signal 120 according to the FDP prediction concept according to an example is shown. DETAILED DESCRIPTION

[0030] Figure 1 An encoder 100 for encoding a current frame of an audio signal from one or more previous frames of the audio signal according to an embodiment is shown.

[0031] The one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain.

[0032] To generate an encoding of the current frame, the encoder 100 will determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. Furthermore, the encoder 100 will determine the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of spectral coefficients of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames of the audio signal.

[0033] The most previous frame can for example be the most previous with respect to the current frame.

[0034] For example, the most previous frame can be (referred to as) an immediately previous frame. For example, the immediately previous frame can directly precede the current frame.

[0035] The current frame comprises one or more harmonic components of the audio signal. Each of the one or more previous frames can comprise one or more harmonic components of the audio signal. It is assumed that the fundamental frequencies of the one or more harmonic components of the current frame and of the one or more previous frames are the same.

[0036] According to an embodiment, the encoder 100 can for example be configured to estimate the two harmonic parameters of each of the one or more harmonic components of the most previous frame without using a second set of spectral coefficients of one or more other spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames.

[0037] According to an embodiment, the encoder 100 can for example be configured to determine the gain factor and the residual signal as an encoding of the current frame from the fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames and from the estimate of the two harmonic parameters of each of the one or more harmonic components of the most previous frame. The encoder 100 can for example be configured to generate the encoding of the current frame such that the encoding of the current frame comprises the gain factor and the residual signal.

[0038] In an embodiment, the encoder 100 can for example be configured to determine the estimate of the two harmonic parameters of each of the one or more harmonic components of the current frame from the estimate of the two harmonic parameters of each of the one or more harmonic components of the most previous frame and from the fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames. It can for example be assumed that the fundamental frequency is constant over the current frame and the one or more previous frames.

[0039] According to an embodiment, the two harmonic parameters of each of the one or more harmonic components are a first parameter of a cosine sub-component for each of the one or more harmonic components and a second parameter of a sine sub-component for each of the one or more harmonic components.

[0040] In an embodiment, the encoder 100 can for example be configured to estimate the two harmonic parameters of each of the one or more harmonic components of the most previous frame by solving a system of linear equations comprising at least three equations, wherein each of the at least three equations depends on spectral coefficients of a first set of spectral coefficients of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames.

[0041] According to an embodiment, the encoder 100 can for example be configured to solve the system of linear equations using a least mean square algorithm.

[0042] According to an embodiment, the system of linear equations is defined by:

[0043]

[0044] wherein,

[0045]

[0046] wherein γ1 indicates a first spectral band of one or more harmonic components of the one or more harmonic components of the most previous frame having a lowest harmonic component frequency of the one or more harmonic components, wherein γ H indicates a second spectral band of one or more harmonic components of the one or more harmonic components of the most previous frame having a highest harmonic component frequency of the one or more harmonic components, wherein r is an integer, r≥0.

[0047] In an embodiment, r≥1.

[0048] According to an embodiment,

[0049]

[0050] wherein,

[0051]

[0052] wherein a h is a parameter for a cosine subcomponent of the h-th harmonic component of the most previous frame, wherein b h is a parameter for a sine subcomponent of the h-th harmonic component of the most previous frame, wherein, for each integer value of 1≤h≤H:

[0053]

[0054] wherein,

[0055]

[0056] wherein,

[0057]

[0058] wherein f(n) is a window function in the time domain, wherein DFT is a discrete Fourier transform, wherein,

[0059]

[0060] wherein,

[0061]

[0062] wherein f0is a fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames, wherein f s is a sampling frequency, and wherein N depends on a length of a transform block used for transforming the time-domain audio signal into the frequency domain or the spectral domain.

[0063] In an embodiment, the system of linear equations can be solved according to:

[0064]

[0065] wherein, is a first vector comprising an estimate of two harmonic parameters for each of one or more harmonic components of the most previous frame, wherein X m-1 is a second vector comprising a first group of three or more spectral coefficients of a plurality of spectral coefficients of each of one or more previous frames, wherein U + is a Moore-Penrose inverse matrix of U = [U1, U2,..., U H ], wherein U comprises a plurality of third matrices or third vectors, wherein each of the third matrices or third vectors indicates, together with an estimate of two harmonic parameters of a harmonic component of the one or more harmonic components of the most previous frame, an estimate of said harmonic component, wherein H indicates a number of harmonic components of the one or more previous frames.

[0066] In an embodiment, the encoder 100 can encode, for example, a fundamental frequency of a harmonic component, a window function, a gain factor, and a residual signal.

[0067] According to an embodiment, the encoder 100 can be configured to determine, for example, a number of one or more harmonic components of a most previous frame and a fundamental frequency of the one or more harmonic components of the most previous frame before estimating two harmonic parameters of each of the one or more harmonic components of the most previous frame using a first group of three or more spectral coefficients of a plurality of spectral coefficients of each of the one or more previous frames of the audio signal.

[0068] According to an embodiment, the encoder 100 can be configured to determine, for example, one or more groups of harmonic components from the one or more harmonic components and to apply a prediction of the audio signal on the one or more groups of harmonic components, wherein the encoder 100 can be configured to encode, for example, an order of each of the one or more groups of harmonic components of the most previous frame.

[0069] In an embodiment, the encoder 100 can be configured to apply, for example:

[0070] c h = a h cos(ω h N) + b h sin(ω h N), and

[0071] wherein the encoder 100 can be configured to apply, for example:

[0072] d h = a h sin(ω h N) + b h cos(ωh

[0073] wherein a h is a parameter for a cosine subcomponent of an h-th harmonic component of the one or more harmonic components of the most previous frame, wherein b h is a parameter for a sine subcomponent of the h-th harmonic component of the one or more harmonic components of the most previous frame, wherein c h is a parameter for a cosine subcomponent of an h-th harmonic component of the one or more harmonic components of the current frame, wherein d h is a parameter for a sine subcomponent of the h-th harmonic component of the one or more harmonic components of the current frame, wherein N depends on a length of a transform block used for transforming a time-domain audio signal into a frequency domain or spectral domain, and wherein

[0074]

[0075] wherein f0is a fundamental frequency of the one or more harmonic components of the most previous frame, which is a fundamental frequency of the one or more harmonic components of the current frame, wherein f s is a sampling frequency, and wherein h is an index indicating one of the one or more harmonic components of the most previous frame.

[0076] According to an embodiment, the encoder 100 can be configured to determine a residual signal from a plurality of spectral coefficients of the current frame in the frequency domain or transform domain and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the current frame, and wherein the encoder 100 can be configured to encode the residual signal, for example.

[0077] In an embodiment, the encoder 100 can be configured to determine a spectral prediction of one or more of the plurality of spectral coefficients of the current frame from the estimate of the two harmonic parameters for each of the one or more harmonic components of the current frame. The encoder 100 can be configured to determine a residual signal and a gain factor from the plurality of spectral coefficients of the current frame in the frequency domain or transform domain and from the spectral prediction of three or more of the plurality of spectral coefficients of the current frame; wherein the encoder 100 can be configured to generate an encoding of the current frame such that the encoding of the current frame comprises the residual signal and the gain factor, for example.

[0078] According to an embodiment, the encoder 100 can be configured to determine a residual signal of the current frame from

[0079]

[0080] wherein m is a frame index, wherein k is a frequency index, wherein R m ​(k) indicates a k-th sample of the residual signal in the spectral domain or in the transform domain, wherein X m (k) indicates a k-th sample of the spectral coefficients of the current frame in the spectral domain or in the transform domain, wherein (k) indicates a k-th sample of the spectral prediction of the current frame in the spectral domain or in the transform domain, and wherein g is a gain factor.

[0081] Figure 2 A decoder 200 for reconstructing a current frame of an audio signal according to an embodiment is shown.

[0082] One or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain.

[0083] The decoder 200 is to receive an encoding of the current frame.

[0084] Further, the decoder 200 is to determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames. The two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first group of spectral coefficients consisting of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal.

[0085] Further, the decoder 200 is to reconstruct the current frame from the encoding of the current frame and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.

[0086] The most previous frame can be, for example, the most previous with respect to the current frame.

[0087] For example, the most previous frame can be (referred to as) an immediately previous frame. For example, the immediately previous frame can directly precede the current frame.

[0088] The current frame comprises one or more harmonic components of the audio signal. Each of the one or more previous frames can comprise one or more harmonic components of the audio signal. It is assumed that the fundamental frequencies of the one or more harmonic components of the current frame and of the one or more previous frames are identical.

[0089] According to an embodiment, the two harmonic parameters for each of the one or more harmonic components of the most previous frame do not depend on a second group of spectral coefficients consisting of one or more other spectral coefficients of the plurality of spectral coefficients of the one or more previous frames.

[0090] In an embodiment, the decoder 200 can be configured to determine, for example, an estimate of the two harmonic parameters of each of the one or more harmonic components of the current frame from an estimate of the two harmonic parameters of each of the one or more harmonic components of the one or more previous frames and from a fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames.

[0091] According to an embodiment, the decoder 100 can be configured to receive, for example, an encoding of the current frame comprising a gain factor and a residual signal. The decoder 200 can be configured to reconstruct, for example, the current frame from the gain factor, from the residual signal and from a fundamental frequency of the one or more harmonic components of the current frame and the one or more previous frames. For example, it can be assumed that the fundamental frequency is constant over the current frame and the one or more previous frames.

[0092] According to an embodiment, the two harmonic parameters of each of the one or more harmonic components are a first parameter of a cosine sub-component for each of the one or more harmonic components and a second parameter of a sine sub-component for each of the one or more harmonic components.

[0093] In an embodiment, the two harmonic parameters of each of the one or more harmonic components of the one or more previous frames depend on a linear equation system comprising at least three equations, wherein each of the at least three equations depends on a spectral coefficient of a first group of spectral coefficients of a plurality of reconstructed spectral coefficients of each of the one or more previous frames.

[0094] According to an embodiment, the linear equation system can be solved using a least mean square algorithm.

[0095] According to an embodiment, the linear equation system is defined by:

[0096]

[0097] wherein,

[0098]

[0099] wherein γ1 indicates a first spectral band of one of the one or more harmonic components of the one or more previous frames having a lowest harmonic component frequency of the one or more harmonic components, wherein γ H indicates a second spectral band of one of the one or more harmonic components of the one or more previous frames having a highest harmonic component frequency of the one or more harmonic components, wherein r is an integer, r > 0.

[0100] In an embodiment, r > 1.

[0101] According to an embodiment,

[0102]

[0103] wherein,

[0104]

[0105] wherein a h is a parameter for a cosine subcomponent of the h-th harmonic component of the most previous frame, wherein b h is a parameter for a sine subcomponent of the h-th harmonic component of the most previous frame, wherein, for each integer value of 1 < h < H:

[0106]

[0107]

[0108] wherein,

[0109]

[0110] wherein,

[0111]

[0112] wherein f(n) is a window function in the time domain, wherein DFT is a discrete Fourier transform, wherein,

[0113]

[0114] wherein,

[0115]

[0116] wherein f0is a fundamental frequency of one or more harmonic components of the current frame and of the one or more previous frames, wherein f s is a sampling frequency, and wherein N depends on a length of a transform block used for transforming the time-domain audio signal into the frequency domain or the spectral domain.

[0117] In an embodiment, the system of linear equations can be solved according to:

[0118]

[0119] wherein, is a first vector comprising an estimate of two harmonic parameters for each of one or more harmonic components of the most previous frame, wherein X m-1 is a second vector comprising a first group of three or more spectral coefficients of a plurality of reconstructed spectral coefficients of each of the one or more previous frames, wherein U + is U = [U1, U2,..., U Hthe Moore-Penrose inverse of U, wherein U comprises a plurality of third matrices or third vectors, wherein each of the third matrices or third vectors indicates, together with an estimate of two harmonic parameters of a harmonic component of one or more harmonic components of the most previous frame, an estimate of said harmonic component, wherein H indicates a number of harmonic components of the one or more previous frames.

[0120] In embodiments, wherein the decoder 200 may, for example, be configured to receive a fundamental frequency of a harmonic component, a window function, a gain factor and a residual signal. The decoder 200 may, for example, be configured to reconstruct the current frame from the fundamental frequency of one or more harmonic components of the most previous frame, from an order of the harmonic component, from the window function, from the gain factor and from the residual signal.

[0121] Only the fundamental frequency, the order of the harmonic component, the window function, the gain factor and the residual need to be transmitted. The decoder 200 may, for example, compute U based on this received information and then proceed with the harmonic parameter estimation and the current frame prediction. For example, the decoder can then reconstruct the current frame by adding the transmitted residual spectrum to the predicted spectrum scaled by the transmitted gain factor.

[0122] According to embodiments, the decoder 200 may, for example, be configured to receive a number of one or more harmonic components of the most previous frame and a fundamental frequency of one or more harmonic components of the most previous frame. The decoder 200 may, for example, be configured to decode an encoding of the current frame from the number of one or more harmonic components of the most previous frame and from the fundamental frequency of one or more harmonic components of the current frame and the one or more previous frames.

[0123] According to embodiments, the decoder 200 decodes an encoding of the current frame from one or more groups of harmonic components, wherein the decoder 200 applies a prediction of the audio signal on the one or more groups of harmonic components.

[0124] According to embodiments, the decoder 200 may, for example, be configured to determine two harmonic parameters of each of one or more harmonic components of the current frame from two harmonic parameters of each of said one of one or more harmonic components of the most previous frame.

[0125] In embodiments,

[0126] c h = a h cos(ω h N) + b h sin(ω h N), and

[0127] wherein the decoder 200 may, for example, be configured to apply:

[0128] d h = a h sin(ω h N) + b h cos(ω h N),

[0129] wherein a h is a parameter for a cosine subcomponent of an h-th harmonic component of the one or more harmonic components of the most previous frame, wherein b h is a parameter for a sine subcomponent of the h-th harmonic component of the one or more harmonic components of the most previous frame, wherein c h is a parameter for a cosine subcomponent of the h-th harmonic component of the one or more harmonic components of the current frame, wherein d h is a parameter for a sine subcomponent of the h-th harmonic component of the one or more harmonic components of the current frame, wherein N depends on a length of a transform block used for transforming the time-domain audio signal into the frequency domain or the spectral domain, and wherein

[0130]

[0131] wherein f0is a fundamental frequency of the one or more harmonic components of the most previous frame, which is a fundamental frequency of the one or more harmonic components of the current frame, wherein f s is a sampling frequency, and wherein h is an index indicating one of the one or more harmonic components of the most previous frame.

[0132] According to an embodiment, the decoder 200 can be configured to receive a residual signal, wherein the residual signal depends on a plurality of spectral coefficients of the current frame in the frequency domain or the transform domain, and wherein the residual signal depends on an estimate of two harmonic parameters for each of the one or more harmonic components of the current frame.

[0133] In an embodiment, the decoder 200 can be configured to determine a spectral prediction of one or more of the plurality of spectral coefficients of the current frame from the estimate of the two harmonic parameters for each of the one or more harmonic components of the current frame, and wherein the decoder 200 can be configured to determine the current frame of the audio signal from the spectral prediction of the current frame and from the residual signal and from the gain factor.

[0134] According to an embodiment, wherein the residual signal of the current frame is defined according to:

[0135]

[0136] wherein m is a frame index, wherein k is a frequency index, wherein is the received residual after quantization reconstruction, wherein is a reconstructed current frame, wherein indicates a spectral prediction of the current frame in the spectral domain or in the transform domain, and wherein g is a gain factor.

[0137] Figure 3 A system according to an embodiment is shown.

[0138] The system comprises an encoder 100 for encoding a current frame of an audio signal according to one of the above-mentioned embodiments.

[0139] Further, the system comprises a decoder 200 for decoding an encoding of a current frame of an audio signal according to one of the above-mentioned embodiments.

[0140] Figure 7 An apparatus 700 for frame loss concealment according to an embodiment is shown.

[0141] The one or more previous frames of the audio signal precede the current frame of the audio signal. Each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in the frequency domain or in the transform domain.

[0142] The apparatus 700 will determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, wherein the two harmonic parameters of each of the one or more harmonic components of the most previous frame depend on a first group of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal.

[0143] If the apparatus 700 does not receive the current frame, or if the apparatus 700 receives the current frame in a corrupted state, the apparatus 700 will reconstruct the current frame from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.

[0144] The most previous frame can for example be the most previous with respect to the current frame.

[0145] For example, the most previous frame can be (referred to as) the immediately previous frame. For example, the immediately previous frame can directly precede the current frame.

[0146] The current frame comprises one or more harmonic components of the audio signal. Each of the one or more previous frames can comprise one or more harmonic components of the audio signal. It is assumed that the fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames is the same.

[0147] According to an embodiment, the apparatus 700 can be configured to receive a number of the one or more harmonic components of the most previous frame. The apparatus 700 can be configured to decode the encoding of the current frame, e.g. according to the number of the one or more harmonic components of the most previous frame and according to a fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames.

[0148] In an embodiment, for reconstructing the current frame, the apparatus 700 can be configured to determine an estimate of two harmonic parameters of each of the one or more harmonic components of the current frame from an estimate of two harmonic parameters of each of the one or more harmonic components of the most previous frame.

[0149] In an embodiment, the apparatus 700 will apply:

[0150] c h = a h cos(ω h N) + b h sin(ω h N) and

[0151] wherein the apparatus 700 will apply:

[0152] d h = a h sin(ω h N) + b h cos(ω h N),

[0153] wherein a h is a parameter of a cosine subcomponent of the h-th harmonic component of the one or more harmonic components of the most previous frame, wherein b h is a parameter of a sine subcomponent of the h-th harmonic component of the one or more harmonic components of the most previous frame, wherein c h is a parameter of a cosine subcomponent of the h-th harmonic component of the one or more harmonic components of the current frame, wherein d h is a parameter of a sine subcomponent of the h-th harmonic component of the one or more harmonic components of the current frame, wherein N depends on a length of a transform block used for transforming a time-domain audio signal into a frequency domain or spectral domain, and wherein

[0154]

[0155] wherein f0is a fundamental frequency of the one or more harmonic components of the most previous frame, which is a fundamental frequency of the one or more harmonic components of the current frame, wherein f s is a sampling frequency, and wherein h is an index indicating one of the one or more harmonic components of the most previous frame.

[0156] According to embodiments, the apparatus 700 can be configured to determine, for example, a spectral prediction of three or more spectral coefficients of the plurality of spectral coefficients of the current frame from an estimate of two harmonic parameters for each of the one or more harmonic components of the current frame.

[0157] In the following, preferred embodiments are provided.

[0158] First, a signal model is described.

[0159] Assume the harmonic part in a digital audio signal to be:

[0160]

[0161] where

[0162]

[0163] where f0is the fundamental frequency of the one or more harmonic components and H is the number of harmonic components. Without loss of generality, the expression for the phase component is deliberately split into two parts, where the part denoted by h (N / 2+1 / 2) facilitates mathematical derivations later on when x(n) is MDCT transformed, where N is the MDCT frame length and h is the remaining part of the phase component.

[0164] f s is, for example, the sampling frequency.

[0165] The harmonic components are determined by three parameters: frequency, amplitude and phase. Assume the frequency information h is known, then the estimation of amplitude and phase is a non-linear regression problem. However, this can be turned into a linear regression problem by rewriting equation (1) as:

[0166]

[0167] The unknown parameters of the harmonic are now a h and b h :

[0168] a h = A h cos(φ h ), (4a)

[0169] b h = -A h sin(φ h ). (4b)

[0170] Transform the block of x(n) of length 2N into the MDCT domain:

[0171]

[0172] where

[0173]

[0174] where f(n) is an analysis window function, κ k is the modulation frequency in frequency band k.

[0175] Replacing equation (3) with equation (5) and through some trigonometric based mathematical derivation, we obtain:

[0176]

[0177] where F() is a real valued function obtained by adding a phase shift term to the Fourier transform of the window function:

[0178]

[0179] In the following, the harmonic estimation and prediction are described.

[0180] Based on the assumed signal model described above through equations (3) - (8), by the additional assumption that the frequencies of the harmonic components do not change rapidly between adjacent frames, the proposed FDLMSP method can be divided into three steps. For example, to predict the m-th frame, first the frequency information of all harmonic components in the m-th frame is estimated. This frequency information will later be transmitted as part of the side information to help the prediction at the decoder 200. Then, only the parameters (denoted as a h , b h , where h = [1...H]) of each harmonic component at the m-1-th frame are estimated using only the previous frames.

[0181] Finally, the m-th frame is predicted based on the estimated harmonic parameters. The residual spectrum is then computed and further processed, e.g. quantized and transmitted. The pitch information in each frame can be obtained by a pitch estimator.

[0182] First, the harmonic estimation is described in detail.

[0183] The transform usually has a limited frequency resolution, so each harmonic component will be distributed over several adjacent bins around the frequency band where its center frequency lies. For a harmonic component with frequency ω h in the m-1-th frame, it will be located in the center of the MDCT frequency band (band index is γ h ), where

[0184]

[0185] and distributed over the following bins:

[0186] Γh = γ h - r,..., γ h + r,

[0187] where r is the number of adjacent bins on each side.

[0188] The parameters a h and b h of the harmonic component can be estimated by solving this system of linear equations formed by equation (7):

[0189]

[0190] where

[0191]

[0192] U h is a real-valued matrix independent of the signal x(n) and can be computed once, f0, N and the window function f(n) are known.

[0193] Assuming that the frequency information of all harmonic components in a frame is known, by combining equation (9) over all harmonic components the following system of linear equations is obtained:

[0194]

[0195] where

[0196]

[0197]

[0198] The matrix U and the MDCT coefficients are real-valued, so there is a system of real-valued linear equations. The estimation of the harmonic parameters p can be obtained by solving the system of linear equations with the pseudo-inverse of U in the least mean square (LMS) as follows:

[0199]

[0200] U + is for example the Moore-Penrose inverse of U.

[0201] (U + is for example the pseudo-inverse of U.)

[0202] is for example the estimation of the harmonic parameters p.

[0203] With respect to the combination of equation (9) over all harmonic components, likewise, while equation (10b) remains unchanged, equations (10a) and (10c) become:

[0204]

[0205] Since Λ≠Γ h , the dimensions of U h and γ change.

[0206] The estimate of in equation (10b) can be referred to as, for example:

[0207]

[0208] In case the number of parameters to be estimated exceeds the number of MDCT frequency bins of the harmonic span, this would result in an underdetermined system of linear equations. This is avoided by vertically stacking the matrix U and horizontally stacking the vector X with corresponding values from more previous frames. However, no additional delay is introduced, since the (most) previous frames are already in the buffer. Instead, with this extension, the proposed approach is applicable to scenarios of very low frequency resolution where the harmonic components are densely spaced. A scaling factor can be applied on the number of previous frames employed to guarantee an overdetermined system of linear equations, which also enhances the robustness of the prediction concept to noise in the signal.

[0209] Now, the prediction is described in detail.

[0210] Assuming that the frequency and amplitude of the sinusoidal wave are constant, the m-th frame in the time domain can be written as:

[0211]

[0212] where

[0213] c h = a h cos(ω h N) + b h sin(ω h N), (15a)

[0214] d h = a h sin(ω h N) + b h cos(ω h N). (15b)

[0215] With the estimate of the harmonic parameters of each of one or more harmonic components in the m-1-th frame at hand, the prediction of the current MDCT frame is based on equations (5) - (9):

[0216]

[0217] where

[0218]

[0219] For frequency bands that have not been predicted, the predicted value is set to zero.

[0220] However, due to the non-stationarity of the signal, the amplitude of the harmonics can vary slightly between consecutive frames. A gain factor is introduced to accommodate this amplitude change and is sent to the decoder 200 as part of the auxiliary information.

[0221] The residual spectrum is then:

[0222]

[0223] The above concepts will be evaluated below.

[0224] To evaluate the performance of the proposed FDLMSP concept, it has been based on Figure 4 The encoder environment was built in Python. The concepts provided are implemented following the description above, where r equals 2. For comparison, TDLTP and FDP have been reimplemented based on references [2], [5]. This was intended to evaluate the three prediction concepts from three different perspectives using experiments: (i) performance with respect to different frequency resolutions of MDCT coefficients, (ii) sensitivity to incoherence of the test material [7], and (iii) overall performance and ability to compare with each other in the same coding scenario. Incoherence of tones usually means that its higher-order harmonics are no longer evenly spaced. Since harmonics in higher frequency bands are less perceptually important [8], the effect of using different prediction bandwidths has been evaluated.

[0225] For the experiments, a sampling frequency of 16 kHz and MDCT frame lengths of 64, 128, 256, and 512 were used. Predictions were performed on finite bandwidths of 1 kHz, 2 kHz, 4 kHz, and 8 kHz. A sinusoidal window was chosen as the analysis window because it satisfies the constraints of perfect reconstruction [9]. The method can also handle asymmetric windows when switching between different frame lengths. To improve the accuracy of harmonic estimation, the F(ω) function is computed on the interpolation transfer function of the analysis window. In TDLTP, for each frame, a 3-tap prediction filter is computed based on the concept of autocorrelation using the fully reconstructed data and the original time-domain signal. The fact that the pitch lag may not be an integer multiple of the sampling interval is also taken into account when searching for the pitch lag of the previously fully reconstructed data from the buffered data. The number of time or spectral adjacent frequency bands is limited to 2 in FDP.

[0226] The YIN algorithm

[10] was used for pitch estimation. The fo search range was set to [20,..., 1000] Hz and the harmonic threshold was 0.25. The perceptual model based on a complex infinite impulse response (IIR) filterbank presented in

[11] was used to compute the quantized masking thresholds. The finer pitch search around the YIN estimate (± 0.5 Hz with a step of 0.02 Hz) and the optimal gain factor search in [0.5,..., 2] with a step of 0.01 were jointly done in each frame by minimizing the perceptual entropy (PE) of the quantized residual spectrum which is an approximation of the entropy of the quantized residual spectrum taking into account the perceptual model.

[0227] The encoder has four operating modes: "FDLMSP", "TDLTP", "FDP" and "Adaptive MDCT LTP (AMLTP)" respectively. In "AMLTP" mode, the encoder switches between different prediction concepts on a frame basis, optimizing the PE minimization criterion. For all four operating modes, no prediction is done in a frame if the PE of the residual spectrum is higher than the one of the original signal spectrum.

[0228] For each mode, the encoder was tested on six different materials: three single tone notes with a duration of 1 to 2 seconds: a bass note (f0 around 50 Hz), a piano note (f0 around 88 Hz) and an oboe note (f0 around 290 Hz). These test materials have a relatively regular harmonic structure and a slowly varying temporal envelope. The encoder was also tested on more complex test materials: a trumpet (around 5 seconds long, f0 varying between 300 and 700 Hz), a female voice (around 10 seconds long, f0 varying between 200 and 300 Hz) and a male speech (around 8 seconds long; f0 varying between 100 and 220 Hz). These three test materials have widely varying envelopes and a rapidly changing pitch along time as well as a less regular harmonic structure. During the experiments, it has been noticed that the bass note has a second harmonic much stronger than the first one, leading to a constantly wrong pitch estimation. Therefore, the f0 search range in the YIN pitch estimator has been adjusted for this bass note for a correct pitch estimation.

[0229] The average PE of the quantized residual spectrum and the quantized original signal spectrum has been estimated. Based on the estimated PE, the bit rate saved in the transmitted signal by applying the prediction (BS) [in bits / s] has been computed without taking into account the bit rate consumption of the side information. First, the behavior of each concept has been checked and the comparison has been limited to single tone note prediction for a reasonable extrapolation and analysis. We then compare the performances of the four modes under the same parameter configuration.

[0230] Figure 5It is shown that the bit rate saved on mono tonal prediction with different prediction bandwidths and MDCT lengths using three prediction concepts.

[0231] First, the FDP prediction concept from the related art is described in the following. The FDP prediction concept is described in more detail in [5] and

[13] (WO 2016142357 Al, published September 2016).

[0232] Figure 8 A schematic block diagram of an encoder 101 for encoding an audio signal 102 according to an example of the FDP prediction concept is shown. The encoder 101 is configured to encode the audio signal 102 in a transform domain or filter bank domain 104 (e.g. frequency domain or spectral domain), wherein the encoder 101 is configured to determine spectral coefficients 106_t0_f1 to 106_t0_f6 of the audio signal 102 for a current frame 108_t0 and spectral coefficients 106_t-1_f1 to 106_t-1_f6 of the audio signal for at least one previous frame 108_t-1. Furthermore, the encoder 101 is configured to selectively apply predictive encoding to a plurality of individual spectral coefficients 106_t0_f2 or a group of spectral coefficients 106_t0_f4 and 106_t0_f5, wherein the encoder 101 is configured to determine a pitch value, wherein the encoder 101 is configured to select the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 to which the predictive encoding is applied based on the pitch value.

[0233] In other words, the encoder 101 is configured to selectively apply predictive encoding to a plurality of individual spectral coefficients 106_t0_f2 or a group of spectral coefficients 106_t0_f4 and 106_t0_f5 selected based on a single pitch value transmitted as side information.

[0234] The pitch value can correspond to a frequency (e.g. a fundamental frequency of a harmonic tone of the audio signal 102) which defines the center of all spectral coefficient groups to which the prediction is applied together with its integer multiples: a first group can be centered at this frequency, a second group can be centered at this frequency multiplied by 2, a third group can be centered at this frequency multiplied by 3, and so on. The knowledge of these center frequencies enables the computation of prediction coefficients which are used to predict the corresponding sinusoidal signal components (e.g. the fundamental frequency and the overtones of a harmonic signal). Therefore, the complex and error-prone backward adaptation of the prediction coefficients is no longer necessary.

[0235] In an example, the encoder 101 can be configured to determine one pitch value per frame.

[0236] In an example, the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 can be separated by at least one spectral coefficient 106_t0_f3.

[0237] In an example, the encoder 101 can be configured to apply the predictive coding to a plurality of individual spectral coefficients separated by at least one spectral coefficient, for example to two individual spectral coefficients separated by at least one spectral coefficient. Further, the encoder 101 can be configured to apply the predictive coding to a plurality of groups of spectral coefficients (each group comprising at least two spectral coefficients) separated by at least one spectral coefficient, for example to two groups of spectral coefficients separated by at least one spectral coefficient. Further, the encoder 101 can be configured to apply the predictive coding to a plurality of individual spectral coefficients and / or groups of spectral coefficients separated by at least one spectral coefficient, for example to at least one individual spectral coefficient and at least one group of spectral coefficients separated by at least one spectral coefficient.

[0238] In Figure 8 In the illustrated example, the encoder 101 is configured to determine the six spectral coefficients 106_t0_f1 to 106_t0_f6 of the current frame 108_t0 and the six spectral coefficients 106_t-1_f1 to 106_t-1_f6 of the (most) previous frame 108_t-1. Accordingly, the encoder 101 is configured to selectively apply the predictive coding to the individual second spectral coefficient 106_t0_f2 of the current frame and to the group of spectral coefficients consisting of the fourth spectral coefficient 106_t0_f4 and the fifth spectral coefficient 106_t0_f5 of the current frame 108_t0. It can be seen that the individual second spectral coefficient 106_t0_f2 and the group of spectral coefficients consisting of the fourth spectral coefficient 106_t0_f4 and the fifth spectral coefficient 106_t0_f5 are separated from each other by the third spectral coefficient 106_t0_f3.

[0239] It is noted that the term "selectively" used herein refers to (only) applying the predictive coding to the selected spectral coefficients. In other words, the predictive coding is not necessarily applied to all spectral coefficients, but only to the selected individual spectral coefficients or groups of spectral coefficients, which can be separated from each other by at least one spectral coefficient. In other words, the predictive coding can be disabled for at least one spectral coefficient by which the selected plurality of individual spectral coefficients or groups of spectral coefficients are separated.

[0240] In an example, the encoder 101 can be configured to selectively apply predictive encoding to the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 based on at least the corresponding plurality of individual spectral coefficients 106_t-1_f2 or the group of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1.

[0241] For example, the encoder 101 can be configured to predictively encode the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 by encoding a prediction error between the plurality of predicted individual spectral coefficients 110_t0_f2 or the group of predicted spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 and the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame (or a quantized version thereof).

[0242] In an example, the encoder 101 can be configured to predictively encode the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 by encoding a prediction error between the plurality of predicted individual spectral coefficients 110_t0_f2 or the group of predicted spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame 108_t0 and the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 of the current frame (or a quantized version thereof). Figure 8 In other words, the second spectral coefficient 106_t0_f2 is encoded by encoding a prediction error (or difference) between the predicted second spectral coefficient 110_t0_f2 and the (actual or determined) second spectral coefficient 106_t0_f2, wherein the fourth spectral coefficient 106_t0_f4 is encoded by encoding a prediction error (or difference) between the predicted fourth spectral coefficient 110_t0_f4 and the (actual or determined) fourth spectral coefficient 106_t0_f4, wherein the fifth spectral coefficient 106_t0_f5 is encoded by encoding a prediction error (or difference) between the predicted fifth spectral coefficient 110_t0_f5 and the (actual or determined) fifth spectral coefficient 106_t0_f5.

[0243]

[0244] ​In an example, the encoder 101 can be configured to determine the plurality of predicted individual spectral coefficients 110_t0_f2 or the group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame 108_t0 by the corresponding actual versions of the plurality of individual spectral coefficients 106_t-1_f2 or the group of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1.

[0245] In other words, in the above determination process, the encoder 101 can directly use the plurality of actual individual spectral coefficients 106_t-1_f2 or the group of actual spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1, wherein 106_t-1_f2, 106_t-1_f4 and 106_t-1_f5 represent the original, not yet quantized spectral coefficients or spectral coefficient groups as they are obtained by the encoder 101, so that the encoder can operate in the transform domain or filter bank domain 104.

[0246] For example, the encoder 101 can be configured to determine the predicted second spectral coefficient 110_t0_f2 of the current frame 108_t0 based on the corresponding not yet quantized version of the second spectral coefficient 106_t-1_f2 of the previous frame 108_t-1, to determine the predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on the corresponding not yet quantized version of the fourth spectral coefficient 106_t-1_f4 of the previous frame 108_t-1, and to determine the predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on the corresponding not yet quantized version of the fifth spectral coefficient 106_t-1_f5 of the previous frame.

[0247] By this approach, the prediction encoding scheme and the prediction decoding scheme can exhibit a kind of harmonic shaping of the quantization noise, since the corresponding decoder can only employ the transmitted quantized versions of the plurality of individual spectral coefficients 106_t-1_f2 or the group of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1 for the prediction decoding in the above determination step.

[0248] While such harmonic noise shaping (e.g., which is traditionally performed by long-term prediction (LTP) in the time domain) can be subjectively advantageous for predictive coding, in some cases it can be undesirable as it can lead to the introduction of unwanted, excessive tonality into the decoded audio signal. For this reason, an alternative predictive coding scheme is described below that is fully synchronized with the corresponding decoding and thus only utilizes any possible prediction gain but does not lead to quantization noise shaping. According to this alternative coding example, the encoder 101 can be configured to determine a plurality of predicted individual spectral coefficients 110_t0_f2 or a group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame 108_t0 using corresponding quantized versions of a plurality of individual spectral coefficients 106_t-1_f2 or a group of spectral coefficients 106_t-1_f4 and 106_t-1_f5 of the previous frame 108_t-1.

[0249] For example, the encoder 101 can be configured to determine a predicted second spectral coefficient 110_t0_f2 of the current frame 108_t0 based on a corresponding quantized version of a second spectral coefficient 106_t-1_f2 of the previous frame 108_t-1, to determine a predicted fourth spectral coefficient 110_t0_f4 of the current frame 108_t0 based on a corresponding quantized version of a fourth spectral coefficient 106_t-1_f4 of the previous frame 108_t-1, and to determine a predicted fifth spectral coefficient 110_t0_f5 of the current frame 108_t0 based on a corresponding quantized version of a fifth spectral coefficient 106_t-1_f5 of the previous frame.

[0250] Furthermore, the encoder 101 can be configured to derive prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5 and 114_f5 from the interval values and to calculate a plurality of predicted individual spectral coefficients 110_t0_f2 or a group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame 108_t0 using corresponding quantized versions of a plurality of individual spectral coefficients 106_t-1_f2 and 106_t-2_f2 or a group of spectral coefficients 106_t-1_f4, 106_t-2_f4, 106_t-1_f5 and 106_t-2_f5 of at least two previous frames 108_t-1 and 108_t-2 and using the derived prediction coefficients 112_f2, 114_f2, 112_f4, 114_f4, 112_f5 and 114_f5.

[0251] For example, encoder 101 can be configured to: derive prediction coefficients 112_f2 and 114_f2 of the second spectral coefficient 106_t0_f2 from the spacing value, derive prediction coefficients 112_f4 and 114_f4 of the fourth spectral coefficient 106_t0_f4 from the spacing value, and derive prediction coefficients 112_f5 and 114_f5 of the fifth spectral coefficient 106_t0_f5 from the spacing value.

[0252] For example, the prediction coefficients can be derived as follows: if the spacing value corresponds to frequency f0 or its encoded version, then the center frequency of the Kth spectral coefficient group enabling prediction is fc = K * f0. If the sampling frequency is fs and the transform jump size (offset between consecutive frames) is N, then assuming the sinusoidal signal has a frequency fc, the ideal predictor coefficients in the Kth group are:

[0253] p1=2*cos(N*2*pi*fc / fs)and p2=-1.

[0254] If, for example, the spectral coefficients 106_t0_f4 and 106_t0_f5 are both within this group, then the prediction coefficients are:

[0255] 112_f4=112_f5=2*cos(N*2*pi*fc / fs)and 114_f4=114_f5=-1.

[0256] For stability reasons, a damping factor d can be introduced to modify the prediction coefficients:

[0257] 112_f4'=112_f5'=d*2*cos(N*2*pi*fc / fs), 114_f4'=114_f5'=d 2 .

[0258] Since the spacing values ​​are transmitted in the encoded audio signal 120, the decoder can derive the exact same prediction coefficients 212_f4=212_f5=2*cos(N*2*pi*fc / fs) and 114_f4=114_f5=-1. If a damping factor is used, the coefficients can be corrected accordingly.

[0259] like Figure 8 As shown, encoder 101 can be configured to provide an encoded audio signal 120. Therefore, encoder 101 can be configured to include in the encoded audio signal 120 a quantized version of the prediction error of a group of individual spectral coefficients 106_t0_f2 or 106_t0_f4 and 106_t0_f5 to which prediction coding has been applied. Furthermore, encoder 101 can be configured not to include prediction coefficients 112_f2 to 114_f5 in the encoded audio signal 120.

[0260] Thus, the encoder 101 can use only the prediction coefficients 112_f2 to 114_f5 to calculate the plurality of predicted individual spectral coefficients 110_t0_f2 or the group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 and, therefrom, the prediction error between the predicted individual spectral coefficients 110_t0_f2 or the group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame and the individual spectral coefficients 106_t0_f2 or the group of predicted spectral coefficients 110_t0_f4 and 110_t0_f5 of the current frame, but does not provide the individual spectral coefficient 106_t0_f4 (or a quantized version thereof) or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 (or a quantized version thereof) in the encoded audio signal 120 nor the prediction coefficients 112_f2 to 114_f5. Thus, the decoder can derive the prediction coefficients 112_f2 to 114_f5 from the interval values, which are used to calculate the plurality of predicted individual spectral coefficients or the group of predicted spectral coefficients of the current frame.

[0261] In other words, the encoder 101 can be configured to provide the encoded audio signal 120 comprising a quantized version of the prediction error instead of a quantized version of the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 for the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 to which the prediction encoding is applied.

[0262] Further, the encoder 101 can be configured to provide the encoded audio signal 102 comprising a quantized version of the spectral coefficient 106_t0_f3 by which the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5 is separated such that, for the spectral coefficient 106_t0_f2 or the group of spectral coefficients 106_t0_f4 and 106_t0_f5, there is an alternation between the quantized version of the prediction error comprised in the encoded audio signal 120 and the spectral coefficient 106_t0_f3 or the group of spectral coefficients provided without using prediction encoding.

[0263] In an example, the encoder 101 can be further configured to entropy encode the quantized version of the prediction error and the quantized version of the spectral coefficients 106_t0_f3 from which the plurality of individual spectral coefficients 106_t0_f2 or groups of spectral coefficients 106_t0_f4 and 106_t0_f5 are separated and the encoder 101 can be further configured to include the entropy encoded version (instead of the non-entropy encoded version thereof) in the encoded audio signal 120.

[0264] In an example, the encoder 101 can be configured to select the groups 116_1 to 116_6 of spectral coefficients arranged in the spectrum from a harmonic grid defined by the pitch value used for the predictive encoding. Thus, the harmonic grid defined by the pitch value describes the periodic spectral distribution (equidistantly spaced) of the harmonics in the audio signal 102. In other words, the harmonic grid defined by the pitch value can be a sequence of equidistantly spaced pitch values describing the harmonics of the audio signal.

[0265] Further, the encoder 101 can be configured to select spectral coefficients (e.g., only those spectral coefficients) whose spectral indices are equal to or lie within a range (e.g., predetermined or variable) around the plurality of spectral indices derived based on the pitch value for the predictive encoding.

[0266] From the pitch value, indices (or numbers) of spectral coefficients can be derived which represent harmonics of the audio signal 102. For example, assuming that the fourth spectral coefficient 106_t0_f4 represents the instantaneous fundamental frequency of the audio signal 102 and assuming that the pitch value is 5, then a spectral coefficient with index 9 can be derived based on the pitch value. The spectral coefficient with index 9 thus derived, i.e., the ninth spectral coefficient 106_t0_f9 represents the second harmonic. Similarly, spectral coefficients with indices 14, 19, 24, and 29 can be derived which represent the third harmonic 124_3 to the sixth harmonic 124_6. However, not only spectral coefficients with indices equal to the plurality of spectral indices derived based on the pitch value can be predictively encoded, but also spectral coefficients with indices within a given range around the plurality of spectral indices derived based on the pitch value can be predictively encoded.

[0267] Furthermore, the encoder 101 can be configured to select the groups of spectral coefficients 116_1 to 116_6 (or the individual spectral coefficients) to which the predictive encoding is applied such that there is a periodic alternation (with a periodicity of + / - 1 spectral coefficient tolerance) between the groups of spectral coefficients 116_1 to 116_6 to which the predictive encoding is applied and the spectral coefficients that are separated from the groups of spectral coefficients (or the individual spectral coefficients) to which the predictive encoding is applied. The + / - 1 spectral coefficient tolerance can be required when the distance between two harmonics of the audio signal 102 is not equal to an integer interval value (with respect to the index or number of spectral coefficients) but a fraction or multiple thereof.

[0268] In other words, the audio signal 102 can comprise at least two harmonic signal components 124_1 to 124_6, wherein the encoder 101 can be configured to selectively apply the predictive encoding to a plurality of groups 116_1 to 116_6 of spectral coefficients (or individual spectral coefficients) that represent at least two of the harmonic signal components 124_1 to 124_6 or a spectral environment around the at least two of the harmonic signal components 124_1 to 124_6. The spectral environment around the at least two of the harmonic signal components 124_1 to 124_6 can be, for example, + / - 1, 2, 3, 4, or 5 spectral components.

[0269] Thus, the encoder 101 can be configured to not apply the predictive encoding to groups 118_1 to 118_5 of spectral coefficients (or individual spectral coefficients) that do not represent the at least two of the harmonic signal components 124_1 to 124_6 of the audio signal 102 or a spectral environment of the at least two of the harmonic signal components 124_1 to 124_6 of the audio signal 102. In other words, the encoder 101 can be configured to not apply the predictive encoding to the plurality of groups 118_1 to 118_5 of spectral coefficients (or individual spectral coefficients) that belong to a non-tonal background noise between the signal harmonics 124_1 to 124_6.

[0270] Furthermore, the encoder 101 can be configured to determine a harmonic interval value that is indicative of a spectral separation between the at least two of the harmonic signal components 124_1 to 124_6 of the audio signal 102, the harmonic interval value being indicative of the plurality of individual spectral coefficients or the groups of spectral coefficients that represent the at least two of the harmonic signal components 124_1 to 124_6 of the audio signal 102.

[0271] Furthermore, the encoder 101 can be configured to provide the encoded audio signal 120 such that the encoded audio signal 120 comprises the interval value (e.g., one interval value per frame) or a parameter from which the interval value can be directly derived (alternatively).

[0272] The example addresses the above two problems of the FDP method by introducing a harmonic spacing value into the FDP process, which is signaled from the encoder (sender) 101 to the respective decoder (receiver) so that both can operate in a fully synchronized manner. The harmonic spacing value can be used as an indicator of the instantaneous fundamental frequency (or pitch) of one or more spectra associated with the frame to be encoded and identifies which spectral segments (spectral coefficients) should be predicted. More specifically, only those spectral coefficients around the harmonic signal components that are located at (in terms of their index) integer multiples of the fundamental pitch (defined by the harmonic spacing value) should be predicted.

[0273] Figure 9 A schematic block diagram of a decoder 201 for decoding an encoded signal 120 of the FDP prediction concept according to an example is shown. The decoder 201 is configured to decode the encoded audio signal 120 in a transform domain or filter bank domain 204, wherein the decoder 201 is configured to parse the encoded audio signal 120 to obtain encoded spectral coefficients 206_t0_f1 to 206_t0_f6 of the audio signal for a current frame 208_t0 and to obtain encoded spectral coefficients 206_t-1_f0 to 206_t-1_f6 for at least one previous frame 208_t-1, and wherein the decoder 201 is configured to selectively apply a prediction decoding to a plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients separated by at least one encoded spectral coefficient.

[0274] In an example, the decoder 201 can be configured to apply the prediction decoding to a plurality of individual encoded spectral coefficients separated by at least one encoded spectral coefficient, e.g. to two individual encoded spectral coefficients separated by at least one encoded spectral coefficient. Further, the decoder 201 can be configured to apply the prediction decoding to a plurality of groups of encoded spectral coefficients separated by at least one encoded spectral coefficient (each group comprising at least two encoded spectral coefficients), e.g. to two groups of encoded spectral coefficients separated by at least one encoded spectral coefficient. Further, the decoder 201 can be configured to apply the prediction decoding to a plurality of individual encoded spectral coefficients and / or groups of encoded spectral coefficients separated by at least one encoded spectral coefficient, e.g. to at least one individual encoded spectral coefficient and at least one group of encoded spectral coefficients separated by at least one encoded spectral coefficient.

[0275] In Figure 9In the illustrated example, the decoder 201 is configured to determine six encoded spectral coefficients 206_t0_f1 to 206_t0_f6 of the current frame 208_t0 and six encoded spectral coefficients 206_t-1_f1 to 206_t-1_f6 of the previous frame 208_t-1. Thus, the decoder 201 is configured to selectively apply the predictive decoding to the individual second encoded spectral coefficient 206_t0_f2 of the current frame and to the group of encoded spectral coefficients consisting of the fourth encoded spectral coefficient 206_t0_f4 and the fifth encoded spectral coefficient 206_t0_f5 of the current frame 208_t0. As can be seen, the individual second encoded spectral coefficient 206_t0_f2 and the group of encoded spectral coefficients consisting of the fourth encoded spectral coefficient 206_t0_f4 and the fifth encoded spectral coefficient 206_t0_f5 are separated from each other by the third encoded spectral coefficient 206_t0_f3.

[0276] It is noted that the term "selectively" as used herein refers to applying the predictive decoding to (only) selected encoded spectral coefficients. In other words, the predictive decoding is not applied to all encoded spectral coefficients, but only to selected individual encoded spectral coefficients or groups of encoded spectral coefficients, the selected individual encoded spectral coefficients and / or groups of encoded spectral coefficients being separated from each other by at least one encoded spectral coefficient. In other words, the predictive decoding is not applied to the at least one encoded spectral coefficient by which the selected plurality of individual encoded spectral coefficients or groups of encoded spectral coefficients are separated.

[0277] In the example, the decoder 201 can be configured to not apply the predictive decoding to the at least one encoded spectral coefficient 206_t0_f3 by which the individual encoded spectral coefficient 206_t0_f2 or the group of spectral coefficients 206_t0_f4 and 206_t0_f5 are separated.

[0278] The decoder 201 can be configured to entropy-decode the encoded spectral coefficients to obtain quantized prediction errors for the spectral coefficients 206_t0_f2, 2016_t0_f4 and 206_t0_f5 to which the predictive decoding is to be applied and to obtain a quantized spectral coefficient 206_t0_f3 for the at least one spectral coefficient to which the predictive decoding is not to be applied. Thus, the decoder 201 can be configured to apply the quantized prediction errors to the plurality of predicted individual spectral coefficients 210_t0_f2 or the group of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 to obtain, for the current frame 208_t0, decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f2, 206_t0_f4 and 206_t0_f5 to which the predictive decoding has been applied.

[0279] For example, the decoder 201 can be configured to obtain a second quantized prediction error for the second quantized spectral coefficient 206_t0_f2 and apply the second quantized prediction error to the predicted second spectral coefficient 210_t0_f2 to obtain a second decoded spectral coefficient associated with the second encoded spectral coefficient 206_t0_f2; wherein the decoder 201 can be configured to obtain a fourth quantized prediction error for the fourth quantized spectral coefficient 206_t0_f4 and apply the fourth quantized prediction error to the predicted fourth spectral coefficient 210_t0_f4 to obtain a fourth decoded spectral coefficient associated with the fourth encoded spectral coefficient 206_t0_f4; and wherein the decoder 201 can be configured to obtain a fifth quantized prediction error for the fifth quantized spectral coefficient 206_t0_f5 and apply the fifth quantized prediction error to the predicted fifth spectral coefficient 210_t0_f5 to obtain a fifth decoded spectral coefficient associated with the fifth encoded spectral coefficient 206_t0_f5.

[0280] Further, the decoder 201 can be configured to determine the plurality of predicted individual spectral coefficients 210_t0_f2 or the group of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 for the current frame 208_t0 based on the corresponding plurality of individual encoded spectral coefficients 206_t-1_f2 of the previous frame 208_t-1 (e.g., using the plurality of previously decoded spectral coefficients associated with the plurality of individual encoded spectral coefficients 206_t-1_f2) or the group of encoded spectral coefficients 206_t-1_f4 and 206_t-1_f5 (e.g., using the group of previously decoded spectral coefficients associated with the group of encoded spectral coefficients 206_t-1_f4 and 206_t-1_f5).

[0281] For example, the decoder 201 can be configured to determine the second predicted spectral coefficient 210_t0_f2 for the current frame 208_t0 using a previously decoded (quantized) second spectral coefficient associated with the second encoded spectral coefficient 206_t-1_f2 of the previous frame 208_t-1, to determine the fourth predicted spectral coefficient 210_t0_f4 for the current frame 208_t0 using a previously decoded (quantized) fourth spectral coefficient associated with the fourth encoded spectral coefficient 206_t-1_f4 of the previous frame 208_t-1, and to determine the fifth predicted spectral coefficient 210_t0_f5 for the current frame 208_t0 using a previously decoded (quantized) fifth spectral coefficient associated with the fifth encoded spectral coefficient 206_t-1_f5 of the previous frame 208_t-1.

[0282] Further, the decoder 201 can be configured to derive prediction coefficients from the interval values, and wherein the decoder 201 can be configured to calculate the plurality of predicted individual spectral coefficients 210_t0_f2 or the group of predicted spectral coefficients 210_t0_f4 and 210_t0_f5 of the current frame 208_t0 using the corresponding plurality of previously decoded individual spectral coefficients or the group of previously decoded spectral coefficients of the at least two previous frames 208_t-1 and 208_t-2 and using the derived prediction coefficients.

[0283] For example, the decoder 201 can be configured to derive prediction coefficients 212_f2 and 214_f2 for the second encoded spectral coefficient 206_t0_f2 from the interval values, to derive prediction coefficients 212_f4 and 214_f4 for the fourth encoded spectral coefficient 206_t0_f4 from the interval values, and to derive prediction coefficients 212_f5 and 214_f5 for the fifth encoded spectral coefficient 206_t0_f5 from the interval values.

[0284] It is noted that the decoder 201 can be configured to decode the encoded audio signal 120 to obtain quantized prediction errors instead of obtaining a plurality of individual quantized spectral coefficients or a group of quantized spectral coefficients of a plurality of individual encoded spectral coefficients or a group of encoded spectral coefficients to which the prediction decoding is applied.

[0285] Further, the decoder 201 can be configured to decode the encoded audio signal 120 to obtain quantized spectral coefficients from which the plurality of individual spectral coefficients or the group of spectral coefficients is separated such that there is an alternation of the encoded spectral coefficient 206_t0_f5 or the group of encoded spectral coefficients 206_t0_f4 and 206_t0_f5 and the encoded spectral coefficient 206_t0_f3 or the group of encoded spectral coefficients, to obtain quantized prediction errors for the encoded spectral coefficient 206_t0_f5 or the group of encoded spectral coefficients 206_t0_f4 and 206_t0_f5, and to obtain quantized spectral coefficients for the encoded spectral coefficient 206_t0_f3 or the group of encoded spectral coefficients.

[0286] The decoder 201 can be configured to provide the decoded audio signal 220 using the decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f2, 206_t0_f4 and 206_t0_f5 to which the prediction decoding is applied and using the entropy-decoded spectral coefficients associated with the encoded spectral coefficients 206_t0_f1, 206_t0_f3 and 206_t0_f6 to which the prediction decoding is not applied.

[0287] In an example, the decoder 201 can be configured to obtain the pitch value, wherein the decoder 201 can be configured to select the group of the plurality of individual encoded spectral coefficients 206_t0_f2 or 206_t0_f4 and 206_t0_f5 to which the predictive decoding is applied based on the pitch value.

[0288] As already mentioned above with respect to the corresponding encoder 101, the pitch value can be, for example, the pitch (or distance) between two characteristic frequencies of the audio signal. Further, the pitch value can be the integer number of spectral coefficients (or indices of spectral coefficients) which approximates the pitch between two characteristic frequencies of the audio signal. Naturally, the pitch value can also be a fraction or multiple of the integer number of spectral coefficients which describes the pitch between two characteristic frequencies of the audio signal.

[0289] The decoder 201 can be configured to select the individual spectral coefficients or the group of spectral coefficients which are arranged in the spectrum according to the harmonic grid defined by the pitch value used for the predictive decoding. The harmonic grid defined by the pitch value can describe the periodic spectral distribution (equidistant spacing) of harmonics in the audio signal 102. In other words, the harmonic grid defined by the pitch value can be a sequence of pitch values which describes the equidistant spacing of harmonics of the audio signal 102.

[0290] Further, the decoder 201 can be configured to select spectral coefficients (e.g. only those spectral coefficients) whose spectral indices are equal to or lie within a range (e.g. predetermined or variable) around a plurality of spectral indices derived based on the pitch value for the predictive decoding. Accordingly, the decoder 201 can be configured to set the width of the range depending on the pitch value.

[0291] In an example, the encoded audio signal can comprise the pitch value or an encoded version thereof (e.g. a parameter from which the pitch value can be derived directly), wherein the decoder 201 can be configured to extract the pitch value or the encoded version thereof from the encoded audio signal to obtain the pitch value.

[0292] Alternatively, the decoder 201 can be configured to determine the pitch value by itself, i.e. the encoded audio signal does not comprise the pitch value. In that case, the decoder 201 can be configured to determine the instantaneous fundamental frequency (of the encoded audio signal 120 representing the audio signal 102) and to derive the pitch value from the instantaneous fundamental frequency or a fraction or multiple thereof.

[0293] In an example, the decoder 201 can be configured to select the plurality of individual spectral coefficients or the group of spectral coefficients to which the predictive decoding is applied such that there is a periodic alternation (with a periodicity of + / - 1 spectral coefficient with a tolerance) between the plurality of individual spectral coefficients or the group of spectral coefficients to which the predictive decoding is applied and spectral coefficients which are separated from the plurality of individual spectral coefficients or the group of spectral coefficients to which the predictive decoding is applied.

[0294] In an example, the audio signal 102 represented by the encoded audio signal 120 comprises at least two harmonic signal components, wherein the decoder 201 is configured to selectively apply the predictive decoding to those groups of the plurality of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 representing at least the two harmonic signal components of the audio signal 102 or the spectral environment around the at least two harmonic signal components of the audio signal 102. The spectral environment around the at least two harmonic signal components can be, for example, + / - 1, 2, 3, 4 or 5 spectral components.

[0295] Thus, the decoder 201 can be configured to identify the at least two harmonic signal components and to selectively apply the predictive decoding to those groups of the plurality of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 associated with the identified harmonic signal components (e.g., representing the identified harmonic signal components or around the identified harmonic signal components).

[0296] Alternatively, the encoded audio signal 120 can comprise information (e.g., a pitch value) identifying the at least two harmonic signal components. In that case, the decoder 201 can be configured to selectively apply the predictive decoding to those groups of the plurality of individual encoded spectral coefficients 206_t0_f2 or encoded spectral coefficients 206_t0_f4 and 206_t0_f5 associated with the identified harmonic signal components (e.g., representing the identified harmonic signal components or around the identified harmonic signal components).

[0297] In both alternatives described above, the decoder 201 can be configured to not apply the predictive decoding to those groups of the plurality of individual encoded spectral coefficients 206_t0_f3, 206_t0_f1 and 206_t0_f6 or encoded spectral coefficients 206_t0_f3, 206_t0_f1 and 206_t0_f6 not representing the at least two harmonic signal components of the audio signal 102 or the spectral environment of the at least two harmonic signal components.

[0298] In other words, the decoder 201 can be configured to not apply the predictive decoding 102 to those groups of the plurality of individual encoded spectral coefficients 206_t0_f3, 206_t0_f1, 206_t0_f6 or encoded spectral coefficients 206_t0_f3, 206_t0_f1, 206_t0_f6 belonging to a non-tonal background noise between signal harmonics of the audio signal.

[0299] The idea of specific embodiments is now to provide an encoder and a decoder having different modes of operation.

[0300] According to an embodiment, the encoder 100 is, for example, capable of operating in a first mode and is, for example, capable of operating in at least one of a second mode and a third mode and a fourth mode.

[0301] If the encoder 100 is in the first mode, the encoder 100 can for example be configured to encode the current frame by determining an estimate of two harmonic parameters for each of one or more harmonic components of the most previous frame using a first group of three or more spectral coefficients of a plurality of spectral coefficients of each of one or more previous frames of the audio signal.

[0302] If the encoder 100 is in the second mode, the encoder 100 can for example be configured to encode the audio signal in a transform domain or a filter bank domain, and the encoder can for example be configured to determine a plurality of spectral coefficients 106_t0_f1:106_t0_f6; 106_t-1_f1:106_t-1_f6 of the audio signal 102 for the current frame 108_t0 and at least for the previous frame 108_t-1, wherein the encoder 100 can for example be configured to selectively apply predictive encoding to the plurality of individual spectral coefficients 106_t0_f2 or to the group of spectral coefficients 106_t0_f4, 106_t0_f5, the encoder 100 can for example be configured to determine a pitch value, the encoder 100 can for example be configured to select the plurality of individual spectral coefficients 106_t0_f2 or the group of spectral coefficients 106_t0_f4, 106_t0_f5 to which to apply the predictive encoding based on the pitch value.

[0303] In embodiments, in each of the first mode and the second mode and the third mode and the fourth mode, the encoder 100 can for example be configured to refine the fundamental frequency on a frame basis to obtain a refined fundamental frequency and to adapt the gain factors to obtain adapted gain factors according to a minimization criterion. Further, the encoder 100 can for example be configured to encode the refined fundamental frequency and the adapted gain factors instead of the original fundamental frequency and the gain factors.

[0304] In embodiments, the encoder 100 can for example be configured to set itself to the first mode or to at least one of the second mode and the third mode and the fourth mode depending on a current frame of the audio signal. The encoder 100 can for example be configured to encode regardless of whether the current frame has been encoded in the first mode or in the second mode or in the third mode or in the fourth mode.

[0305] With respect to the decoder, according to embodiments, the decoder 200 can for example be operable in the first mode and can for example be operable in at least one of the second mode and the third mode and the fourth mode.

[0306] If the decoder 200 is in the first mode, the decoder 200 can for example be configured to determine an estimate of two harmonic parameters for each of one or more harmonic components of a most previous frame, wherein the two harmonic parameters for each of the one or more harmonic components of the most previous frame depend on a first group of three or more spectral coefficients of a plurality of reconstructed spectral coefficients of each of one or more previous frames of the audio signal, and the decoder 200 can for example be configured to decode the encoding of the current frame in dependence on the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.

[0307] If the decoder 200 is in the second mode, the decoder 200 can for example be configured to parse the encoding of the audio signal 120 to obtain encoded spectral coefficients 206_t0_f1:206_t0_f6; 206_t-1_f1:206_t-1_f6 of the audio signal 120 for the current frame 208_t0 and at least for at least the most previous frame 208_t-1, and the decoder 200 can for example be configured to selectively apply a predictive decoding to the plurality of individual encoded spectral coefficients 206_t0_f2 or the group of encoded spectral coefficients 206_t0_f4, 206_t0_f5, wherein the decoder 200 can for example be configured to obtain a pitch value, wherein the decoder 200 can for example be configured to select the plurality of individual encoded spectral coefficients 206_t0_f2 or the group of encoded spectral coefficients 206_t0_f4, 206_t0_f5 to which the predictive decoding can for example be applied based on the pitch value.

[0308] If the decoder 200 is in the third mode, the decoder 200 can for example be configured to decode the audio signal by employing a time-domain long-term prediction.

[0309] If the decoder 200 is in the fourth mode, the decoder 200 can for example encode the audio signal by employing an adaptive modified discrete cosine transform long-term prediction, wherein, if the decoder 200 employs the adaptive modified discrete cosine transform long-term prediction, the decoder 200 can for example be configured to select on a frame basis as a prediction method either a time-domain long-term prediction or a frequency-domain prediction or a frequency-domain minimum mean square prediction in dependence on a minimization criterion.

[0310] According to embodiments, in each of the first mode and the second mode and the third mode and the fourth mode, the decoder 200 can for example be configured to decode the audio signal in dependence on a refined fundamental frequency and an adapted gain factor, wherein the refined fundamental frequency and the adapted gain factor have been determined on a frame basis.

[0311] In an embodiment, the decoder 200 may, for example, receive and decode an encoding comprising an indication that the current frame has been encoded in the first mode or in the second mode or in the third mode or in the fourth mode. The decoder 200 can set itself to the first mode or to the second mode or to the third mode or to the fourth mode depending on the indication.

[0312] In Figure 5 It can be seen that the BS of all three concepts drops significantly for flute notes when the frame length increases, as the redundancy in the original signal has been largely removed by the transform itself. The performance of the FDP can drop significantly for low-pitched bass notes due to the highly overlapping harmonics on the MDCT coefficients. The performance of the TDLTP is generally good. However, it can drop when the frame length is large, where a large delay is needed to find a matching previous pitch period. The FDLMSP provides relatively good and stable performance for different notes and different frame lengths. Figure 5 It is also shown that the BS drops when the prediction bandwidth is increased to 8 kHz, due to the inharmonicity of the tones in the higher frequency band. As the inharmonicity depends on the spectral characteristics of each individual sound material, a pre-computation and comparison of the bit rate consumption can be done over the frequency bands to obtain higher coding efficiency. A prediction decision can then be made and signalled as side information in each frame.

[0313] Figure 6 The bit rate savings in four different operating modes are shown on six different projects with a bandwidth limit of 4 kHz, MDCT frame length of 64 and 512.

[0314] As Figure 6 shown, the FDLMSP outperforms the TDLTP and the FDP in many scenarios and provides good performance overall. The AMLTP performs the best and selects the FDLMSP or the TDLTP in most cases, indicating that the FDLMSP can be combined with the TDLTP to greatly enhance the BS.

[0315] A novel approach for LTP in the MDCT domain has been provided. The novel approach models each MDCT frame as a hypothesis of harmonic components and estimates the parameters of all the harmonic components from previous frames using the LMS concept. Prediction is then made based on the estimated harmonic parameters. The approach provides competitive performance compared to similar concepts and can also be used jointly to improve audio coding efficiency.

[0316] The above concepts can be used, for example, to analyze the influence of pitch information accuracy on the prediction, e.g. by using different pitch estimation algorithms or by applying different quantization steps. The above concepts can also be used to determine or refine the pitch information of an audio signal on a frame basis using a minimization criterion. For example, the influence of inharmonicity and other complex signal characteristics on the prediction can be considered. The above concepts can be used, for example, for error concealment.

[0317] While some aspects have been described in the context of an apparatus, it is clear that separate aspects also correspond to a description of a corresponding method, where a block or apparatus corresponds to a method step or a feature of a method step. Analogously, aspects described in the context of a method step also represent a description of a corresponding block or item or feature of a corresponding apparatus. Some or all of the method steps can be executed by (or using) a hardware apparatus, like for example, a microprocessor, a programmable computer or electronic circuit. In some embodiments, one or more of the most important method steps can be executed by such an apparatus.

[0318] Depending on certain implementation requirements, embodiments of the application can be implemented in hardware or in software, or in a combination of hardware and software. The implementation can be performed using a digital storage medium, for example a floppy disk, a DVD, a Blu-Ray, a CD, a ROM, a PROM, an EPROM, an EEPROM or a FLASH memory, having electronically readable control signals stored thereon, which cooperate (or are capable of cooperating) with a programmable computer system such that the respective method is performed. Therefore, the digital storage medium can be computer readable.

[0319] Some embodiments according to the application comprise a data carrier having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein is performed.

[0320] Generally, embodiments of the present application can be implemented as a computer program product with a program code, the program code being operative for performing one of the methods when the computer program product runs on a computer. The program code can for example be stored on a machine readable carrier.

[0321] A further embodiment comprises a computer having installed thereon the computer program according to the application.

[0322] In other words, an embodiment of the inventive method is, therefore, a computer program having a program code for performing one of the methods described herein, when the computer program runs on a computer.

[0323] A further embodiment of the inventive method is, therefore, a data carrier (or a digital storage medium, or a computer-readable medium) having electronically readable control signals, which are capable of cooperating with a programmable computer system, such that one of the methods described herein can be performed.

[0324] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein. The data stream or the sequence of signals can for example be configured to be transferred via a data communication connection, for example via the Internet.

[0325] A further embodiment comprises a processing means, such as a computer, or a programmable logic device, configured to, or adapted to, perform one of the methods described herein.

[0326] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0327] A further embodiment comprises an apparatus or a system configured to transfer (for example, electronically or optically) a computer program for performing one of the methods described herein to a receiver. The receiver may, for example, be a computer, a mobile device, a memory device or the like. The apparatus or system may, for example, comprise a file server for transferring the computer program to the receiver.

[0328] In some embodiments, a programmable logic device (for example a field programmable gate array) can be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor in order to perform one of the methods described herein. Generally, the methods are preferably performed by any hardware apparatus.

[0329] The apparatuses described herein can be implemented using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0330] The methods described herein can be performed using a hardware apparatus, or using a computer, or using a combination of hardware and computer.

[0331] The above-described embodiments are merely illustrative for the principles of the present application. It is understood that modifications and variations of the arrangements and the details described herein will be apparent to others skilled in the art. It is the intent, therefore, to be limited only by the scope of the appended patent claims and not by the specific details presented by way of description and explanation of the embodiments herein.

[0332] References:

[0333] [1] Jürgen Herre and Sascha Dick, "Psychoacoustic models for perceptual audio coding a tutorial review," Applied Sciences, vol. 9, pp. 2854, ITT 2019.

[0334] [2] Juha Mauri and Lin Yin, "Long Term Predictor for Transform Domain Perceptual Audio Coding," in Audio Engineering Society Convention 107, Sep 1999.

[0335] [3] Hendrik Fuchs, "Improving mpeg audio coding by backward adaptive linear stereo prediction," in Audio Engineering Society Convention 99, Oct 1995.

[0336] [4] J. Princen, A. Johnson, and A. Bradley, "Subband / transform coding using filter bank designs based on time domain aliasing cancellation," in ICASSP '87. IEEE International Conference on Acoustics, Speech, and Signal Processing, April 1987, vol. 12, pp. 2161-2164.

[0337] [5] Christian Helmrich, Efficient Perceptual Audio Coding Using Cosine and Sine Modulated Lapped Transforms, doctoral thesis, Friedrich-Alexander- Erlangen-Nürnberg (FAU), 2017, Chapter 3.3: Frequency-Domain Prediction with Very Low Complexity.

[0338] [6] J. Rothweiler, "Polyphase quadrature filters - a new subband coding technique," in ICASSP'83. IEEE International Conference on Acoustics, Speech, and Signal Processing, April 1983, vol. 8, pp. 1280-1283.

[0339] [7] Albrecht Schneider and Klaus Frieler, "Perception of harmonic and inharmonic sounds: Results from ear models," in Computer Music Modeling and Retrieval. Genesis of Meaning in Sound and Music, Ystad, Richard Kronland-Martinet, and Kristoffer Jensen, Eds., Berlin, Heidelberg, 2009, pp. 18-44, Springer Berlin Heidelberg.

[0340] [8] Hugo Fastl and Eberhard Zwicker, Psychoacoustics: Facts and Models, Springer- Verlag, Berlin, Heidelberg, 2006, Chapter 7.2: Just-Noticeable Changes in Frequency.

[0341] [9] John P. Princen and Alan Bernard Bradley, "Analysis / synthesis filterbank design based on time domain aliasing cancellation," IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 34, no. 5, pp. 1153-1161, October 1986.

[0342]

[10] Alain de Cheveign and Hideki Kawahara, "Yin, a fundamental frequency estimator for speech and music," The Journal of the Acoustical Society of America, vol. 111, pp. 1917-30, 05 2002.

[0343]

[11] Armin Taghipour, Psychoacoustics of detection of tonality and asymmetry of masking: implementation of tonality estimation methods in a psychoacoustic model for perceptual audio coding, doctoral thesis, Friedrich-Alexander- Universität Erlangen-Nürnberg (FAU), 2016, Chapter 4: The Psychoacoustic model.

[0344]

[12] J. D. Johnston, "Estimation of perceptual entropy using noise masking criteria," in ICASSP-88, International Conference on Acoustics, Speech, and Signal Processing, April 1988, pp. 2524--2527 vol. 5.

[0345]

[13] WO 2016 142357 Al, published September 2016.

Claims

1. An encoder (100) for encoding a current frame of an audio signal from one or more previous frames of the audio signal, wherein the one or more previous frames precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, characterized in that In order to generate an encoding of the current frame, the encoder (100) is to determine an estimate of two harmonic parameters for each of one or more harmonic components of a most previous frame of the one or more previous frames, the two harmonic parameters being a first parameter for a cosine subcomponent and a second parameter for a sine subcomponent, wherein the encoder (100) is to determine the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame using a first group of spectral coefficients of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames of the audio signal; wherein, in order to generate the encoding of the current frame, the encoder (100) is to determine a gain factor and a residual signal as the encoding of the current frame from a fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein the encoder (100) is to generate the encoding of the current frame such that the encoding of the current frame comprises the gain factor and the residual signal.

2. The encoder (100) of claim 1, wherein The encoder (100) is to determine the estimate of the two harmonic parameters for each of the one or more harmonic components of the current frame from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame and from a fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames.

3. The encoder (100) of claim 1, wherein The encoder (100) is to estimate the two harmonic parameters for each of the one or more harmonic components of the most previous frame by solving a linear equation system comprising at least three equations, wherein each of the at least three equations depends on a spectral coefficient of a first group of spectral coefficients of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames.

4. The encoder (100) of claim 3, wherein The encoder (100) is to solve the linear equation system using a least mean square algorithm.

5. The encoder (100) of claim 3, wherein The linear equation system is defined by wherein, wherein, a first spectral band indicative of a harmonic component of the one or more harmonic components of the most previous frame having a lowest harmonic component frequency of the one or more harmonic components, wherein, a second spectral band indicative of a harmonic component of the one or more harmonic components of the most previous frame having a highest harmonic component frequency of the one or more harmonic components, wherein, r is an integer, r > 0; wherein indicating the first matrix, , wherein indicating a second matrix, .

6. The encoder (100) of claim 5, wherein r≥1。 7. The encoder (100) of claim 3, wherein, The linear equation system is solvable according to wherein, is a first vector comprising an estimate of two harmonic parameters for each of one or more harmonic components of the most recent frame, wherein, is a second vector comprising a first set of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames, wherein is the Moore-Penrose inverse matrix of wherein, comprises a plurality of third matrices or third vectors, wherein each of the third matrix or the third vector indicates, together with the estimate of the two harmonic parameters of a harmonic component of the one or more harmonic components of the most previous frame, an estimate of the harmonic component, wherein H indicates the number of harmonic components of the one or more previous frames.

8. The encoder (100) according to claim 1, wherein the encoder (100) encodes the fundamental frequency of a harmonic component, a window function, the gain factor and the residual signal.

9. The encoder (100) according to claim 8, wherein the encoder (100) determines the number of the one or more harmonic components of the most previous frame before estimating the two harmonic parameters of each of the one or more harmonic components of the most previous frame using a first group of spectral coefficients consisting of three or more spectral coefficients of a plurality of spectral coefficients of each of the one or more previous frames of the audio signal.

10. The encoder (100) according to claim 9, wherein the encoder (100) determines one or more groups of harmonic components from the one or more harmonic components and applies a prediction of an audio signal on the one or more groups of harmonic components, wherein the encoder (100) encodes an order of each of the one or more groups of harmonic components of the most previous frame.

11. The encoder (100) according to claim 1, wherein the encoder (100) determines the two harmonic parameters of each of the one or more harmonic components of the current frame from the two harmonic parameters of each of the harmonic components of the one or more harmonic components of the most previous frame.

12. The encoder (100) according to claim 11, wherein the encoder (100) applies: and wherein the encoder (100) applies: , wherein a h is a parameter of a cosine subcomponent of the hth harmonic component of the one or more harmonic components for the most previous frame, wherein b h is a parameter of a sinusoidal subcomponent of the hth harmonic component of the one or more harmonic components for the most previous frame, wherein c h is a parameter of a cosine subcomponent of the hth harmonic component of the one or more harmonic components for the current frame, wherein d h is a parameter of a sinusoidal subcomponent of the h-th harmonic component of the one or more harmonic components of the current frame, wherein N depends on a length of a transform block used for transforming a time domain audio signal into the frequency domain or spectral domain, and wherein , wherein f0is a fundamental frequency of the one or more harmonic components of the most previous frame and f0is a fundamental frequency of the one or more harmonic components of the current frame, where f s is the sampling frequency, and wherein h is an index indicating a harmonic component of the one or more harmonic components of the most previous frame.

13. The encoder (100) according to claim 1, wherein the encoder (100) determines the residual signal from a plurality of spectral coefficients of the current frame in the frequency domain or the transform domain and from the estimate of the two harmonic parameters of each of the one or more harmonic components of the current frame, and wherein the encoder (100) encodes the residual signal.

14. The encoder (100) according to claim 13, wherein the encoder (100) determines a spectral prediction of one or more spectral coefficients of the plurality of spectral coefficients of the current frame from the estimate of the two harmonic parameters of each of the one or more harmonic components of the current frame, and wherein the encoder (100) is configured to determine the residual signal and a gain factor from a plurality of spectral coefficients in the frequency domain or the transform domain according to the current frame and from a spectral prediction of three or more spectral coefficients of the plurality of spectral coefficients of the current frame, wherein the encoder (100) is configured to encode an order of each of one or more groups of harmonic components of the most previous frame.

15. The encoder (100) according to claim 14, wherein the encoder (100) is configured to determine the residual signal of the current frame according to wherein m is a frame index, wherein k is a frequency index, wherein N depends on a length of a transform block used for transforming a time domain audio signal into the frequency domain or spectral domain, wherein, indicating a kth sample of the residual signal in the spectral domain or in the transform domain, wherein, indicates the kth sample of the spectral coefficients of the current frame in the spectral domain or in the transform domain, wherein, indicates the kth sample in the spectral domain or the transform domain of the spectral prediction of the current frame, and wherein g is the gain factor.

16. The encoder (100) according to claim 1, wherein, the encoder (100) is operable in a first mode and is operable in at least one of a second mode, a third mode and a fourth mode, wherein, if the encoder (100) is in the first mode, the encoder (100) is configured to encode the current frame by determining an estimate of two harmonic parameters of each of one or more harmonic components of the most previous frame using a first group of spectral coefficients consisting of three or more spectral coefficients of a plurality of spectral coefficients of each of the one or more previous frames of the audio signal, wherein, if the encoder (100) is in the second mode, the encoder (100) is configured to encode the audio signal in the transform domain or in a filter bank domain and the encoder is configured to determine a plurality of spectral coefficients (106_t0_f1:106_t0_f6; 106_t-1_f1:106_t-1_f6) of the audio signal (102) for the current frame (108_t0) and at least for the most previous frame (108_t-1), wherein the encoder (100) is configured to selectively apply predictive encoding to a plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5), the encoder (100) is configured to determine a pitch value, the encoder (100) is configured to select the plurality of individual spectral coefficients (106_t0_f2) or groups of spectral coefficients (106_t0_f4, 106_t0_f5) to which predictive encoding is applied based on the pitch value, wherein, if the encoder (100) is in the third mode, the encoder (100) is configured to encode the audio signal by employing time domain long term prediction, and wherein, if the encoder (100) is in the fourth mode, the encoder (100) encodes the audio signal by employing adaptive improved discrete cosine transform long-term prediction, wherein, if the encoder (100) employs adaptive improved discrete cosine transform long-term prediction, the encoder (100) is configured to select, on a frame basis, a time-domain long-term prediction or a frequency-domain prediction or a frequency-domain minimum mean square prediction as a prediction method according to a minimization criterion.

17. The encoder (100) according to claim 16, wherein In each of the first mode, the second mode, the third mode and the fourth mode, the encoder (100) is to refine a fundamental frequency on a frame basis according to a minimization criterion to obtain a refined fundamental frequency and to adapt a gain factor to obtain an adapted gain factor, wherein the encoder (100) is to encode the refined fundamental frequency and the adapted gain factor instead of an original fundamental frequency and gain factor.

18. The encoder (100) according to claim 16, wherein the encoder (100) is to set itself in the first mode or in the second mode and in at least one of the third mode and the fourth mode, and wherein the encoder (100) is to encode irrespective of whether the current frame was encoded in the first mode or in the second mode or in the third mode or in the fourth mode.

19. A decoder (200) for reconstructing a current frame of an audio signal, wherein one or more previous frames of the audio signal precede the current frame, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, characterized in that the decoder (200) is to receive an encoding of the current frame, the encoding of the current frame comprising a gain factor and a residual signal, wherein the decoder (200) is to reconstruct the current frame from the gain factor, from the residual signal and from a fundamental frequency of the one or more harmonic components of the current frame and one or more previous frames, wherein the decoder (200) is to determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, the two harmonic parameters being a first parameter for a cosine subcomponent and a second parameter for a sine subcomponent, wherein the two harmonic parameters of each of the one or more harmonic components of the most previous frame depend on a first group of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal, wherein the decoder (200) is to reconstruct the current frame from the encoding of the current frame and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame.

20. The decoder (200) according to claim 19, wherein The two harmonic parameters of each of the one or more harmonic components of the most previous frame depend on a system of linear equations comprising at least three equations, wherein each of the at least three equations depends on spectral coefficients of a first group of three or more spectral coefficients of a plurality of reconstructed spectral coefficients of each of the one or more previous frames.

21. The decoder (200) according to claim 20, wherein The system of linear equations can be solved using a least mean square algorithm.

22. The decoder (200) according to claim 20, wherein The system of linear equations is defined by: wherein, wherein, a first spectral band indicative of a harmonic component of the one or more harmonic components of the most previous frame having a lowest harmonic component frequency of the one or more harmonic components, wherein, a second spectral band indicative of a harmonic component of the one or more harmonic components of the most previous frame having a highest harmonic component frequency of the one or more harmonic components, wherein r is an integer, r > 0, wherein indicating the first matrix, , wherein indicating a second matrix, .

23. The decoder (200) of claim 22, wherein r≥1。 24. The decoder (200) according to claim 20, wherein, The system of linear equations can be solved according to: wherein, is a first vector comprising an estimate of two harmonic parameters for each of one or more harmonic components of the most recent frame, wherein, is a second vector comprising a first set of three or more spectral coefficients of the plurality of reconstructed spectral coefficients of each of the one or more previous frames, wherein is the Moore-Penrose inverse matrix of wherein, comprises a plurality of third matrices or third vectors, wherein each of the third matrix or third vector indicates, together with an estimate of two harmonic parameters of a harmonic component of the one or more harmonic components of the most previous frame, an estimate of the harmonic component, wherein H indicates a number of harmonic components of the one or more previous frames.

25. The decoder (200) according to claim 19, wherein, The decoder (200) is to receive a fundamental frequency of a harmonic component, a window function, the gain factor and the residual signal, wherein the decoder (200) is to reconstruct the current frame from the fundamental frequency of the one or more harmonic components of the most previous frame, from the window function, from the gain factor and from the residual signal.

26. The decoder (200) according to claim 25, wherein The decoder (200) is to receive a number of the one or more harmonic components of the most previous frame, and wherein the decoder (200) is to decode the encoding of the current frame from the number of the one or more harmonic components of the most previous frame.

27. The decoder (200) according to claim 26, wherein The decoder (200) is to decode the encoding of the current frame from one or more groups of harmonic components, wherein the decoder (200) is to apply a prediction of the audio signal on the one or more groups of harmonic components.

28. The decoder (200) according to claim 19, wherein The decoder (200) is to determine two harmonic parameters of each of the one or more harmonic components of the current frame from two harmonic parameters of each of the harmonic components of the one or more harmonic components of the most previous frame.

29. The decoder (200) according to claim 28, wherein The decoder (200) is to apply: and wherein the decoder (200) is to apply: , wherein a h is a parameter for a cosine subcomponent of the hth harmonic component of the one or more harmonic components of the most previous frame, wherein b h is a parameter of a sinusoidal subcomponent of the hth harmonic component of the one or more harmonic components of the most previous frame, wherein c h is a parameter for a cosine subcomponent of the hth harmonic component of one or more harmonic components of the current frame, wherein d h is a parameter of a sinusoidal subcomponent of the h-th harmonic component of one or more harmonic components of the current frame, wherein N depends on a length of a transform block used for transforming a time domain audio signal into the frequency domain or spectral domain, and wherein, , wherein f0is a fundamental frequency of the one or more harmonic components of the most previous frame and f0is a fundamental frequency of the one or more harmonic components of the current frame, where f s is the sampling frequency, and wherein h is an index indicating one of the one or more harmonic components of the most previous frame.

30. The decoder (200) according to claim 19, wherein The decoder (200) will receive the residual signal, wherein the residual signal depends on a plurality of spectral coefficients of the current frame in the frequency domain or in the transform domain, and wherein the residual signal depends on an estimate of two harmonic parameters of each of one or more harmonic components of the current frame.

31. The decoder (200) of claim 30, wherein The decoder (200) will determine a spectral prediction of one or more spectral coefficients of the current frame from the estimate of two harmonic parameters of each of one or more harmonic components of the current frame, and wherein the decoder (200) will determine the current frame of the audio signal from the spectral prediction of the current frame and from the residual signal and from the gain factor.

32. The decoder (200) of claim 31, wherein The residual signal of the current frame is defined according to: wherein m is a frame index, wherein k is a frequency index, wherein is the received quantized reconstruction residual, wherein, is the reconstructed current frame, wherein, indicates a spectral prediction in the spectral domain or in the transform domain of the current frame, and wherein g is the gain factor.

33. The decoder (200) of claim 19, wherein The decoder (200) is operable in a first mode and is operable in at least one of a second mode, a third mode and a fourth mode, wherein, if the decoder (200) is in the first mode, the decoder (200) will determine the estimate of two harmonic parameters of each of one or more harmonic components of the most previous frame, wherein the two harmonic parameters of each of one or more harmonic components of the most previous frame depend on a first group of three or more spectral coefficients of a plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal, and the decoder (200) will decode the encoding of the current frame from the estimate of two harmonic parameters of each of one or more harmonic components of the most previous frame, wherein, if the decoder (200) is in the second mode, the decoder (200) will parse the encoding of the audio signal (120) to obtain encoded spectral coefficients (206_t0_f1:206_t0_f6; 206_t-1_f1:206_t-1_f6) of the audio signal (120) for the current frame (208_t0) and at least for the most previous frame (208_t-1), and the decoder (200) is configured to selectively apply a prediction decoding to a plurality of individual encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5), wherein the decoder (200) is configured to obtain a pitch value, wherein the decoder (200) is configured to select the plurality of individual encoded spectral coefficients (206_t0_f2) or groups of encoded spectral coefficients (206_t0_f4, 206_t0_f5) to which to apply the prediction decoding based on the pitch value, wherein, if the decoder (200) is in the third mode, the decoder (200) will decode the audio signal by employing time domain long-term prediction, and wherein, if the decoder (200) is in the fourth mode, the decoder (200) will decode the audio signal by employing adaptive modified discrete cosine transform long-term prediction, wherein, if the decoder (200) employs adaptive modified discrete cosine transform long-term prediction, the decoder (200) is configured to select, on a frame basis, time domain long-term prediction or frequency domain prediction or frequency domain minimum mean square prediction as the prediction method depending on a minimization criterion.

34. The decoder (200) according to claim 33, wherein In each of the first mode, the second mode, the third mode and the fourth mode, the decoder (200) will decode the audio signal depending on a refined fundamental frequency and depending on an adapted gain factor, wherein the refined fundamental frequency and the adapted gain factor have been determined on a frame basis.

35. The decoder (200) according to claim 33, wherein The decoder (200) will receive and decode an encoding, the encoding comprising an indication that the current frame has been encoded in the first mode or in the second mode or in the third mode or in the fourth mode, and wherein the decoder (200) will set itself to the first mode or to the second mode or to the third mode or to the fourth mode depending on the indication.

36. An apparatus (700) for frame loss concealment, wherein one or more previous frames of an audio signal preceding a current frame of the audio signal, wherein each of the current frame and the one or more previous frames comprises one or more harmonic components of the audio signal, wherein each of the current frame and the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, wherein the apparatus (700) will determine an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, wherein the two harmonic parameters of each of the one or more harmonic components of the most previous frame depend on a first group of three or more spectral coefficients out of the plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal, wherein, if the apparatus (700) does not receive the current frame or if the apparatus (700) receives the current frame in a corrupted state, the apparatus (700) will reconstruct the current frame depending on the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein, for reconstructing the current frame, the apparatus (700) will determine an estimate of two harmonic parameters for each of the one or more harmonic components of the current frame depending on the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein the apparatus (700) is to determine the two harmonic parameters for each of the one or more harmonic components of the current frame from the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein the apparatus (700) is to apply: and wherein the apparatus (700) is to apply: , wherein a h is a parameter of a cosine subcomponent of the hth harmonic component of the one or more harmonic components of the most previous frame, where b h is a parameter of a sinusoidal subcomponent of the hth harmonic component of the one or more harmonic components of the most previous frame, where c h is a parameter for a cosine subcomponent of the hth harmonic component of one or more harmonic components of the current frame, where d h is a parameter of a sinusoidal subcomponent of the hth harmonic component of one or more harmonic components of the current frame, wherein N depends on a length of a transform block used for transforming a time-domain audio signal into the frequency domain or spectral domain, and wherein , wherein f0is a fundamental frequency of the one or more harmonic components of the most previous frame, and f0is a fundamental frequency of the one or more harmonic components of the current frame, where f s is the sampling frequency, and wherein h refers to an index indicating one of the one or more harmonic components of the most previous frame.

37. A system for encoding a current frame of an audio signal and for reconstructing the current frame of the audio signal, comprising: an encoder (100) according to claim 1 for encoding the current frame of the audio signal, and a decoder (200) according to claim 19 for decoding the encoding of the current frame of the audio signal.

38. A method for encoding a current frame of an audio signal from one or more previous frames of the audio signal, wherein the one or more previous frames precede the current frame, wherein the current frame and each of the one or more previous frames comprises one or more harmonic components of the audio signal, wherein the current frame and each of the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, characterized in that for generating an encoding of the current frame, the method comprises determining an estimate of two harmonic parameters for each of the one or more harmonic components of a most previous frame of the one or more previous frames, the two harmonic parameters being a first parameter for a cosine subcomponent and a second parameter for a sine subcomponent, wherein determining the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame is performed using a first group of spectral coefficients of three or more spectral coefficients of the plurality of spectral coefficients of each of the one or more previous frames of the audio signal, wherein for generating the encoding of the current frame, the method comprises determining a gain factor and a residual signal as the encoding of the current frame from a fundamental frequency of the one or more harmonic components of the current frame and of the one or more previous frames and from the estimate of the two harmonic parameters for each of the one or more harmonic components of the most previous frame, wherein the encoding of the current frame is generated such that it comprises the gain factor and the residual signal.

39. A method for reconstructing a current frame of an audio signal, wherein, the one or more previous frames of the audio signal precede the current frame, wherein the current frame and each of the one or more previous frames comprises one or more harmonic components of the audio signal, wherein the current frame and each of the one or more previous frames comprises a plurality of spectral coefficients in a frequency domain or in a transform domain, characterized in that the method comprises receiving an encoding of the current frame, the encoding of the current frame comprising a gain factor and a residual signal, wherein the current frame is to be reconstructed from the gain factor, from the residual signal and from a fundamental frequency of one or more harmonic components of the current frame and one or more previous frames, wherein the method comprises determining an estimate of two harmonic parameters of each of one or more harmonic components of a most previous frame of the one or more previous frames, the two harmonic parameters being a first parameter for a cosine subcomponent and a second parameter for a sine subcomponent, wherein the two harmonic parameters of each of the one or more harmonic components of the most previous frame depend on a first group of three or more spectral coefficients of a plurality of reconstructed spectral coefficients of each of the one or more previous frames of the audio signal, wherein the method comprises reconstructing the current frame from the encoding of the current frame and from the estimate of the two harmonic parameters of each of the one or more harmonic components of the most previous frame.

40. A non-transitory computer readable medium comprising a computer program for implementing the method of claim 38 or 39 when executed by a computer or signal processor.

Citation Information

Patent Citations

  • Audio encoder and decoder

    CN105247614A

  • Audio encoder, audio decoder, method for encoding an audio signal and method for decoding an encoded audio signal

    WO2016142357A1